<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scaling Up Bibliographic Data Science</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mikko Tolon</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Helsinki</institution>
          ,
          <addr-line>Helsinki</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Turku</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <fpage>450</fpage>
      <lpage>456</lpage>
      <abstract>
        <p>Bibliographic data science is an emerging research paradigm in digital humanities. It aims at systematic quanti cation of the trends in knowledge production based on large-scale analysis of bibliographic metadata collections and the methods of modern data science. Compared to the earlier related attempts in book history and sociology of literature, advances in data processing and quality control are now making it possible for the rst time to scale up the analysis to millions of print products while at the same time paying attention to data quality, representativity and completeness. This provides a new quantitative method that can support the analysis of classical research questions in intellectual history. Here, we discuss the methodological challenges that we have encountered in such studies and how to scale up the solutions based on collaborative research e orts.</p>
      </abstract>
      <kwd-group>
        <kwd>Library catalogues history</kwd>
        <kwd>Open science</kwd>
        <kwd>Data science ecosystem</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Bibliographic data collections contain a vast array of data on publishing
activities, and they have been traditionally used as a tool for information retrieval.
A systematic analysis of bibliographic catalogues can generate rich information
on historical patterns in knowledge production, and their research potential has
been debated for at least half centuries now [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Large-scale integration of data
from bibliographies and supporting information sources across long time periods,
geographical areas, and genres has the potential to shed new light on classical
hypotheses as well as to uncover previously overlooked historical trends. Such
analyses depend critically not only on data quality and completeness but also
on understanding the historical context. Hence, seamless collaboration between
data scientists and historians is crucial for obtaining robust conclusions and for
developing e cient research methods for such analysis. Whereas earlier studies in
analytical bibliography and related elds have discussed these opportunities,
emphasized the role of quanti cation [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref17 ref4 ref6">12, 13, 17, 6, 11, 4</xref>
        ] and charted the long-term
developments in the history in books [
        <xref ref-type="bibr" rid="ref1 ref2 ref5 ref7">5, 1, 2, 7</xref>
        ], systematic quantitative research
use of large bibliographic metadata collections has proven to be challenging.
We recently proposed the concept of bibliographic data science (BDS) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and
provided the rst case studies to demonstrate the research potential of this
approach. In this paper we complement the previous work by discussing practical
solutions to scaling up these e orts in order to integrate data across dozens of
bibliographic collections that comprise altogether to millions of print products
over several centuries, genres, and languages.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methodological aspects</title>
      <p>
        Bibliographic data science (BDS) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] facilitates research use of library catalogues
by developing systematic quantitative methods to ensure data quality, and
support analysis and interpretation. It has been suggested that scaling up
bibliographic data science to cover dozens of catalogues and millions of entries could
substantially expand the depth and scope of such studies but this depends on our
ability to develop scalable solutions for reliable data harmonization and analysis.
Such e orts can remarkably bene t from the latest developments in open data
science (see e.g. [
        <xref ref-type="bibr" rid="ref3 ref8">8, 3</xref>
        ]). Compared to other data science projects with related
challenges, unique to our e orts is the central role of historical interpretation
and the necessity of incorporating prior knowledge on the data collection
processes which may introduce unexpected biases in the analyses. Here, we brie y
discuss some of the key elements that can help to scale up bibliographic data
science. These include automation, standardization, use of supporting data sources,
quality monitoring, machine learning, and open collaboration models.
      </p>
      <p>Harmonization of the original records is the rst step in facilitating reliable
research use. Bibliographic metadata is often manually entered in the databases,
and seldom su ciently standardized and readily amenable to quantitative
analysis. Biases, inaccuracies, gaps, and varying standards and languages pose
challenges for data integration both within and across catalogues. The heterogeneity
of the data elds, spanning from time intervals to persons, physical dimensions,
or geographical locations is also a notable challenge. The scale of these issues
greatly exceeds our ability to manually verify and correct the entries. Essential
for this undertaking is to move towards automated analysis work ows that
enable standardized treatment of the data and e cient implementation of similar
work ows across multiple catalogues and iterative data corrections.</p>
      <p>
        The integration of data across catalogues enables the analysis of publishing
activity beyond what is accessible by the use of individual national bibliographies
alone, as we have recently suggested in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Since all bibliographies follow the
conventions that are described in the MARC cataloguing format, we have been
able to apply largely identical processing across all collections, thus facilitating
transparent and scalable data processing. We have extensively harmonized
selected elds of the Finnish and Swedish National Bibliographies (FNB and SNB,
respectively), the English Short-Title Catalogue (ESTC), and the Heritage of the
Printed Book database (HPBD). Altogether, these four bibliographies cover over
6 million entries of print products printed in Europe and elsewhere, and 2.64
million harmonized entries in the period 1500-1800, ranging from the 16365 entries
in the FNB to 2.1 million entries in HPBD, which is a compilation of 45 smaller,
mostly national, bibliographies3.
      </p>
      <p>Standardization and reproducibility can bene t from the design of
dedicated software packages and unit tests. Hence, we have implemented
systematic algorithmic approaches that help to assess and improve data reliability in a
semi-automated fashion. These work ows combine various procedures to remove
spelling errors, disambiguate and standardize terms, augment missing values,
and incorporate manually curated information such as known pseudonymes or
synonymous entries. Some steps are straightforward, including removal of extra
periods or spaces; in other cases, custom algorithms have to be developed. One
example is the page count information, which follows a speci c MARC notation4
that separates paginations in di erent parts of the document, such as preface,
contents, gures and tables, and in separate volumes. Speci cally designed
algorithms are necessary for scalable interpretation of such information.</p>
      <p>Remarkable portions of information is typically missing in library catalogues
but can be readily augmented based on information that is already contained
elsewhere in the catalogue, or based on external information sources. For
instance, missing country information can be lled in when the publication place
is available and uniquely mappable to a country; and exact physical dimensions
of a document can be often inferred with a good accuracy based on the
gatherings information that is more frequently available. In such cases, information
is however retained whether the value is based on the original entries, or later
estimated based on other information. We have also complemented the
original records with entirely new derived elds, such as print area, which quanti es
the paper consumption per unique physical copy of a document, or title. The
combined print area across all unique titles then re ects the overall breadth of
printing activities per time period, genre, or geographic location, rather than the
mere volume of print products. Hence, it complements total paper consumption
which quanti es the overall volume of print products but does not consider the
uniqueness or diversity across titles. Estimating the overall paper consumption
is more di cult as it requires information on print run sizes and this is often not
readily available. The estimated print run size of one thousand copies is often
used for early-modern books but notable variations over time and geography
have been reported. The print area and paper consumption can hence provide
complementary ways to assess trends in knowledge production.</p>
      <p>Monitoring the quality of the harmonization process is an important part of
such e orts. We are routinely generating conversion tables that show how the
original raw entries have been converted into the nal harmonized versions, list
of common entries in each eld (e.g. authors, publishers, languages), and
statistical summaries such as average document dimensions by format, histograms
of publication years and author life spans, or changes in gender distribution</p>
      <sec id="sec-2-1">
        <title>3 https://www.cerl.org/resources/hpb/content</title>
      </sec>
      <sec id="sec-2-2">
        <title>4 see the Library of Congress web document https://www.loc.gov/marc/bibliographic/</title>
        <p>for the full description of the MARC21 format.
over time. Such overviews have been invaluable in spotting unexpected events
that may in turn facilitate the detection of remaining inaccuracies or biases in
the harmonized data sets. The quality monitoring has been greatly facilitated
by such reproducible overviews, and the most up-to-date summaries for each
bibliographic catalogue that we are working on can be accessed via Helsinki
Computational History Group website5. By openly sharing the summaries and
algorithms, we are aiming to improve the transparency, reliability, and overall
quality of our work by providing the broader audience with the means to detect
and report potential inaccuracies.</p>
        <p>Advances in machine learning and arti cial intelligence are providing further
means to scale up the analysis. We are routinely utilizing methods from natural
language processing, feature selection, clustering, and classi cation to facilitate
duplicate identi cation and quality monitoring. Information in the publisher
eld, for instance, often contains lengthy verbal explanations that have to be
subjected to named entity recognition. Moreover, detecting alternative spellings for
author or publisher names can be greatly accelerated by string distance
matching algorithms that can rank potential duplicates and facilitate semi-automated
data curation. Page number estimation is a prominent example of the potential
of this approach, as the original text entries can be converted into a single
wellde ned number, which is the nal page count estimate. The overall accuracy of
such fully automated estimates can be quanti ed to a great extent by examining
the most common conversions and representative sets of random examples.
Future developments could take increasing advantage of such statistical techniques
in order to reduce the need for human input, thus improving the overall
scalability of data harmonization. The already curated data sets provide ample training
material for supervised machine learning techniques.</p>
        <p>Finally, the research community can bene t from open sharing of the data
and algorithms. The lack of open data availability is forming a major
bottleneck for collaborative development of bibliographic data science but this might
be gradually changing. The National Library of Finland, for instance, recently
released the complete MARC entries of the FNB6 under an open data license
that allows the modi cation, reuse, and sharing of derivative versions. We have
made the key algorithms for harmonization and analysis openly available in the
bibliographica R package7, and our harmonized versions of the FNB data set
can be accessed via Helsinki Computational History Group website. The
harmonized data sets can be further veri ed, investigated, and enriched by others, and
they could be integrated into Linked Open Data and other popular formats in
order to utilize the vast pool of existing software tools. Combining such
largescale harmonization with existing data management infrastructures could open
up new doors for research on national bibliographies.</p>
      </sec>
      <sec id="sec-2-3">
        <title>5 https://www.helsinki. /en/researchgroups/computational-history</title>
      </sec>
      <sec id="sec-2-4">
        <title>6 http://data.nationallibrary. /</title>
      </sec>
      <sec id="sec-2-5">
        <title>7 https://github.com/COMHIS/bibliographica</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Emerging applications</title>
      <p>
        Our recent analyses based on four large bibliographies provides examples of the
research potential of this approach [
        <xref ref-type="bibr" rid="ref16 ref9">9, 16</xref>
        ]. One example is the analysis of the
long-term development of book formats and languages across Europe, which
reects changes in public communication. We have reported that on a general
European level the rise of the octavo format is particularly strong during the
eighteenth century, and supported by the data in all four catalogues, where
octavo holds the largest share of the print area by the end of the eighteenth
century. All four metadata collections that we have analysed show a steadily
declining, parallel trend in the share of publications in Latin in the period
15001800. Such observations can highlight material aspects of vernacularization in
the early modern period and re ect European-wide transformations that took
place predominantly during the hand-press era. Moreover, the analyses can also
highlight local variations in the publication pro les of individual European cities
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Therefore joint analysis and comparison of multiple catalogues can be
useful also in terms of assessing the historical representativity of the data, thus
demonstrating the value of bibliographic data science and paving the way for
new research and guidelines for future data integration in this eld.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>We have conceptualized a new approach, bibliographic data science, to expand
the research potential of bibliographic records. This derives from the already
established eld of data science and associates this general paradigm speci cally
with quantitative analysis of bibliographic metadata and related information
sources. While having a speci c scope, BDS is opening up pragmatically oriented
and substantial new research opportunities in the digital humanities.</p>
      <p>Drawing valid conclusions critically depends on data quality, representativity
and completeness. Automation and quality control are essential when the data
collections may contain information on millions of documents, and the
overall data science ecosystem integrates a number of distinct work ows that are
dedicated to harmonizing speci c subsets of the data. We have indicated how
speci cally tailored open data analytical ecosystems can help to address this
challenge. Our approach has potential for wider implementation in related
studies, and provides guidelines for more extensive integration of national collections.</p>
      <p>
        Our future work envisions continued harmonization and data integration for
the HPBD as well as further, related data resources such as the Universal Short
Title Catalogue.8, in order to expand the study to cover public communication
more broadly. We have incorporated best practices and tools from data science,
such as unit tests, tidy data [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and reproducible work ows [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Future
developments could take increasing advantage of machine learning in order to reduce the
need for human input, thus improving the overall scalability of data
harmonization. When combined with a proper quality control, such approaches can have
      </p>
      <sec id="sec-4-1">
        <title>8 https://ustc.ac.uk/</title>
        <p>
          potential for wider implementation in related studies in the digital humanities.
As we have extracted and harmonized publisher information from imprints from
ESTC and FNB [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], it is possible to connect that data to full-text collections
such as the ECCO, and to study how the materiality of printing is related to
developments in newspapers [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Modern statistical techniques are essential not
only the harmonization and quality control of the data but also in investigating
and characterizing the overall spatio-temporal trends, networks, and dynamics in
knowledge production. Taking full advantage of the developments in open sharing
of research data and analysis methods can support collaborative and cumulative
research e orts. Automated harmonization can enhance the overall reliability
and commensurability between independently maintained metadata collections,
thus complementing linked open data and other technologies that primarily
focus on data management and distribution. Hence, bibliographic data science can
help to ll an important gap in the eld by aiming to signi cantly improve the
quality and reliability of the currently available bibliographic records.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baten</surname>
            , J., van Zanden,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Book production and the onset of modern economic growth</article-title>
          .
          <source>Journal of Economic Growth</source>
          <volume>13</volume>
          (
          <issue>3</issue>
          ),
          <volume>217</volume>
          {
          <fpage>235</fpage>
          (
          <year>2008</year>
          ). https://doi.org/10.1007/s10887-008-9031-9
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barnard</surname>
          </string-name>
          , J.:
          <source>Provisional Count of STC Titles, 14751640. Publishing History</source>
          <volume>31</volume>
          (
          <issue>1</issue>
          ),
          <volume>4764</volume>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Borgman</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          :
          <article-title>Big data, little data, no data : scholarship in the networked world</article-title>
          . The MIT Press, Cambridge, Massachusetts; London, England (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bozzolo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ornato</surname>
          </string-name>
          , E.:
          <article-title>Pour une histoire du livre manuscrit au Moyen Age : trois essais de codicologie quantitative</article-title>
          .
          <article-title>Equipe de recherche sur l'humanisme francais des XIVe et XVe siecles, Editions du Centre national de la recherche scienti que</article-title>
          , Paris (
          <year>1980</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Buringh</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanden</surname>
            ,
            <given-names>J.L.V.</given-names>
          </string-name>
          :
          <article-title>Charting the \Rise of the West": Manuscripts and Printed Books in Europe A Long-Term Perspective from the Sixth through Eighteenth Centuries</article-title>
          .
          <source>The Journal of Economic History</source>
          <volume>69</volume>
          (
          <issue>02</issue>
          ),
          <volume>409</volume>
          (
          <year>2009</year>
          ). https://doi.org/10.1017/s0022050709000837
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Giesecke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Der Buchdruck in der frhen Neuzeit : eine historische Fallstudie uber die Durchsetzung neuer Informations-</article-title>
          und
          <string-name>
            <surname>Kommunikationstechnologien</surname>
          </string-name>
          . Suhrkamp,
          <source>Frankfurt am Main</source>
          (
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Horstb ll, H.:
          <article-title>Menigmands medie: det folkelige bogtryk i Danmark 1500-1840: en kulturhistorisk unders gelse. Danish humanist texts and studies</article-title>
          , volume
          <volume>19</volume>
          ,
          <string-name>
            <surname>Det</surname>
            <given-names>Kongelige Bibliotek</given-names>
          </string-name>
          &amp; Museum
          <string-name>
            <surname>Tusculanum</surname>
          </string-name>
          , Copenhagen (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lahti</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Open data science</article-title>
          .
          <source>In: Advances in Intelligent Data Analysis XVII. Lecture Notes in Computer Science 11191</source>
          . vol.
          <volume>11191</volume>
          . Springer, India (
          <year>October 2018</year>
          ),
          <source>conference proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lahti</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marjanen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roivainen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tolonen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Bibliographic data science and the history of the book (c</article-title>
          . 15001800). Cataloging &amp; Classi cation Quarterly pp.
          <volume>1</volume>
          {
          <issue>19</issue>
          (
          <year>January 2019</year>
          ). https://doi.org/10.1080/01639374.
          <year>2018</year>
          .
          <volume>1543747</volume>
          , special issue.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Marjanen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vaara</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roivainen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Makela, E.,
          <string-name>
            <surname>Lahti</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tolonen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Analysing the language, location and form of newspapers in nland</article-title>
          ,
          <fpage>1771</fpage>
          -
          <lpage>1910</lpage>
          . Tech. rep.,
          <source>Digital Humanities in the Nordics, Gothenburg</source>
          (
          <year>2017</year>
          ), conference abstract.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Neddermeyer</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Von der Handschrift zum gedruckten Buch : Schriftlichkeit und Leseinteresse im Mittelalter und in der fruhen Neuzeit : quantitative und qualitative Aspekte. Buchwissenschaftliche Beitrge aus dem Deutschen Bucharchiv Mnchen</article-title>
          , Harrassowitz, Wiesbaden (
          <year>1998</year>
          ), 1: Text ; 2: Anlagen.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Suarez</surname>
            ,
            <given-names>M.F.</given-names>
          </string-name>
          :
          <article-title>Towards a bibliometric analysis of the surviving record 1701{1800</article-title>
          . In: Suarez,
          <string-name>
            <given-names>M.F.</given-names>
            ,
            <surname>Turner</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.L</surname>
          </string-name>
          . (eds.) The Cambridge History of the Book in Britain, pp.
          <volume>37</volume>
          {
          <fpage>65</fpage>
          . Cambridge University Press (
          <year>2009</year>
          ). https://doi.org/10.1017/chol9780521810173.003
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Suarez</surname>
            <given-names>SJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M.F.</surname>
          </string-name>
          :
          <article-title>Book history from descriptive bibliographies</article-title>
          . In: Howsam,
          <string-name>
            <surname>L</surname>
          </string-name>
          . (ed.) The Cambridge Companion to the
          <source>History of the Book</source>
          , pp.
          <volume>199</volume>
          {
          <fpage>218</fpage>
          . Cambridge University Press (
          <year>2014</year>
          ). https://doi.org/10.1017/cco9781139152242.015
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tanselle</surname>
          </string-name>
          , G.T.:
          <article-title>Bibliography and science</article-title>
          .
          <source>Studies in Bibliography</source>
          <volume>27</volume>
          ,
          <issue>55</issue>
          {
          <fpage>90</fpage>
          (
          <year>1974</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Tolonen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lahti</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roivainen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ilomki</surname>
          </string-name>
          , N.:
          <article-title>Printing in a Periphery: a Quantitative Study of Finnish Knowledge Production</article-title>
          ,
          <fpage>1640</fpage>
          -
          <lpage>1828</lpage>
          . In: Digital Humanities 2016: Conference Abstracts. pp.
          <volume>383</volume>
          {
          <fpage>385</fpage>
          . Jagiellonian University &amp; Pedagogical University, Krakw (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Tolonen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lahti</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roivainen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marjanen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A quantitative approach to book-printing in sweden and nland, 16401828</article-title>
          . Historical Methods:
          <source>A Journal of Quantitative and Interdisciplinary</source>
          History pp.
          <volume>1</volume>
          {
          <issue>22</issue>
          (
          <year>2018</year>
          ). https://doi.org/10.1080/01615440.
          <year>2018</year>
          .1526657
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Weedon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The Uses of Quanti cation</article-title>
          . In: Eliot,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Rose</surname>
          </string-name>
          ,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (eds.)
          <article-title>A Companion to the History of the Book</article-title>
          , pp.
          <volume>33</volume>
          {
          <fpage>49</fpage>
          . Blackwell Publishing Ltd, London (
          <year>2008</year>
          ). https://doi.org/10.1002/9780470690949.ch3
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Wickham</surname>
          </string-name>
          , H.:
          <article-title>Tidy data</article-title>
          .
          <source>Journal of Statistical Software</source>
          <volume>59</volume>
          (
          <issue>10</issue>
          ),
          <volume>1</volume>
          {
          <fpage>23</fpage>
          (
          <year>2014</year>
          ). https://doi.org/10.18637/jss.v059.i10
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Wilson,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Bryan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Cranston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Kitzes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Nederbragt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Teal</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.K.</surname>
          </string-name>
          :
          <article-title>Good enough practices in scienti c computing</article-title>
          .
          <source>PLoS Computational Biology</source>
          <volume>13</volume>
          (
          <issue>6</issue>
          ),
          <year>e1005510</year>
          (
          <year>2017</year>
          ). https://doi.org/10.1371/journal.pcbi.1005510
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>