<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>What Shall We Do With the Unseen Sailor? Estimating the Size of the Dutch East India Company Using an Unseen Species Model</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Melvin Wevers</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Folgert Karsdorp</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jelle van Lottum</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>KNAW Huygens Institute</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>KNAW Meertens Institute</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Amsterdam</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <fpage>189</fpage>
      <lpage>197</lpage>
      <abstract>
        <p>Historians base their inquiries on the sources that are available to them. However, not all sources that are relevant to the historian's inquiry may have survived the test of time. Consequently, the resulting data can be biased in unknown ways, possibly skewing analyses. This paper deals with the Dutch East India Company its digitized ledgers of contracts. We apply an unseen species model, a method from ecology, to estimate theactual number of unique seafarers contracted. We 昀椀nd that the lower bound of actual seafarers is much higher than what the remaining contracts indicate: at least, thirty-six percent of the seafarers is unknown. Moreover, we 昀椀nd that even in periods when few records survived, we can still credibly estimate a lower bound on the unique number of seafarers.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Computational History</kwd>
        <kwd>Survivor Bias</kwd>
        <kwd>Unseen Species Model</kwd>
        <kwd>Sampling Without Replacement</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction: Historical Records and Survivor Bias</title>
      <p>
        Historians can only rely on the archival records that have survived the test of time. That a
substantial share of historical records has not survived may be due to natural causes, such
as 昀椀res, decisions on the level of archival policy making, but also content production
biases [17]. For instance, whether or not particular sources were retained can depend on
socioeconomical factors [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], as data representative of lower classes were long deemed less relevant
by archivists [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        As historians are working with data that is hampered by many possible types of bias, they
need to critically evaluate to what extent the remaining data is representative of the collection
or historical period from which its stems1[4]. Put di昀erently, historians need to re昀氀ect on how
transmitted archival records relate to the actual historical past. Especially now when historical
records are rapidly and continuously being digitized, we have to be even more cautious of
potential biases in archives. The speed with which we can analyze data combined with the
increased distance between the researcher and the source material makes it easier to overlook
how bias impacts the historical inferences we make [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Evidently, we can only digitize the
records that have survived. Even if all surviving records would be digitized, biases will remain
to exist.
      </p>
      <p>
        Yet, at the same time, the fact that data is digitized also facilitates the use of statistical and
computational methods which help chart and possibly deal with the blind spots of the data.
Studies have already applied statistical methods to expose how bias can lead to overestimating
the impact of historical events [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], misrepresentations of the standard of living 2[0], or the
underestimation of wartime causalities 1[3].
      </p>
      <p>
        In this paper, we focus on estimating the number of unique seafarers employed by the Dutch
East India Company (Verenigde Oost-Indisch Compagnie, VOC). The company kept detailed
personnel administration records (pay ledgers), which have been digitized in the VOCOP datase1t.
While the dataset, which has been used extensively by historians and social scientists in the
context of 昀椀nancial [
        <xref ref-type="bibr" rid="ref23">22</xref>
        ] and maritime history [
        <xref ref-type="bibr" rid="ref24">23</xref>
        ], contains a sizeable amount of data, we also
know that records have been lost and thus have not been digitized. This can impact the
qualitative and quantitative historical study of the VOC. More speci昀椀cally, we do not know how the
loss of records is distributed temporally, and whether the surviving records can give a credible
estimate of the number of unique individuals hired by the VOC. Having better information on
the representativeness of this data allows us to better study aspects such as career mobility and
the 昀椀nancial position of VOC employees.
      </p>
      <p>
        For the estimation, we draw on unseen species models 7[, 6], which aim to estimate the
number of unique species living in a given environment. Beyond ecology, these models have
been successfully applied to a wide array of cases, ranging from estimating the number of
classes of stone tools in archaeology 1[2], the number of bugs in so昀琀ware code [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the number
of stars in the Pleiades [3], the size of an author’s vocabulary [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and, more recently, to
estimate the number of lost medieval literary works [16]. Here, we apply a modi昀椀cation of the
model for samples without replacement [9], which has not yet been applied in the context of
humanities research.2
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Data and preprocessing</title>
      <p>This paper uses two di昀erent data sources: VOC: CAREERS (VOCCAR) and Dutch Asiatic
Shipping in the 17th and 18th centuries (DAS).</p>
      <p>
        VOCCAR is an enriched version of the VOCOP dataset, which contains digitized pay ledgers
of the VOC. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] The dataset contains 774,200 contracts between 1633 and 1795, with the
majority of records stemming from the 1 8th century. The contracts specify, among other
1These records have been digitized by volunteers working for the National Archives of the Netherlands, and can
be accessed here: https://www.nationaalarchief.nl/onderzoeken/index/nt00444?activeTab=nt
2The data and code used in this paper have been registered under:https://doi.org/10.5281/zenodo.7268250
things, the name, rank, place of birth of the contractee, the date of sailing, and the ship
on which they sailed. The original ledgers from which these records have been digitized
could contain multiple contracts belonging to the same person. In VOCCAR, the records
have been clustered around unique individuals, which allows us to count how o昀琀en they
appeared in the records.
      </p>
      <p>We only focus on the records that have been clustered, resulting in 546,973 records ( ), of
which 460,274 are unique seafarers ( ). Figure 1 provides an overview of the clustered
records in the VOCCAR dataset. We see a sharp increase in records during the 1t7h
century as well as noticeable gaps in the data for the 18th century.</p>
      <p>DAS provides an overview of the number of ships that sailed out from the Dutch republi3c.</p>
      <p>This data is almost complete, with only a few voyages missing from the data4.
We learn from DAS that the VOC sailed out 4,352 times between 1633 and 1795. For
about 91 percent of these voyages, DAS provides information on the number of people
that boarded the vessel. However, because this data contains some noticeable outliers,
we decided to calculate the mean voyagers per ship for periods of twenty-昀椀ve years
rather than imputing merely the missing values. Next, for each period, we multiplied this
mean by the number of voyages, including those that lack information on the number of
voyagers. This provides us with an estimation of the total number of records: 952,1457.
3http://resources.huygens.knaw.nl/das/EnglishIntro
4The data quality is discussed here:http://resources.huygens.knaw.nl/retroboeken/das/
5We use DAS to calculate the total number of records rather than VOCCAR because DAS is much more complete.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Method: Estimating the Number of Unique Individuals under</title>
    </sec>
    <sec id="sec-4">
      <title>Sampling without Replacement</title>
      <p>To estimate the number of unique seafarers of the VOC, we employ an unseen species model.
The model was originally developed in ecology, where researchers are o昀琀en confronted with
incomplete data as a result of undersampling. Due to such data incompleteness, it is possible
that important statistics such as biodiversity are estimated to be much lower than they actually
are. To combat such estimation biases, it is an important research question in ecology how the
resulting di昀erence between the number of observed and the true number of unique species
can be reliably estimated. A potential solution is given by the Chao1 estimator, developed by
Anne Chao [7, 6].</p>
      <p>The Chao1 estimator is a non-parametric unseen species model that estimates a universally
valid lower bound on the number of unseen entities (e.g., seafarers; call that 0), based on
entities that have been observed once or twice (call those 1 and 2). Theoretically, we can
calculate the number of unseen entities by taking the product of the average relative frequency
of unseen entities ( 0) and the number of unseen entities ( 0), divided by 0. However, 0 cannot
be calculated directly. What we do know is that the average relative frequency of unseen
entities ( 0) is probably lower than that of entities occurring once (1), i.e., 0 ≤ 1. It then
follows that 000 must be at least equal to or greater than 010 (hence, 0 is a lower bound). That
latter expression is computable and can be rewritten into the Chao1 estimator [cf8.]:
̂ =
where refers to the observed number of unique entities, to the sum of their occurrences, and
̂ to the bias-corrected lower bound. It is important to note that when 0 ≈ 1, that is, when
unseen entities have approximately the same average relative frequency as entities occurring
once, Chao1 becomes an unbiased point estimator 8[].</p>
      <p>
        The Chao1 estimator was developed assuming that samples are formed with replacement.
This means that during each sampling moment, the same individuals can be observed multiple
times. It also means that observations are independent of each other, and that the observation
of one individual does not a昀ect the observation of the next. In other words, the covariation
between successive observations is zero. Thus, sampling with replacement essentially assumes
an in昀椀nite population. For example, a snippet of text can be seen as a sample of an author’s
in昀椀nite stream of words. And if we apply the Chao1 estimator to this snippet, a lower bound
on the vocabulary of the author is also exactly what is estimated 1[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>We can also think of the snippet as a sample of the 昀椀nite space of the snippet’s
encompassing book. Treating the snippet as such would imply that the sample was createdwithout
replacement. In such samples, observations are not independent, nor is the covariation between
successive observations zero. Crucially, however, because of its assumption that samples are
formed with replacement and are thus drawn from an in昀椀nite population, the Chao1 estimator
does not estimate the number of unique words in the book encompassing the snippet. Thus,
even though we know a given sample to come from a 昀椀nite population, Chao1 always treats it
as coming from an in昀椀nite one.</p>
      <p>The VOC records of this study should be conceptualized as samples created without
replacement. There has been a 昀椀nite population of seafarers with the VOC of which the records show
a sample without replacement.6 The problem, however, is that when we apply the Chao1
estimator to this sample, we do not obtain an estimate of the number of unique individuals in
the total, 昀椀nite population, but rather that of a potential population of seafarers, which is not
what we are a昀琀er. To estimate the number of unique individuals in the 昀椀nite population of
employees of the VOC, we employ a modi昀椀ed Chao1 estimator developed by Chao and Lin for
samples without replacement [9]. This modi昀椀ed estimator assumes we know the size of the
total population, and thus know the ratio of the observed sample size to the total population:
̂wor =
+
,
(2)
(3)
where = /( − 1)and = /(1 − .)Note that when approaches zero, Eq. 2 reduces to the
standard Chao1 estimator in Eq.1. We refer to the modi昀椀ed estimator as Chao1 wor. Con昀椀dence
intervals for Chao1wor can be computed based on the variance estimator 9[]:
var( ̂wor) = 0̂ +</p>
      <p>Based on the total number of records we derived from DAS (̂ = 952, 147 , see above), we
calculate the sample fraction by dividing by ̂ . For the complete dataset ≈ 0.57. Note
that ̂ ≠ ̂ , since individuals may have been shipped out multiple times.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Results: There are many more unique seafarers than the records show</title>
      <p>At least thirty-six percent of the seafarers is unknown Based on the observed
abundances, i.e. how many times each unique individual was s“ighted” in the data, and the sample
fraction estimated from ̂ , we calculate with Chao1wor the lower bound on the number of
actual unique individuals in the VOC population (̂ ) to be 716,818 (95%CI: 715,439 to 718,203).
This number suggests that we should account for a survival rate of / ̂ ≈ 64%, or conversely,
that of the original VOC population, at least 36% of the individuals is unknown.
The loss of records impacted the number of unique individuals in the records To get
a better understanding of the coverage of the data across time, we applied the same approach
to successive periods of twenty-昀椀ve years. For each period, we calculated the mean number
of voyagers on a journey and multiplied this with the total number of journeys in that period,
thus estimating the actual number of seafarers (see Table 1). Figure 2 displays the observed
number of unique seafarers against the estimated number ̂ over time. The gray overlay
6This 昀椀nite population can be constrained by many di昀erent things, about which we can now only speculate: the
total number of ships, skills required to be enlisted, etc.
Chao1wor for time spans of 25 years.
represents the 95% con昀椀dence intervals of the estimates. The plot shows that, especially in
the 17th century, the lack of data has led to a severe underestimation of the number of unique
individuals (ranging from 45% in 1683–1708 to 96% in 1633–1658). With an average of≈ 23%,
the gap between the observed and the actual number of unique individuals is smaller in the
18th century but still considerable.
Overview of data for twenty-five-year periods. The last period only spans twelve years.</p>
      <p>period
just below 1 million. The same is true for the shorter periods of 25 years. Here too, Chao1
systematically produces impossible estimates. As we explained above, the Chao1 estimate might
be considered thepotential rather than the actual number of seafarers that could have worked
for the VOC. By contrast, the estimate of the Chao1wor estimator is compatible with the upper
limit of ̂ and thus supports our approach of conceptualizing these sightings as samples
without replacement. More generally, these results emphasize the need to understand the sampling
process underlying the data, and to exercise caution when applying the estimators. When data
are sampled without replacement, but the estimator assumes otherwise, Chao1 is not
guaranteed to produce a lower bound, which puts any reliable interpretation of the results into
question.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusion</title>
      <p>This short paper is the 昀椀rst to quantify the scale and extent of the assumed data loss and lack
of representativeness of the archives of the the Dutch East India Company (the VOC). We
applied the Chao1wor estimator to a database of employees of the VOC, and found that we can
make credible predictions on the lower bound of the number of unique seafarers that have
been employed by the company. Moreover, even when relatively small fractions of the records
have survived, the estimates appear to be robust. For the entire archival period, we estimate
that at least forty percent of unique seafarers are not recorded in the archives. Put di昀erently,
the actual number of unique seafarers was much higher than the surviving records indicate.
Moreover, the estimated increase in the number of unique seafarers in the 17th century is not
as steep in actuality as the empirical, observed records suggest. Finally from the 1t8h century
onward, the di昀erence between the observed and the actual number of unique seafarers is
smaller but still considerable. More generally, our results show how unseen species models
from ecology can be used to obtain a clearer perspective on the parts of historical archives that
are lost.</p>
      <p>
        This paper adds to a series of recent studies exploring the applicability of unseen species
models to cultural data [
        <xref ref-type="bibr" rid="ref10">12, 15, 10, 16</xref>
        ]. While these prior studies primarily investigate
samples from in昀椀nite populations, the present paper explored the applicability of Chao1 without
replacement [9] in the context of cultural data sampled from 昀椀nite populations. The case study
of the VOC underscored the importance of a proper conceptualization and understanding of
the sampling process underlying the data. Without such understanding, or when the
assumptions about the sampling process of the model do not correspond to the actual sampling process
underlying the observed data, the estimates may no longer be reliable – or, more precisely –
they do not match what we hope to estimate. For example, when data are sampled without
replacement, Chao1 is no longer guaranteed to estimate a lower bound. The records of the VOC
should be conceptualized as a sample without replacement, for which the modi昀椀ed Chao1wor
estimator can, by contrast, adequately estimate a credible lower bound. An important remaining
issue with the application of unseen species models to cultural data (whether they are sampled
with or without replacement) is that the data are assumed to be homogeneous and thus that all
entities (e.g., seafarers) are equally likely to be observed. The consequence of this simplifying
assumption is that the unseen species estimators reduce from a point-estimate to a lower bound
of the actual population size. In a series of studies, Böhning and colleagues present generalized
unseen species models that show how adding information about the origins of heterogeneity
of the data can reduce some of the bias of the estimates 4[, 2]. In future work, we aim to re昀椀ne
our estimates by incorporating such covariate information in these generalized unseen species
models. For example, the current analysis o昀ers no information on whether factors such as
rank or origin impacted the loss of certain records. It is quite conceivable, however, that the
scrupulousness of the log 昀椀les may vary between records of high-ranking o昀케cials from the
Dutch republic and those of seafarers from further away. One may also wonder whether the
e昀ect of rank or origin 昀氀uctuates over time, possibly relating to periods of social unrest.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Benjamin</surname>
          </string-name>
          . Race A昀琀er Technology:
          <article-title>Abolitionist Tools for the New Jim Code . 1st edition</article-title>
          . Medford, MA: Polity, June 17,
          <year>2019</year>
          . 172 pp.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Bohning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vidal-Diez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lerdsuwansri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Viwatwongkasem</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Arnold</surname>
          </string-name>
          .
          <article-title>“A Generalization of Chao's Estimator for Covariate Information”</article-title>
          .
          <source>InB: iometrics 69</source>
          (
          <year>2013</year>
          ), pp.
          <fpage>1033</fpage>
          -
          <lpage>1042</lpage>
          . doi:
          <volume>10</volume>
          .1111/biom.12082.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Böhning</surname>
          </string-name>
          . “
          <article-title>Chao's Lower Bound Estimator and the Size of the Pleiades”</article-title>
          .
          <source>InE:nvironmental and Ecological Statistics 27.1</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>171</fpage>
          -
          <lpage>173</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10651-020-00440- w.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Böhning</surname>
          </string-name>
          and
          <string-name>
            <surname>P. G. M. van der Heijden.</surname>
          </string-name>
          “
          <article-title>A Covariate Adjustment for Zero-Truncated Approaches to Estimating the Size of Hidden and Elusive Populations”</article-title>
          .
          <source>InT:he Annals of Applied Statistics 3.2</source>
          (
          <year>2009</year>
          ). doi:
          <volume>10</volume>
          .1214/
          <fpage>08</fpage>
          -
          <lpage>aoas214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Briand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>El Emam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Freimut</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Laitenberger</surname>
          </string-name>
          .
          <article-title>“A Comprehensive Evaluation of Capture-Recapture Models for Estimating So昀琀ware Defect Content”</article-title>
          .
          <source>In: IEEE Transactions on So昀琀ware Engineering 26.6</source>
          (
          <issue>2000</issue>
          ), pp.
          <fpage>518</fpage>
          -
          <lpage>540</lpage>
          . doi:
          <volume>10</volume>
          .1109/32.852741.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          . “
          <article-title>Estimating Population Size for Sparse Data in Capture-Recapture Experiments”</article-title>
          .
          <source>In: Biometrics 45.2</source>
          (
          <issue>1989</issue>
          ), pp.
          <fpage>427</fpage>
          -
          <lpage>438</lpage>
          . doi:
          <volume>10</volume>
          .2307/2531487.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          . “
          <article-title>Nonparametric Estimation of the Number of Classes in a Population”</article-title>
          .
          <source>InS:candinavian Journal of Statistics 11.4</source>
          (
          <issue>1984</issue>
          ), pp.
          <fpage>265</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>“Deciphering the Enigma of Undetected Species, Phylogenetic, and Functional Diversity Based on Good-Turing Theory”</article-title>
          .
          <source>In:Ecology</source>
          <volume>98</volume>
          .11 (
          <year>2017</year>
          ), pp.
          <fpage>2914</fpage>
          -
          <lpage>2929</lpage>
          . doi:
          <volume>10</volume>
          .1002/ec y.
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.-W.</given-names>
            <surname>Lin</surname>
          </string-name>
          . “
          <article-title>Nonparametric Lower Bounds for Species Richness and Shared Species Richness under Sampling without Replacement”</article-title>
          .
          <source>In: Biometrics 68.3</source>
          (
          <issue>2012</issue>
          ), pp.
          <fpage>912</fpage>
          -
          <lpage>921</lpage>
          . doi:
          <volume>10</volume>
          .1111/j.1541-
          <fpage>0420</fpage>
          .
          <year>2011</year>
          .
          <volume>01739</volume>
          .x.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Colwell</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          . “
          <article-title>Measuring and comparing class diversity in archaeological assemblages: A brief guide to the history and state-of-the-art in diversity statistics</article-title>
          .” In: De昀椀ning and Measuring Diversity in Archaeology.
          <source>Another Step Toward an Evolutionary Synthesis of Culture</source>
          . Ed. by
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Eren</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Buchanan</surname>
          </string-name>
          . New York - Oxford: Berghahn,
          <year>2020</year>
          , pp.
          <fpage>263</fpage>
          -
          <lpage>294</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Efron</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Thisted</surname>
          </string-name>
          . “
          <article-title>Estimating the Number of Unseen Species: How Many Words Did Shakespeare Know?”</article-title>
          <source>In: Biometrika 63.3</source>
          (
          <issue>1976</issue>
          ), p.
          <fpage>435</fpage>
          . doi:
          <volume>10</volume>
          .2307/2335721.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Eren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-H.</given-names>
            <surname>Hwang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Colwell</surname>
          </string-name>
          . “
          <article-title>Estimating the Richness of a Population When the Maximum Number of Classes Is Fixed: A Nonparametric Solution to an Archaeological Problem”</article-title>
          .
          <source>In:PLoS ONE 7.5</source>
          (
          <year>2012</year>
          ). Ed. by
          <string-name>
            <given-names>A.</given-names>
            <surname>Mesoudi</surname>
          </string-name>
          ,
          <year>e34179</year>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0034179</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Gillespie</surname>
          </string-name>
          . “
          <article-title>Estimating the Number of Casualties in the American Indian War: A Bayesian Analysis Using the Power Law Distribution”</article-title>
          .
          <source>In:The Annals of Applied Statistics 11.4</source>
          (
          <issue>2017</issue>
          ), pp.
          <fpage>2357</fpage>
          -
          <lpage>2374</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>K.</given-names>
            <surname>Inwood</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Maxwell-Stewart</surname>
          </string-name>
          .
          <article-title>“Selection Bias and Social Science History”</article-title>
          .
          <source>InS:ocial Science History 44.3</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>411</fpage>
          -
          <lpage>416</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Karsdorp</surname>
          </string-name>
          . “
          <article-title>Estimating the Loss of Medieval Literature with an Unseen Species Model from Ecodiversity”</article-title>
          .
          <source>In:Workshop on Computational Humanities Research</source>
          . Amsterdam: Ceur-ws,
          <year>2020</year>
          , pp.
          <fpage>44</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Sawyer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Sleiderink</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          . “Forgotten Books:
          <article-title>The Application of Unseen Species Models to the Survival of Culture”</article-title>
          .
          <source>In:Science 375.6582</source>
          (
          <year>2022</year>
          ), pp.
          <fpage>765</fpage>
          -
          <lpage>769</lpage>
          . doi:
          <volume>10</volume>
          .1126 /science.abl7655.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Lee</surname>
          </string-name>
          . “
          <article-title>The Library of Babel: How (and How Not) to Use Archival Sources in Political Science”</article-title>
          .
          <source>In: Journal of Historical Political Economy 2.3</source>
          (
          <issue>2022</issue>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>L.</given-names>
            <surname>Mordechai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eisenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. P.</given-names>
            <surname>New昀椀eld</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Izdebski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Kay</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Poinar</surname>
          </string-name>
          . “
          <source>The Justinianic Plague: An Inconsequential Pandemic?” In:Proceedings of the National Academy of Sciences 116.51</source>
          (
          <year>2019</year>
          ), pp.
          <fpage>25546</fpage>
          -
          <lpage>25554</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>L.</given-names>
            <surname>Petram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koolen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wevers</surname>
          </string-name>
          , R. van Koert,
          <article-title>and</article-title>
          <string-name>
            <surname>J. van Lottum. “</surname>
          </string-name>
          <article-title>Data on the Maritime Workforce of the Dutch East India Company in the 18th Century”</article-title>
          . In: (Forthcoming).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>B.</given-names>
            <surname>Quanjer</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Kok</surname>
          </string-name>
          .
          <article-title>“Dra昀琀ing the Dutch: Selection Biases in Dutch Conscript Records in the Second Half of the Nineteenth Century”</article-title>
          .
          <source>In: Social Science History 44.3</source>
          (
          <issue>2020</issue>
          ), pp.
          <fpage>501</fpage>
          -
          <lpage>524</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>M.-R. Trouillot</surname>
            and
            <given-names>H. V.</given-names>
          </string-name>
          <string-name>
            <surname>Carby</surname>
          </string-name>
          .
          <article-title>Silencing the Past: Power and the Production of History</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Boston</surname>
          </string-name>
          , Massachusetts: Beacon Press,
          <year>2015</year>
          . 190 pp.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [22]
          <string-name>
            <surname>C. Van Bochove</surname>
            and
            <given-names>T. Van Velzen.</given-names>
          </string-name>
          “Loans to Salaried
          <source>Employees: The Case of the Dutch East India Company</source>
          ,
          <fpage>1602</fpage>
          -
          <lpage>1794</lpage>
          ”.
          <source>In: European Review of Economic History 18.1</source>
          (
          <issue>2014</issue>
          ), pp.
          <fpage>19</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23]
          <string-name>
            <surname>J. van Lottum</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Petram</surname>
          </string-name>
          . “In Search of Strayed Englishmen.
          <article-title>English Seamen Employed in the Dutch East India Company in the Late Seventeenth and Eighteenth Centuries”</article-title>
          . In:
          <article-title>Anglo-Dutch Connections in the Early Modern World</article-title>
          . Ed. by E. van Raamsdonk,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levelt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Rose</surname>
          </string-name>
          . London: Taylor &amp; Francis, Forthcoming. Forthcoming.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>