<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Antwerp, Belgium
∗Corresponding author.
£ folgert@karsdorp.i(oF. Karsdorp);e.m.a.manjavacas.arevalo@hum.leidenuniv.n(lE. Manjavacas);
l.fonteyn@hum.leidenuniv.n(lL. Fonteyn)
ç https://www.karsdorp.io/(F. Karsdorp);
https://www.universiteitleiden.nl/medewerkers/enrique-manjavacas-areva(lEo./Manjavacas);
https://www.universiteitleiden.nl/en/staffmembers/lauren-fonte(yLn. Fonteyn)
ȉ</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Introducing Functional Diversity: A Novel Approach to Lexical Diversity in (Historical) Corpora</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>FolgertKarsdorp</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>EnriqueManjavacas</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>LaurenFonteyn</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>KNAW Meertens Institute</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Leiden University</institution>
          ,
          <addr-line>Leiden</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The question how we can reliably estimate the lexical diversity of a particular text (collection) has o昀琀en been asked by linguists and literary scholars alike. This short paper introduces a way of operationalizing functional diversity measurements by means of token-based embeddings, and argues that functional diversity is not only a practically advantageous, but also a theoretically relevant addition to the Computational Humanities Research toolkit. By means of an experiment on the historical ARCHER corpus, we show that lexical diversity at the level of functional groups is less sensitive to orthographic variation, and provides insight into an important and o昀琀en disregarded dimension of vocabulary diversity in textual data.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Lexical diversity</kwd>
        <kwd>Functional diversity</kwd>
        <kwd>Historical text</kwd>
        <kwd>Hill numbers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        knows at di昀erent ages [
        <xref ref-type="bibr" rid="ref24 ref4">5, 25</xref>
        ]. In a similar vein, researchers have also attempted to estimate
(and compare) the richness of the active vocabulary of particular authors [1e6. g,.18] or
literary works across time [e.g.2,3], or the ‘productivity’ of linguistic structures (i.e., how many
di昀erent word types are used in a particular linguistic contex2t,[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]) for di昀erent individuals
[e.g., 24, 1] or across time [e.g.,21]. To attain these goals, researchers o昀琀en resort to corpus
research, using text (excerpt) collections of varying sizes with diversity measures that rely on
the number of word tokens, unique word types, and/or hapax legomena (i.e., words that
occur only once), such as (variations on) Mean Word Frequency (MWF) and Type-Token Ratio
(TTR) [for examples, see27], realized/potential/expanding productivity2][, or measures that
originate in Shannon entropy2[
        <xref ref-type="bibr" rid="ref5">6</xref>
        ].
      </p>
      <p>There is, however, a practical problem that arises with any measure of diversity that relies on
hapaxes and/or unique types. In many digitized text corpora, the number of unique character
strings cannot be equated to the number of unique words. A substantial amount of variation
in how word types are represented in a corpus may be due to OCR errors (e.g., in historical
texts, the long S character &lt;ſ&gt; is o昀琀en mistaken for &lt;f&gt; or &lt;l&gt;, which means the word type
strength could be represented in a corpus as at least three di昀erent character strings: &lt;ſ罴rength&gt;,
&lt;frength&gt; and &lt;lrength&gt;). Furthermore, some types of corpora contain texts where authors
do not (consistently) adhere to (present-day) standard spelling conventions, such as corpora of
(informal) language on social media or any historical corpora that pre-date the establishment of
uniform spelling conventions. This introduces a dimension of variation that makes it di昀케cult to
accurately count the number of actual hapax legomena or unique types. Of course, at least some
of this unwanted variation can be tackled in corpus pre-processing through (semi-)automated
spelling normalisation, but this too can prove challenging given that neither OCR errors nor
non-standard spelling variation are entirely or even largely systematic.</p>
      <p>
        In this paper, we argue that there are substantial advantages to relying on functional
diversity measures (rather than, or as a complement to lexical diversity measures) to estimate and
compare the ‘lexical richness’ of (collections of) text. More speci昀椀cally:
• We demonstrate that functional diversity estimates are a昀ected to a much lesser extent
by spelling errors and inconsistencies than lexical diversity estimates. As such, there is
a clearpractical advantage to relying on functional diversity.
• We suggest that, even in corpora that are free from orthographic noise, there
tisheaoretical advantage to examining higher-order diversity at the level of functional groups.
We propose that a theoretically relevant distinction can be made when making claims
about ‘vocabulary richness’ or lexical diversity by taking the semantic similarity between
words into account1. This higher-order, functional-semantic dimension of diversity is
theoretically relevant, as it helps characterize diversity in terms of depth and width, and
o昀ers a perspective on diversity that is not captured by more traditional, exclusively
categorical measures.
1The distinction between lower-order and higher-order diversity proposed here is reminiscent of the distinction
between ‘productivity’ and ‘schematicity’ in2[
        <xref ref-type="bibr" rid="ref1 ref12">1, 13</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Measuring Diversity</title>
      <p>
        Functional Diversity For our measurements of functional diversity in (historical) corpora,
we apply the framework of attribute diversity, which was originally developed in the context
of ecological diversity8,[
        <xref ref-type="bibr" rid="ref5">6</xref>
        ]. In ecology too, it is important to not only account for categorical
diversity (the taxonomic model of species diversity), but also for attribute variation between
and within species. A昀琀er all, certain species (e.g., ducks vs. geese) are more similar to each
other than others (e.g., ducks vs. sheep). This is not captured by taxonomic diversity, which
treats all species as equally distant.
      </p>
      <p>
        In the framework of attribute diversity8,[
        <xref ref-type="bibr" rid="ref5">6</xref>
        ], categorical diversity is considered a special case
of functional diversity, where each type (or species) is considered its own functional group and
all groups are functionally equally di昀erent. In this extreme case, each functional di昀erence
results in the de昀椀nition of a new functional group which is equivalent to a categorical type.
More precisely, the thresholdfor de昀椀ning a new functional group is set to the smallest
pairwise distance between types. The framework allows researchers to specify functional groups
at higher distinctiveness thresholds. then speci昀椀es the distance threshold beyond which
types are considered equally distant and thus belong to di昀erent functional groups. Atsends
to in昀椀nity, types become functionally indistinct and belong to the same functional group.
      </p>
      <p>Each type ÿ contributes to the frequency of a functional group. Leÿt be the frequency of
type ÿ and ÿ the frequency of a functional group, thenÿ( )can be de昀椀ned as the proportional
contribution of typeÿ to a group for a given threshold leve.l Functional diversity, then, is
de昀椀ned as the sum over the proportional contributionsÿ( )of each type ÿ = 1, 2, … , ā:
where ÿ( ) = ÿ/ ÿ. Note that when each type belongs to its own functional group, i.e., when
the de昀椀nition of functional groups and types coincide, ÿ equals unity. In this case, ÿ = ÿ( )
and thus the functional diversity is equal to the number of typāe.s When functional groups
and types do not coincide, certain functional groups consist of more than one type, which in
turn may belong to more than one group. To account for such many-to-many type-function
relations, the abundance ÿ at threshold is computed as the number of tokens of typ eÿ plus a
fraction of the tokens of any other typeĀ that is functionally indistinctive from typÿ:e
ā
ÿ( ) = ÿ + ∑ (1 − ÿĀ( )) Ā</p>
      <p>Ā≠ÿ
here, ÿĀ( )refers to the distance between typeÿ and Ā, which is set to if ÿĀ &gt;
otherwise.</p>
      <p>
        Functional Hill numbers Eq. 1 describes the functional richness of a collection, or the
number of functional groups given a distinctiveness threshol.dRichness is just one of many
diversity measures which treats each functional group as equally important. However, certain
functional groups may be more prominent than others which better is captured by other
diversity measures, like Shannon entropy or the Gini-Simpson index. To account for other aspects
of diversity, Chao and colleagues8,[
        <xref ref-type="bibr" rid="ref5">6</xref>
        ] integrate functional diversity into a mathematically
uni昀椀ed family of diversity indexes called Hill Number1s4[]. Hill numbers are parameterized
only by , which determines the sensitivity to the relative frequencyÿ of variant typeÿ [
        <xref ref-type="bibr" rid="ref13 ref16 ref6 ref8">14, 7,
17, 9</xref>
        ]:
ā
(( 1, … , ā)) =(∑ ÿ )
ÿ=1
The diversity values at certain orderscorrespond to well-known diversity indices. The
number of unique types (also called the ‘richness’ of a sample is equal0 to. With = 0 no weight
is given to the relative frequency of the types, or, conversely, maximum weight is given to rare
types. By setting to 1, the weight of each type is proportional to its relative frequency. Note,
however, that1 is unde昀椀ned. Yet, the limit lim →1 exists, which is equal to the exponent of
Shannon entropy [cf.7, 17, 9]. With &gt; 1, disproportionally more weight is given to more
frequent types. For instance, the Hill number of orde=r 2 is equal to the inverse of the
GiniSimpson index, which expresses the probability that two random tokens are of the same type.
An interesting property of Hill numbers is that all diversity indices are expressed in terms of
the e昀ective number of types: the number of equally frequent types required to obtain a
particular observed diversity value. Because the indices are on the same scale, they can easily be
represented in ‘diversity pro昀椀les’, which chart the diversity at di昀erent order. These pro昀椀les,
then, can be used to characterized the evenness of some collection. Pro昀椀les with steep declines
indicate a large disparity in the frequencies of the types, wheres 昀氀at pro昀椀les indicate a more
even distribution among types.
      </p>
      <p>
        By incorporating functional diversity into the Hill number framework, Chao, Chiu and
colleagues [
        <xref ref-type="bibr" rid="ref5 ref7">8, 6</xref>
        ] show how to estimate the e昀ective number of equally distinct functional groups
at a given distinctiveness threshold and diversity order. The ‘e昀ective number’, sometimes
called ‘true diversity’, represents the number of types in an idealized reference sample that all
have the same frequency and distance between them of at least . Expanding on Eq. 1, the
functional diversity of orderis de昀椀ned as follows:
      </p>
      <p>ā
(Δ( )) =(∑ ÿ( )( ÿ( )) )
ÿ=1
1
1−
(3)
(4)
where refers to the total number of tokens in the collection.</p>
      <p>Example To obtain a better intuition of what functional diversity measures entail, and
specifically how the measure responds to the parameter, we present the following exampl2e.
Consider these four words and their corresponding frequenciaepsr:icot ( 1 = 20),
pineapple ( 2 = 15), digital ( 3 = 10), information ( 4 = 5). For each word, Table1 lists whether
it co-occurs with any of ten context words. Each word can thus be represented as a binary
2Our example is a translation of6][ into a linguistic context.
context vector, which can be used to compute the distance between two words. For example,
computing the pairwise distances between all four words using the Jaccard distance yields the
following distance matriΔx:
apricot 0
pineapple £¤0.4
information ¤ 1
digital ¥ 1</p>
      <p>In Figure1, we calculate functional diversity for di昀erent distinctiveness threshold. sWe
begin in the top row with = min, which is equal to the minimum distance between di昀erent
word types (i.e., intra-type distances are not considered in this example). Amtin, equals
0.4, which means that word types with at least a distance of 0.4 between them are considered
functionally equally distant. This translates by truncating all distances greater th=an0.4 to
0.4 in the distance matrix (cf. the matrix in Figur1e). In this scenario, each word is functionally
equally distant and thus each type has a proportional contributi oÿnof unity to its functional
group, or, in other words, each type makes up for its own functional group. This is illustrated
in Figure1 with the circles whose size is proportional to their frequency. The circles do not
overlap, which illustrates that they each comprise their own functional group. The functional
diversity at = 0 , then, is = 4 , which is simply the sum over the proportional contributions
ÿ( )of each type ÿ to a functional group (cf. Eq1.).</p>
      <p>
        As the threshold value increases, an increasing number of types becomes functionally
indistinguishable. In other words, with higher values o, ffunctional groups consist of more
types. Chiu and Chao [
        <xref ref-type="bibr" rid="ref5 ref7">8, 6</xref>
        ] suggest to use Rao’s quadratic entropy for , which is a
similaritysensitive diversity measure representing the average distance between two randomly selected
instances in a collection2[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. , herea昀琀er denoted as mean, is expressed as:
=
      </p>
      <p>ā ā
mean = ∑ÿ=1 ∑Ā=1 ÿĀ ÿ Ā ,
(5)
where ÿĀ refers to the distance between typesÿ and Ā, and ÿ and Ā to their relative frequencies.</p>
      <p>As shown in Figure 1, setting at mean = 0.54 decreases the functional diversity to
= 3.58 . At the threshold of 0.54,apricot and pineapple become functionally less distinct,
contributing to a shared functional group (illustrated by the overlapping circles). By contrast,
at = mean, digital and information remain functionally equally distant and as such belong to
their own functional group. Note that with = 3.58 , the functional diversity at = mean
is larger than 3. This is because the co-occurrence pro昀椀le oafpricot and pineapple is
considered partially overlapping but not identical. Whenis set to the maximum distance in the
distance matrix ( = max = 1, however, digital and information contribute to a shared
functional group. Note that when there are many di昀erent word types, = max is less informative
than = mean, because functional diversity is then o昀琀en close to unity [cf6.]. Finally, as
tends to in昀椀nity, all words become part of the same functional group, which is expressed by a
functional diversity of = 1 .</p>
    </sec>
    <sec id="sec-3">
      <title>3. Data and pre-processing</title>
      <p>
        Archer Corpus For our experiments, we use ARCHER 3.22[
        <xref ref-type="bibr" rid="ref7">8</xref>
        ], a corpus of historical English
registers (3.3M words). The corpus covers a period of almost 400 years (1600-1999), and
contains texts from 12 di昀erent genres or registers: advertisements, drama, 昀椀ction, sermons,
journals, legal text, medicine, news, early prose, science, letters, and diaries. In terms of spelling,
ARCHER 3.2 contains the original spelling of published editions normalized with VAR4D].2 [
In contrast to many other historical corpora, ARCHER 3.2 is a well-balanced, cleaned (and
relatively small) corpus, and hence it constitutes the ideal starting point for our experiment.
Simulating Errors To mimic di昀erent degrees of text noise, we ‘pollute’ each text in the
clean ARCHER corpus by simulating errors. In this simulation procedure, each token of each
text is modi昀椀ed with a probability . The modi昀椀cation involves replacing each letter by a
random ASCII letter with probabilit y. With = 0.2, a word likediversity is replaced with
diversizy. We experiment with ∈ 0, 0.1, 0.2, 0.35, 0.5, 0.75 , and chart the import of having a
more distorted text on the stability of the diversity measur3esI.n all experiments is set to 0.2.
With this procedure, each text is manipulated 昀椀ve times per value. The reported diversity
measurements are computed by taking the mean diversity over these 昀椀ve di昀erent texts.
Embeddings For the present study, we use token-based embeddings to obtain semantic
similarity estimates between the words in a given text. These embeddings are computed on the
basis of MacBERTh [
        <xref ref-type="bibr" rid="ref18 ref19">20, 19</xref>
        ], a Large Language Model that follows the architecture of BERT-base
uncased [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ], which is pre-trained on historical English (1450-1950) using a custom vocabulary.
Token-based embeddings are expected to be more robust than type-based embeddings in the
presence of noise, since they take the sentential context in which the target word appears into
account. This means that they can associate (even lower frequency) variants of the same word
with each other, where the sentential context is expected to match. Moreover, thanks to the
built-in adaptive tokenization approach, MacBERTh is also able to compute embeddings for
words that were not seen during training, which is an invaluable feature for texts with large
amounts of orthographic variatio4n.
      </p>
      <p>
        In order to compute the type-level distance matrix between all word types in a corpus, we
3Note that texts resulting from = 0.75 are perhaps less realistic than lower values. To illustrate, the OCR error
rate in Eighteenth Century Collections Online (ECCO) has, for instance, been estimated to at approximately 25%
[
        <xref ref-type="bibr" rid="ref14">15</xref>
        ].
4The purpose of this paper is to introduce the attribute diversity framework into lexical diversity research. As a
椀昀rst operationalization, we resorted to token-based embeddings, which is a theoretically sound choice (as these
models are sensitive to the fact that words can have multiple meanings) that comes with certain practical
advantages (with respect to lower-frequency and ‘unseen’ items). We are, however, interested in trying out other ways
of operationalizing the concept of functional groups in future work. One possibility, for instance, would be to
test and compare di昀erent implementations of implement semantic similarity, comparing type and token-based
approaches.
椀昀rst compute the token-embeddings of all words it contain5s. If the same token appears
multiple times in the input corpus, we compute a single embedding by averaging over the
embeddings of all occurrences. Finally, we rely on the cosine distance function in order to obtain a
distance value between 0 and 26.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. Functional diversity is a昀ected less by increased orthographic variation</title>
        <p>Figure2 shows the relative change in the number of functional groups a昀琀er modifying texts
with a text modi昀椀cation probability with respect to their original, unmodi昀椀ed counterparts
( = 0 ). The le昀琀 panel shows the values for = 0 at three di昀erent thresholds of . As expected,
the number of functional entities of functional entities at = min increases more or less
5Note that due to the tokenization approach of MacBERTh, input words are o昀琀en split into smaller units
(subtokens). In order to compute a single embedding in such cases, we average over the embeddings of the di昀erent
sub-tokens.
6More speci昀椀cally, the cosine distance is de昀椀ned as 1 minus the cosine similarity of two given vectors—the latter
being bounded between -1 and 1.
linearly with the probability of modifying words. Indeed, the probability of a modi昀椀cation
yielding a orthographically unique letter combination is high, and each unique combination
is taken to account for a new word type (i.e., a new functional group). The relative change
in the number of functional groups is much less strong for= mean, where the number of
functional groups at extreme values ofis still relatively close to the number of groups=at0 .
Note that the same holds for = max, which also remains stable with larger values of noise.
However, as explained above, with = max, estimates are o昀琀en close to unity, which makes
the stability of = max less surprising. The two right panels present the same results for
higher diversity orders. These plots show that when more weight is given to high frequency
entities, functional diversity is also better able to cope with orthographic variation than lexical
diversity at = min.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Functional diversity is a theoretically relevant complement to lexical diversity</title>
        <p>To get a 昀椀rmer grip on what could be gained from integrating functional diversity estimates
into discussions of lexical richness, we automatically identi昀椀ed text pairs with approximately
the same number of unique word types (min), but a diverging number of functional groups at
mean. In each of these text pairs, one text is functionally less ‘condensed’, using the same
number of unique lexical items to cover a broader functional range. A commonly occurring type
of text pairing, in that respect, is that of a 昀椀ction text with a text containing a collection of
advertisements, where advertisements consistently cover a smaller number of functional groups
despite being as lexically diverse as the paired 昀椀ction text. The relatively strong reduction from
min to mean in advertising, illustrated in Figu3r,eis intuitive, as advertisements o昀琀en present
a list of (functionally closely related) services and/or goods (see Fi4g)u, rreesulting in a more
condensed diversity that suggests depth rather than breadth. For 昀椀ction, by contrast, there is
no reason to expect a similar reduction.</p>
        <p>Pairings of two texts from the same genre also emerged. A telling example is the pairing of
Isabel Clarendon (1886), a 昀椀ction text by naturalist/realist author George Gissing, witChaprice
(1917) by Ronald Firbank (see Figur5e). The excerpts in the corpus from both texts have roughly
the same number of unique word typesC(aprice: 1374 vs. Isabel Clarendon: 1377), but the types
in Caprice – a minimalist novel that, unlike realist work, predominantly consists of dialogue
and contains only limited descriptions of setting and character – cover a considerably smaller
number of functional groups at = mean (213 vs. 368). Interestingly, with 5260 word tokens,
the excerpt ofIsabel Clarendon has a lower TTR than the excerpt ofCaprice, which comprises
3753 tokens. Hence, the TTRs would suggest thatCaprice covers more ground in fewer words.
The functional diversity estimate, however, paints a di昀erent picture, which adds a theoretically
relevant dimension to investigations into the lexical richness of texts.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this short paper, we introduce a way of incorporating the notion of functional diversity into
lexical diversity measurements in (historical) corpora by means of token-based embeddings.</p>
      <p>Our experiment shows that considering lexical diversity at the level of functional groups has
the practical advantage of being less sensitive to orthographic noise in the data, and the
theoretical advantage of adding an important and o昀琀en disregarded dimension (capturing depth vs.
breadth) of vocabulary diversity in textual data. As such, the framework of attribute diversity
commonly used in Ecology should be considered an important addition to the Computational
Humanities research toolkit.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The training of MacBERTh has been made possible by the Platform Digital Infrastructure
(Social Sciences and Humanities) fund (PDI-SSH). We thank Melvin Wevers (University of
Amsterdam) for his constructive feedback.
[4] A. Baron and P. Rayson. “VARD2: A tool for dealing with spelling variation in historical
corpora”. In:Postgraduate conference in corpus linguistics. 2008.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Anthonissen</surname>
          </string-name>
          . Individuality in Language Change. Berlin, Boston: De Gruyter Mouton,
          <year>2021</year>
          . doi: doi:10.1515/9783110725841.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R. H.</given-names>
            <surname>Baayen</surname>
          </string-name>
          . “
          <article-title>Corpus linguistics in morphology: Morphological productivity”C.oIrn-: pus Linguistics: An International Handbook</article-title>
          . Ed. by
          <string-name>
            <given-names>A.</given-names>
            <surname>Lüdeling</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kytö</surname>
          </string-name>
          . Vol.
          <volume>2</volume>
          . Berlin, New York: De Gruyter Mouton,
          <year>2009</year>
          , pp.
          <fpage>899</fpage>
          -
          <lpage>919</lpage>
          . doid:
          <source>oi:10.1515/9783110213881.2</source>
          .899.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Barðdal</surname>
          </string-name>
          .Productivity:
          <article-title>Evidence from Case and Argument Structure in Icelandic</article-title>
          . Amsterdam, Philadelphia: John Benjamins,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Brysbaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stevens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mandera</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Keuleers.</surname>
          </string-name>
          “
          <article-title>How Many Words Do We Know? Practical Estimates of Vocabulary Size Dependent on Word De昀椀nition, the Degree of Language Input and the Participant's Age”</article-title>
          .
          <source>In:Frontiers in Psychology</source>
          <volume>7</volume>
          (
          <year>2016</year>
          ). doi: 10.3 389/fpsyg.
          <year>2016</year>
          .
          <volume>01116</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Chiu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Villéger</surname>
            ,
            <given-names>I.-F.</given-names>
          </string-name>
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Thorn</surname>
            ,
            <given-names>Y.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-M. Chiang</surname>
            , and
            <given-names>W. B. Sherwin. “</given-names>
          </string-name>
          <article-title>An Attribute-diversity Approach to Functional Diversity, Functional Beta Diversity, and Related (Dis)Similarity Measures”</article-title>
          .
          <source>IEnc:ological Monographs 89.2</source>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .1002/ecm.1343.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Gotelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. C.</given-names>
            <surname>Hsieh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Sander</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. H.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Colwell</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Ellison</surname>
          </string-name>
          . “
          <article-title>Rarefaction and Extrapolation with Hill Numbers: A Framework for Sampling and Estimation in Species Diversity Studies”</article-title>
          .
          <source>InE:cological Monographs 84.1</source>
          (
          <issue>2014</issue>
          ), pp.
          <fpage>45</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.-H.</given-names>
            <surname>Chiu</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Chao</surname>
          </string-name>
          . “
          <article-title>Distance-Based Functional Diversity Measures and Their Decomposition: A Framework Based on Hill Numbers”</article-title>
          .
          <source>IPnL:oS ONE 9.7</source>
          (
          <year>2014</year>
          ). Ed. by F. de Bello,
          <year>e100014</year>
          .
          <source>doi:10.1371/journal.pone.010001.4</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Daly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Baetens</surname>
          </string-name>
          , and B. De Baets. “
          <article-title>Ecological diversity: measuring the unmeasurable”</article-title>
          .
          <source>In:Mathematics 6</source>
          .7 (
          <issue>2018</issue>
          ), p.
          <fpage>119</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          . “BERT:
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding”. PInro:ceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)</article-title>
          . Minneapolis, Minnesota: Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
          <year>do1i</year>
          :
          <fpage>0</fpage>
          .18653/v1/
          <fpage>N19</fpage>
          -1423.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Efron</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Thisted</surname>
          </string-name>
          . “
          <article-title>Estimating the Number of Unseen Species: How Many Words Did Shakespeare Know?”</article-title>
          <source>In:Biometrika 63.3</source>
          (
          <issue>1976</issue>
          ), p.
          <fpage>435</fpage>
          . doi:
          <volume>10</volume>
          .2307/2335721.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ellegård</surname>
          </string-name>
          . “Estimating Vocabulary Size”.
          <source>IWn:ord 16.2</source>
          (
          <issue>1960</issue>
          ), pp.
          <fpage>219</fpage>
          -
          <lpage>244</lpage>
          . doi:
          <volume>10</volume>
          .108 0/00437956.
          <year>1960</year>
          .
          <volume>11659728</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L.</given-names>
            <surname>Fonteyn</surname>
          </string-name>
          and
          <string-name>
            <surname>E. Manjavacas.</surname>
          </string-name>
          “
          <article-title>Adjusting scope: a computational approach to casedriven research on semantic change”</article-title>
          .
          <source>InP:roceedings of the Workshop on Computational Humanities Research (CHR</source>
          <year>2021</year>
          ). Vol.
          <volume>2898</volume>
          . CEUR Workshop Proceedings. Amsterdam,
          <year>2021</year>
          , pp.
          <fpage>280</fpage>
          -
          <lpage>298</lpage>
          . url: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2989</volume>
          /long%5C%
          <fpage>5Fpaper26</fpage>
          .p.df
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M. O.</given-names>
            <surname>Hill</surname>
          </string-name>
          . “
          <article-title>Diversity and Evenness: A Unifying Notation and Its Consequences”</article-title>
          .
          <source>In: Ecology 54.2</source>
          (
          <issue>1973</issue>
          ), pp.
          <fpage>427</fpage>
          -
          <lpage>432</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Hill</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Hengchen</surname>
          </string-name>
          . “
          <article-title>Quantifying the Impact of Dirty OCR on Historical Text Analysis: Eighteenth Century Collections Online as a Case Study”</article-title>
          .
          <source>IDni:gital Scholarship in the Humanities 34.4</source>
          (
          <issue>2019</issue>
          ), pp.
          <fpage>825</fpage>
          -
          <lpage>843</lpage>
          . doi:
          <volume>10</volume>
          .1093/llc/fqz02 4.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Hoover</surname>
          </string-name>
          . “
          <article-title>Another Perspective on Vocabulary Richness”</article-title>
          .
          <source>ICno:mputers and the Humanities 37.2</source>
          (
          <issue>2003</issue>
          ), pp.
          <fpage>151</fpage>
          -
          <lpage>178</lpage>
          . doi:
          <volume>10</volume>
          .1023/a:
          <fpage>1022673822140</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Jost</surname>
          </string-name>
          . “
          <article-title>Entropy and diversity”</article-title>
          .
          <source>InO:ikos 113.2</source>
          (
          <issue>2006</issue>
          ), pp.
          <fpage>363</fpage>
          -
          <lpage>375</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kubát</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Milička</surname>
          </string-name>
          . “
          <article-title>Vocabulary Richness Measure in Genres”</article-title>
          .
          <source>IJno:urnal of Quantitative Linguistics 20.4</source>
          (
          <issue>2013</issue>
          ), pp.
          <fpage>339</fpage>
          -
          <lpage>349</lpage>
          . doi:
          <volume>10</volume>
          .1080/09296174.
          <year>2013</year>
          .
          <volume>830552</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Fonteyn</surname>
          </string-name>
          . “
          <article-title>Adapting vs. Pre-training Language Models for Historical Languages”</article-title>
          .
          <source>In: Journal of Data Mining &amp; Digital Humanities Nlp4dh</source>
          (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .462 98/jdmdh.9152.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Fonteyn</surname>
          </string-name>
          . “
          <article-title>MacBERTh: Development and Evaluation of a Historically Pre-trained Language Model for English (1450-1950)”</article-title>
          .
          <source>PInro:ceedings of the Workshop on NLP4DH ICON</source>
          <year>2021</year>
          .
          <article-title>online: NLP Association of India (NLPAI</article-title>
          ),
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>F.</given-names>
            <surname>Perek</surname>
          </string-name>
          . “
          <article-title>Recent change in the productivity and schematicity of thweay -construction: A distributional semantic analysis”</article-title>
          .
          <source>InC:orpus Linguistics and Linguistic Theory</source>
          <volume>14</volume>
          .1 (
          <issue>2018</issue>
          ), pp.
          <fpage>65</fpage>
          -
          <lpage>97</lpage>
          . doi:
          <volume>10</volume>
          .1515/cllt-2016-0014.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Rao</surname>
          </string-name>
          . “
          <article-title>Diversity and dissimilarity coe昀케cients: a uni昀椀ed approach”</article-title>
          .
          <source>In: Theoretical population biology 21.1</source>
          (
          <issue>1982</issue>
          ), pp.
          <fpage>24</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Riba</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ginebra</surname>
          </string-name>
          . “
          <article-title>Diversity of vocabulary and homogeneity of literary style”</article-title>
          .
          <source>In: Journal of Applied Statistics 33.7</source>
          (
          <issue>2006</issue>
          ), pp.
          <fpage>729</fpage>
          -
          <lpage>741</lpage>
          . doi:
          <volume>10</volume>
          .1080/02664760600708970.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [24]
          <string-name>
            <surname>H.-J. Schmid</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mantlik</surname>
          </string-name>
          . “
          <article-title>Entrenchment in Historical Corpora? Reconstructing Dead Authors' Minds from their Usage Pro昀椀les”</article-title>
          .
          <source>InA:nglia 133.4</source>
          (
          <issue>2015</issue>
          ), pp.
          <fpage>583</fpage>
          -
          <lpage>623</lpage>
          . doi: doi: 10.1515/ang-2015-0056.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Segbers</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Schroeder</surname>
          </string-name>
          . “
          <article-title>How many words do children know? A corpus-based estimation of children's total vocabulary size”</article-title>
          .
          <source>ILna:nguage Testing 34.3</source>
          (
          <issue>2017</issue>
          ), pp.
          <fpage>297</fpage>
          -
          <lpage>320</lpage>
          . doi:
          <volume>10</volume>
          .1177/0265532216641152.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Shannon</surname>
          </string-name>
          .
          <article-title>“A Mathematical Theory of Communication”</article-title>
          .
          <source>InM: obile Computing and Communications Review 5 (I</source>
          <year>1948</year>
          ), p.
          <fpage>53</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Tweedie</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. H.</given-names>
            <surname>Baayen</surname>
          </string-name>
          . “
          <article-title>How Variable May a Constant be? Measures of Lexical Richness in Perspective”</article-title>
          .
          <source>In:Computers and the Humanities 32.5</source>
          (
          <issue>1998</issue>
          ), pp.
          <fpage>323</fpage>
          -
          <lpage>352</lpage>
          . doi:
          <volume>10</volume>
          .1023/a:
          <fpage>1001749303137</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>N.</given-names>
            <surname>Yáñez-Bouza</surname>
          </string-name>
          .
          <article-title>ARCHER 3.2: A Representative Corpus of Historical English Registers</article-title>
          . h ttps : / / www . projects . alc . manchester . ac . uk / archer / wp - content / uploads / 2020 / 06 /ARCHER_poster.pdf.
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>