<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ZipfExplorer: A Tool for the Comparison of Shared Lexis</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>English, University of Oulu</institution>
          ,
          <addr-line>90014 Oulu</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <fpage>145</fpage>
      <lpage>155</lpage>
      <abstract>
        <p>Word frequency statistics and lexical diversity measures can provide insights into discourse differences between texts. The ZipfExplorer, a tool and online app for the interactive visualization and comparison of word frequencies in two texts, shows side-by-side rank-frequency profiles and interactive tables of shared lexis, enabling keyword analysis and shedding light on discourse differences. Four lexical diversity measures (type-token ratio, Gini coefficient, powerlaw alpha parameter, and Shannon entropy) are calculated for the shared word types. Word frequency information is provided for a selection of mainly literary texts, and users can upload their own files. This paper provides an overview of the visualization of word frequency distributions, describes the functionality of the ZipfExplorer tool and demonstrates some of its features, and briefly discusses the lexical diversity measures calculated by the tool.</p>
      </abstract>
      <kwd-group>
        <kwd>Word Frequencies</kwd>
        <kwd>Visualization</kwd>
        <kwd>Lexical Diversity</kwd>
        <kwd>Zipf</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Word frequencies are a fundamental starting point for many analytical procedures in
corpus-based linguistic, literary, or cultural analysis and for natural language
processing tasks.1 The study of word frequency distributions and their statistical properties
continues to be an active topic of research in computational linguistics [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5 ref6 ref7 ref8">2, 3, 4, 5, 6,
7,8</xref>
        ], and in recent years, the analysis of word frequencies has been facilitated by the
availability of large corpora or other data sets and open access to data via platforms
such as CLARIN, GitHub, or the Center for Open Science as well as by dedicated
libraries of scripting functions in popular programming languages such as R or Python
[
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref9">9, 10, 11, 12</xref>
        ]. The representation of word frequencies in an interactive visualization
format, however, has not generally been a primary focus, despite the fact that interactive
visualizations can facilitate exploratory data analysis, enhance pedagogy, and
complement textual presentation of research [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ].
      </p>
      <p>
        In language and linguistics or literary or cultural studies, the comparison of word
frequencies in two texts or between a selected text and a reference corpus is a primary
method for gaining insight into differences in discourse content. The ZipfExplorer2 is
an online tool for the interactive visualization of word frequencies in texts or corpora,
1 This paper, an expanded version of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], includes a more detailed discussion of the Zipf
distribution and the lexical diversity measures calculated by the ZipfExplorer. In addition,
some code changes have been made to enhance the useability of the tool.
2 https://zipfexplorer.herokuapp.com
      </p>
      <p>
        Copyright © 2021 for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0).
named after Zipf’s Law [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ], the fact that for most longer natural language texts or
corpora, the frequency of a given word type is approximately inversely proportional to
its rank in a sorted list of the frequencies of word types for the text. The ZipfExplorer
provides an interactive means to show the concept of “keyness” [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ], or the extent
to which a lexical item occurs more often or less often than would be expected in
comparison to a reference text. The tool shows word frequency distributions for the textual
overlap of two texts, or the word types that they share, a text aspect that may also be of
theoretical interest in terms of its relationship to the concept of textual entailment, or
recognizing, given two text fragments, whether the meaning of one text can be inferred
(entailed) from the other [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], as well as to word error rate and derived measures of
textual similarity used in speech recognition [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. In addition, the tool, built using the
Bokeh module in Python [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], calculates several lexical diversity measures (type-token
ratio, Gini coefficient, power-law alpha parameter, and Shannon entropy). The code for
the tool is publicly available.3
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>Among the first to systematically study lexical type frequencies was the early
20th-century American Germanist George Zipf, who noted that when the words of a text are
ordered in decreasing frequency, the relationship between a the frequency and the rank
for a word of rank r can be expressed as   ≈   −1, where C represents a constant. A
Zipfian rank-frequency profile, when plotted in double logarithmic space, is typically
close to a straight line, but the shape of a frequency distribution for only those lexical
types that are shared with a comparison text or reference corpus depends not only on
the frequency information of the particular texts under consideration, but also on the
degree of textual overlap between the two texts. Visualizations of shared lexis, in
addition to highlighting discourse similarities and differences between texts through the
examination of particular word types, can also give insight into the interplay between
frequencies, derived lexical diversity measures, and the shape of discrete frequency
distributions in general.</p>
      <p>Following Zipf [15, pp. 45–48; 16, p. 25], word frequency distributions are typically
displayed in double logarithmic space, with frequency on the y-axis and frequency rank
on the x-axis, as in the top right quadrant of Figure 1, which shows four visualizations
of the word frequency distribution for Charles Dickens’ 1859 novel A Tale of Two
Cities. Each circle on the plot corresponds to a distinct word type. The most frequent type,
at the top left of the plot, is the word type “the”, occurring 8,058 times in the text,
followed by “and”, “of” and other common words.
3 https://github.com/stcoats/zipf_explorer
The plot in the top left represents the same information in linear space, whereas the
lower left plot is the so-called degree distribution (sometimes also referred to as the
frequency spectrum): Here, the word frequency counts themselves have been binned,
so that the top left circle is the proportion of all word types that occur once in the novel
(the hapax legomena). Hapax comprise 47% of the word types in the novel; words that
occur twice (dis legomena) 16%, and so on. While the information contained in the Zipf
rank-frequency plot and the degree distribution plot is equivalent, the latter plot is more
difficult to interpret in terms of discourse, as points on the plot do not correspond to
individual word types. In the bottom left of Figure 1, the complementary cumulative
distribution function is depicted: the cumulative proportion of types with a frequency
equal to or greater than a given frequency. Thus, 100% of types in the novel occur at
least once, 53% at least twice, 37% at least three times, and so on. The complimentary
cumulative distribution visualization is the reflection of the Zipf double-logarithmic
profile across a line extending from the bottom left to the top right of the subplot.
Because the upper two plots in Figure 1 are intuitively easier to understand, the
ZipfExplorer visualizes rank-frequency utilizes them, rather than the degree distribution or the
complementary cumulative distribution.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Tool Functionality</title>
      <p>
        In Figure 2 the default linear-scale view for the shared vocabulary types in Mary
Shelley’s Frankenstein and H. G. Wells’ War of the Worlds is depicted: Each subplot shows
the rank-frequency profile for the text selected via the dropdown menus to the right of
the plots. Points on the plots show word relative frequency (per 10,000 words) on the
y-axis and type rank in an ordered list of the frequencies of all words in the shared lexis
on the x-axis. Values for the lexical diversity measures type-token ratio, Gini
coefficient, alpha exponent of the best-fit power-law distribution, and Shannon entropy are
shown above the plots. Hovering over a word type will show its rank, frequency,
relative frequency, and the log-likelihood measure [
        <xref ref-type="bibr" rid="ref22 ref23">22, 23</xref>
        ] and associated p-value
compared to the shared lexis of the comparison text.
      </p>
      <p>Words can be highlighted with a hover tool and selected with a box-drawing tool (in
the toolset above the right-hand subplot). Selected words are highlighted in the sortable
tables below the plots; clicking on a word in one of the tables highlights it in the plots.
The tables show frequency rank, the word form, frequency, relative frequency,
difference in relative frequency compared to the other text, and the log-likelihood value:
higher log-likelihood values indicate are calculated for types with larger frequency
differences.</p>
      <p>
        The default texts available for comparison are selectable via a drop-down menu to
the right of the plots. In addition, users can upload their own texts for comparison with
the upload buttons. A ‘Remove most frequent words’ drop-down list removes 0, 10, 20,
50, 100, or 200 of the most frequent words in English, based on the Project Gutenberg
English Corpus from Sketch Engine [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. As many of the most frequent words are
determiners, prepositions, conjunctions, or other function words that bear relatively little
semantic information, removing frequent words can help to highlight content and
discourse differences between the texts. Below the remove words drop-down menu, the
total number of types and tokens in the original texts is shown along with the percentage
of types that are shared in the two texts. To examine the word frequency distribution of
a single text, rather than the distributions of the shared lexis in two texts, the same text
can be selected for both plot windows.
      </p>
      <p>
        The source texts are a selection of mainly literary texts from Project Gutenberg, a
corpus of inaugural addresses of U.S. presidents from NLTK [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], the Brown Corpus
and its subsections [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], and the Freiburg-Brown Corpus of American English [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
3.1
      </p>
      <sec id="sec-3-1">
        <title>Sorting</title>
        <p>The columns in the tables below the subplots can be sorted. They show original word
order in the left-hand text, word form, rank in the frequency table, relative frequency,
difference to the other text in relative frequency, and log-likelihood score. Sorting can
show items that are much more relatively frequent in a text. In Figure 2, the personal
pronouns ‘my’, ‘you’, and ‘I’ are more frequent in Frankenstein, a text with a
firstperson point of view, than in the third-person War of the Worlds.</p>
        <p>The types ‘up, ‘out’, and ‘there’ (Fig. 4) are more relatively frequent in War of the
Worlds; when considered along with other prepositions, place adverbials and location
names, it becomes clear that spatial organization plays a greater role as a narrative
element in War of the Worlds than in Frankenstein.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Hapax Types</title>
        <p>Hapax can also shed light on discourse differences. Highlighting the hapax types in
Frankenstein (in the left-hand rank-frequency profile of Fig. 5) shows their ranks and
relative frequencies in War of the Worlds: although many are also hapax in the other
text, or are found mainly in the tail of the frequency distribution for Frankenstein, the
types ‘smoke’ and ‘red’ are much higher in the profile in War of the Worlds – a
frequency difference that reflects the discourse content of the latter novel.
Using the drop-down menu to the right of the subplots, 10, 20, 50, 100, or 200 of the
most-frequent types in the Gutenberg corpus can be removed from the visualizations
and tables. Because these types, which are mostly function words such as determiners,
pronouns, and relativizers or common verbs, structure texts in important ways but
contribute relatively little to discourse content, removing them may serve to highlight
discourse differences between two texts.</p>
        <p>In terms of the distribution shape and the derived lexical diversity statistics, the
removal of common words has the effect of increasing the relative frequency of the
remaining words. In effect, removing stopwords tends to change the shape of the Zipf
profile in double-logarithmic space to a more curvilinear form – when function words
are no longer considered, word frequencies deviate substantially from a power-law
distribution. As can be expected, removal of common words tends to increase the lexical
diversity of the texts for the remaining shared types, which tend to be more uniformly
distributed in terms of their relative frequencies.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Lexical Diversity</title>
      <p>has a maximum theoretical value of log2(n), for data consisting of n unique types.</p>
      <p>
        The diversity statistics calculated by the tool provide evidence for the sensitivity of
lexical diversity measures to sample size [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ]. For the textual overlap between two
texts or a text and a corpus, the shorter text will likely exhibit lower Gini values and a
4 In these cases, however, word frequencies are unlikely to be distributed according to a power
law, and thus the measure is not necessarily a good diversity indicator. See Clauset, Shalizi
and Newman (2009).
      </p>
      <p>
        The ZipfExplorer displays four lexical diversity measures: the type-token ratio, the Gini
coefficient, the exponent α for the best fit of a power-law distribution, and the Shannon
entropy  [
        <xref ref-type="bibr" rid="ref28 ref3 ref4">3, 4, 28</xref>
        ]. These measures, while related, can be used to highlight different
aspects of lexical diversity. The type-token ratio,
the interval (0,1], with smaller values indicating less lexical diversity.
      </p>
      <p>has a range in
The Gini coefficient, which can be calculated with
=
2

∑  
∑  
−




+ 1
ranges from 0 (no diversity) to 1 (maximum diversity) for n word types with relative
frequencies x.
1 +
1

( =</p>
      <p>−1</p>
      <p>
        The exponent α results from the best-fit line to the degree distribution function
(lower left plot in Fig. 1) for the frequency information for the shared lexical types,
calculated using the powerlaw package in Python [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] with the equation  ( ) ∝  − .
The alpha parameter is related to the slope z of the Zipf rank-frequency profile by  =
1 ) [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. The parameter typically ranges in value between ~1.5 and 3,
although higher or lower values are calculated for the shared vocabulary of extremely
dissimilar or extremely short texts.4
      </p>
      <p>
        Shannon entropy [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], calculated with


=
−
∑
      </p>
      <p>log2  


higher type-token ratio, whereas the longer text will exhibit a smaller α exponent and a
higher  value. Removing frequent words will often increase values for the type-token
ratio and the alpha parameter, and decrease the values for the Gini coefficient and the
Shannon entropy, although this depends on the texts in question, their original
frequencies, and the degree of textual overlap. For texts with a relatively large proportion of
shared types, such as two novels by the same author, and with the removal of frequent
function words, the lexical diversity measures may give insight into topical diversity in
terms of narrative development. For texts that share relatively few types, the
relationship between the measure values and the properties of the underlying original texts is
less straightforward.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The ZipfExplorer enables the interactive exploration of word frequencies in the shared
lexis of two comparison texts or corpora, potentially shedding light on discourse
similarities and differences and properties of frequency distributions. The lexical diversity
measures type-token ratio, Gini coefficient, alpha parameter of the power-law function,
and Shannon entropy, calculated by the tool, vary according to text length and textual
overlap and are also affected by the removal of common function words.</p>
      <p>In a pedagogical context, the ZipfExplorer provides a hands-on way to make
frequency information concrete. Given the increasing importance of artificial intelligence
models not only in linguistics and other sciences, but ultimately in many working-life
and administrative domains and in the contexts of daily life, the tool can serve as a
starting point for understanding how linguistic frequency distributions underlie the
large data sets used train machine learning models.</p>
      <p>The tool may also be useful for the comparison of various discrete distributions in
computational studies of language or digital humanities, and for applied analysis in
literary, historical, or cultural studies in which “distant reading” approaches are
employed. Planned further development of the tool is to allow upload of different file
formats, enable text extraction from URLs, and enable automatic annotation of
part-ofspeech tags or named entities whose frequency distributions may be of interest. It is
also hoped that other researchers will use the code for the tool (or parts thereof),
available at GitHub, in order to create new and exciting ways to visualize linguistic data
such as word frequency information.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Coats</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Comparing word frequencies and lexical diversity with the ZipfExplorer tool</article-title>
          . In Sanita Reinsone, Inguna Skadiņa, Anda Baklāne and Jānis Daugavieti (eds.),
          <source>Proceedings of the 5th Digital Humanities in the Nordic Countries Conference</source>
          , Riga, Latvia,
          <source>October 21- 23</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>225</lpage>
          . CEUR, Aachen, Germany (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baayen</surname>
            ,
            <given-names>R. H.</given-names>
          </string-name>
          :
          <article-title>Word frequency distributions</article-title>
          . Kluwer, Dordrecht (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bérubé</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sainte-Marie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mongeon</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larivière</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Words by the tail: Assessing lexical diversity in scholarly titles using frequency-rank distribution tail fits</article-title>
          .
          <source>PLoS ONE</source>
          <volume>13</volume>
          (
          <issue>7</issue>
          ) (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Clauset</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shalizi</surname>
            ,
            <given-names>C. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>M. E. J.</given-names>
          </string-name>
          :
          <article-title>Power-Law distributions in empirical data</article-title>
          .
          <source>SIAM Review</source>
          <volume>51</volume>
          (
          <issue>4</issue>
          ),
          <fpage>661</fpage>
          -
          <lpage>703</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lü</surname>
            ,
            <given-names>L</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>Z.-K.</given-names>
            ,
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <article-title>Zipf's law leads to Heaps' law: Analyzing their relation in finite-size systems</article-title>
          .
          <source>PLoS One</source>
          <volume>5</volume>
          (
          <issue>12</issue>
          ) e14139 (
          <year>2010</year>
          ). https://doi.org/10.1371/journal.pone.0014139
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Montemurro</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          :
          <article-title>Beyond the Zipf-Mandelbrot law in quantitative linguistics</article-title>
          .
          <source>Physica A: Statistical Mechanics and its Applications</source>
          <volume>300</volume>
          (
          <issue>3-4</issue>
          ),
          <fpage>567</fpage>
          -
          <lpage>578</lpage>
          (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>M. E. J.:</given-names>
          </string-name>
          <article-title>Power laws, Pareto distributions and Zipf's law</article-title>
          .
          <source>Contemporary Physics</source>
          <volume>46</volume>
          (
          <issue>5</issue>
          ),
          <fpage>323</fpage>
          -
          <lpage>351</lpage>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Piantadosi</surname>
            ,
            <given-names>S. T.</given-names>
          </string-name>
          <article-title>Zipf's word frequency law in natural language: A critical review and future directions</article-title>
          .
          <source>Psychonomic Bulletin &amp; Review</source>
          <volume>21</volume>
          (
          <issue>5</issue>
          ),
          <fpage>1112</fpage>
          -
          <lpage>1130</lpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Alstott</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bullmore</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plenz</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Powerlaw: A Python package for analysis of heavy-tailed distributions</article-title>
          .
          <source>PLoS ONE</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ) (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Baayen</surname>
            ,
            <given-names>R. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shafaei-Bajestan</surname>
          </string-name>
          , E.: languageR: Analyzing Linguistic Data: A Practical Introduction to Statistics.
          <source>(R package version 1.5.0)</source>
          . https://CRAN.Rproject.org/package=languageR (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Evert</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>zipfR: Word frequency distributions in R (R package version 0</article-title>
          .
          <fpage>6</fpage>
          -
          <lpage>10</lpage>
          of 2017-
          <volume>08</volume>
          -17).
          <source>In: Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics, Posters and Demonstrations Sessions</source>
          , pp.
          <fpage>29</fpage>
          -
          <lpage>32</lpage>
          , ACL, Stroudsburg, PA (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Gillespie</surname>
            ,
            <given-names>C. S.:</given-names>
          </string-name>
          <article-title>Fitting heavy tailed distributions: The poweRlaw package</article-title>
          .
          <source>Journal of Statistical Software</source>
          <volume>64</volume>
          (
          <issue>2</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          . http://www.jstatsoft.org/v64/i02/ (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Cleveland</surname>
          </string-name>
          , W. S.:
          <article-title>Visualizing data</article-title>
          . Hobart Press, Summit, NJ (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Wilkinson</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>The grammar of graphics</article-title>
          , Springer, New York (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zipf</surname>
            ,
            <given-names>G. K.</given-names>
          </string-name>
          :
          <article-title>The psycho-biology of language</article-title>
          . Routledge, London (
          <year>1936</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zipf</surname>
            ,
            <given-names>G. K.</given-names>
          </string-name>
          :
          <article-title>Human behavior and the principle of least effort</article-title>
          .
          <source>Addison-Wesley</source>
          , Cambridge, MA (
          <year>1949</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tribble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Textual patterns</article-title>
          .
          <source>John Benjamins</source>
          , Amsterdam (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Stubbs</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Three concepts of keywords</article-title>
          . In: Bondi,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Scott</surname>
          </string-name>
          , M. (eds.), Keyness in texts, pp.
          <fpage>21</fpage>
          -
          <lpage>42</lpage>
          . John Benjamins, Amsterdam (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Androutsopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malakasiotis</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A survey of paraphrasing and textual entailment methods</article-title>
          .
          <source>Journal of Artificial Intelligence Research</source>
          <volume>38</volume>
          ,
          <fpage>135</fpage>
          -
          <lpage>187</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Morris</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maier</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Green</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <string-name>
            <surname>From</surname>
            <given-names>WER</given-names>
          </string-name>
          and
          <article-title>RIL to MER and WIL: Improved evaluation measures for connected speech recognition</article-title>
          .
          <source>In: Proceedings of INTERSPEECH 2004 - ICSLP, 8th International Conference on Spoken Language Processing</source>
          , pp.
          <fpage>2765</fpage>
          -
          <lpage>2768</lpage>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21. Bokeh Development Team. Bokeh:
          <article-title>Python library for interactive visualization</article-title>
          . http://www.bokeh.pydata.org,
          <source>last accessed</source>
          <year>2019</year>
          /09/30.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Dunning</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Accurate methods for the statistics of surprise and coincidence</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>19</volume>
          ,
          <fpage>61</fpage>
          -
          <lpage>74</lpage>
          (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Rayson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garside</surname>
          </string-name>
          , R.:
          <article-title>Comparing corpora using frequency profiling</article-title>
          .
          <source>In: WCC '00 proceedings of the workshop on comparing corpora</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . ACM, New York (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baisa</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bušta</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakubíček</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovář</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michelfeit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rychlý</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suchomel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The Sketch Engine: ten years on</article-title>
          .
          <source>Lexicography 1</source>
          ,
          <fpage>7</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loper</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
          </string-name>
          , E.:
          <article-title>Natural language processing with Python updated for NLTK 3.0</article-title>
          . Newton, MA,
          <string-name>
            <given-names>O</given-names>
            <surname>'Reilly</surname>
          </string-name>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Francis</surname>
            ,
            <given-names>W. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kučera</surname>
          </string-name>
          , H.:
          <article-title>A standard corpus of present-day edited American English, for use with digital computers</article-title>
          . Brown University, Providence, RI (
          <year>1979</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Hundt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sand</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skandera</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Manual of information to accompany The Freiburg - Brown Corpus of American English ('Frown')</article-title>
          . Department of English,
          <string-name>
            <surname>Albert-LudwigsUniversität Freiburg</surname>
          </string-name>
          , Freiburg, Germany (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Kunegis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Preusse</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Fairness on the web: Alternatives to the power law</article-title>
          .
          <source>In: Proceedings of WebSci</source>
          <year>2012</year>
          , June 22-24,
          <year>2012</year>
          , pp.
          <fpage>175</fpage>
          -
          <lpage>184</lpage>
          . ACM, New York (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Adamic</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Zipf, power-laws, and Pareto-a ranking tutorial</article-title>
          . https://www.hpl.hp.com/research/idl/papers/ranking/ranking.html,
          <source>last accessed</source>
          <year>2020</year>
          /12/04.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Shannon</surname>
            ,
            <given-names>C. E.</given-names>
          </string-name>
          :
          <article-title>A mathematical theory of communication</article-title>
          .
          <source>Bell System Technical Journal 27</source>
          ,
          <fpage>379</fpage>
          -
          <lpage>423</lpage>
          ;
          <fpage>623</fpage>
          -
          <lpage>656</lpage>
          (
          <year>1948</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>