<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using a Hybrid Algorithm for Lemmatization of a Diachronic Corpus</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Chelyabinsk State University</institution>
          ,
          <addr-line>Chelyabinsk, 454001</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1975</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Lack of lemmatization often undermines the quality of concordances, which is especially relevant for diachronic corpora. A significant part of lemmatizers are designed for Modern English. This paper presents MiddleEnglishLem, an application designed for dictionary-based lemmatization of Middle English texts. We use a hybrid algorithm to lemmatize the Helsinki Corpus of English Texts, a long-time-span diachronic corpus that includes Middle English texts of different genres, a total of 608 570 words. MiddleEnglishLem is capable of associating multiple inflected forms and orthographic varieties with canonical forms. Lemmatization becomes more accurate owing to comprehensive premade dictionaries. The competitiveness of this lemmatizer is proved by the low average errors - less than 2.5 percent, whereas its prebuilt stemmer has a strength of 0.38, a relatively high value. Accuracy of the lemmatization process can be improved by implementing syntagmatic analysis at the part-of-speech identification step. MiddleEnglishLem can be applied to diachronic corpora in order to research the development of English.</p>
      </abstract>
      <kwd-group>
        <kwd>Concordance</kwd>
        <kwd />
        <kwd>Lemmatization</kwd>
        <kwd />
        <kwd>Diachronic Corpus</kwd>
        <kwd />
        <kwd>Computer simulation</kwd>
        <kwd />
        <kwd>Hybrid Algorithm</kwd>
        <kwd />
        <kwd>Software Development</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The authors of this paper have carried out an experiment with the diachronic part of
the Helsinki Corpus of English Texts [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to verify a glottochronology-based
hypothesis. The hypothesis holds that diachronic changes of a vocabulary are predictable
using a specific set of mathematical models. Such calculus-based modelling uses
frequency rank tables [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The experiment is considered invalid as the original
Helsinki Corpus of English Texts is not lemmatized, meaning that concordances list
inflected forms as separate words, whereas we actually have to compare the frequency
of superlemmas in two time-separated states of the vocabulary. This is where the
issue of lemmatization arises, as the researchers are in need of better concordances. It
has, however, been found out that there has been so far developed no algorithm for
lemmatization of Middle English corpora. The existing lemmatizers/stemmers for
Modern English texts are obviously inappropriate for this task due to significant
morpho-logical discrepancies between these two diachronically separate forms of the
English language. It is therefore our additional objective to create a program that
would allow to lemmatize the Helsinki Corpus of English Texts.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Classification of Algorithms. Overview of Existing</title>
    </sec>
    <sec id="sec-3">
      <title>Software</title>
      <p>
        To begin the development of a proprietary algorithm for lemmatizing the Helsinki
Corpus of English Texts, we first of all had to decide which algorithm suits our needs
better and which existing software could probably be applied. V.A. Yatsko [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] states
that inflected words can be stem-associated by means of simple stemming or
lemmatization. Existing stemmers and their types can be classified as follows:
      </p>
      <p>
        Lemmatization, according to Yatsko, differs from stemming in the approach to
part-of-speech identification. Unlike stemmers of any type, lemmatizers identify parts
of speech and take them into account when associating inflected forms with their
respective lemmas. In the English language, words may be homographic yet
belonging to different parts of speech, making lemmatization a considerably more reliable
approach when it comes to concordance-building. Besides, algorithmic stemmers are
actually designed to reduce multiple inflected forms to their stems, and stems are not
always identical to canonical forms. Lemmatization, on the other hand, is always
aimed at returning such forms and not just stems. Stemming is therefore considered a
simplified alternative to lemmatization, which can be unsuitable for some research
objectives. S.Th. Gries and A.L. Berez, however, mention that multiple
stemming/lemmatization technologies can be combined to create a hybrid approach [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
which is going to be our case.
      </p>
      <p>As of today, two popular stemmers used for English corpora are the Porter
Stemming Algorithm by Martin Porter and the Lancaster Stemmer by Chris Paice and
Garth Husk. While comparable in terms of stemming strength when applied to
Modern English, none of these stemmers functions for Middle English, which is explicitly
stated by Mr. Porter himself on his webpage. Yet capable of automatic normalization
of Early New English texts, i.e. correcting their orthography in accordance with the
modern standards, the Porter stemmer will not cope with the complex morphology
and non-codified writing of Middle English. The same applies to MorphAdorner, a
popular lemmatizer which includes both the Porter stemmer and the Lancaster
stemmer as importable modules. Thus, no piece of software we have been able to test
could be used for our research, and an algorithm of our making has become a
necessity.
3</p>
    </sec>
    <sec id="sec-4">
      <title>Research Material: The Helsinki Corpus of English</title>
    </sec>
    <sec id="sec-5">
      <title>Texts</title>
      <p>The experiment mentioned in the introduction used data extracted from the Helsinki
Corpus of English Texts, which has a diachronic part covering the period of
7301710, i.e. from late Old English till Early Modern English. The total volume of the
corpus is about 1.5 million words (or word occurrences), which breaks into circa 450
texts belonging to philosophical, religious, scientific, fictional, educational and
instructive writing as well as private correspondence. Dialectal division is present as
well, with four dialects distinguished for Old English and five distinguished for
Middle English. However, the main criterion used to group text samples together is their
time of origins.</p>
      <p>The Helsinki Corpus of English Texts uses the COCOA tagging standard and is
therefore compatible with the Oxford Concordance Program. However, tagging is
only used in corpora files to indicate some non-linguistic or extra-linguistic
parameters of text samples, i.e. the date of creation, the date of manuscript-making, the
author and their status, the dialect, etc. Within sentences, tags are used to distinguish
Latin citations and editorial commentaries from the main text. That being said, the
corpus as available via the Oxford Text Archives is neither syntactically parsed nor
lemmatized. Due to the lack of lemmatization, the Oxford Concordance Program
counts each inflected or orthographically different word form as a separate lexeme,
which is why it returns incorrect statistical counts.</p>
      <p>
        For our main experiment, we chose two diachronically separate sections of the
corpus, Sections MEI and MEII (see [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for the explanation of such choice).
Therefore, we have a smaller corpus of approximately 210 thousand words to lemmatize,
which is still a too significant amount, rendering any attempts of manual
lemmatization unfeasible. For automatic lemmatization, we decided to use the same sections as
experimental samples to check the actual functionality and appropriateness of our
lemmatizing script. If properly programmed, it should return a lemma list that can be
imported in the Oxford Concordance Program so that the latter can build appropriate
and accurate concordances for the same sections.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Challenges of Middle English Lemmatization.</title>
    </sec>
    <sec id="sec-7">
      <title>Methodology</title>
      <p>The linguistic nature of Middle English is very challenging when it comes to language
processing; this paragraph is to analyze what kind of challenges we are facing and
how we can cope with them while developing our lemmatization algorithm.</p>
      <p>
        The first and most obvious challenge is the phenomenon of suppletion, i.e. use of
“inflected forms” that do not share any stem with their canonical forms. For instance,
we had such forms as us, ur(e), which were not morphological derivatives of the
dictionary form. Suppletion has existed since Proto-Germanic into Old English, changed
a little in Middle English [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and is present in Modern English as well, most strikingly
in pronouns.
      </p>
      <p>
        While manual lemmatization of suppletive forms is certainly not that difficult due
to the small number thereof, this issue is further complicated by another peculiarity of
Middle English, which is its non-codified orthography. Essentially, the graphical
representation of many consonantal and vocalic clusters was not standardized until the
18th century, resulting in single words having multiple orthographic varieties. Those
were not dialectal or even author-dependent, as even one text could contain, for
instance, both scylde and scilde (these are the same word, shield). C could be written
instead of K and vice-a-versa, ou and u were mutually replaceable as well, and the
same applies to such clusters as y/ge/i/ȝ/ghe (the prefix of the Participle II form).
Gradually, th came to replace letters “thorn” and “eth”, but such replacement was not
consistent until much later [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. A simple solution would be to consider such mutually
replaceable elements of writing as equivalent character combinations, but there might
have existed some instances where the use of one such element had a differential
effect. Non-codified orthography, however, is not limited to non-standardized use of
some clusters; there are cases where versions of a single word are barely
recognizable, e.g. bringen, the past (simple) form of which could be abrouhte and bryggte,
further complicating any attempts at automatic processing of such texts.
      </p>
      <p>
        Strong verbs represent another significant challenge, as some of their forms have
the root vowel changed, e.g. the past (simple) plural form of riden is roden. It should
be taken into account, however, that unlike weak verbs, strong verbs did not have
dental suffixes as markers of their past forms, a fact that may help enhance the
algorithm. The aforementioned participle II form of many verbs is another hereto related
problem, as it often had a y/ge/i/ȝ/ghe prefix (a similar phenomenon is observable in
Standard (or High) German of modern times). For instance, y-sungen was the
Participle II for syngen [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Therefore, the morph-truncating algorithm should also include
a verb-exclusive function to remove the prefix. This, however, may malfunction for
verbs like yelden, where y is not a prefix but a part of the root.
      </p>
      <p>The primary issue, however, is part-of-speech determination. On the one hand,
Middle English was rather rich in terms of morphology, and nominal parts of speech
had specific sets of morphs. On the other hand, many morphs did coincide for nouns
and adjectives. Finally, even having a very accurate preset list of morphs will not
enable appropriate differentiation of morphs and morph-like letter clusters. For
instance, for weorde, which is a verb, -de is a suffix, but for Franclonde, which is a
proper noun, it is not. Therefore, part-of-speech tagging is not always possible by
means of simple morph-truncating, meaning that the algorithm will require a preset
dictionary listing already PoS-tagged lemmas.</p>
      <p>
        So far, we have come to a simple yet labor-intensive solution to combine both
truncating and dictionary-based stemming, resulting in the creation of a
dictionarydependent lemmatizer. This is a very conventional and somewhat obsolete approach,
as modern algorithms mostly rely on finite-state transducers. However, we simply do
not have sufficient volumes of data to train and refine an automaton-based machine.
Besides, such an algorithm is far more difficult to develop, and we currently believe
that even the manual preparation of a PoS dictionary makes more sense in our case. A
dictionary like Mayhew and Skeat’s [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] can be used to build such lemma dictionary
("the lexicon") listing both canonical forms and possible orthographic varieties. Using
the lexicon will help distinguish nouns and verbs ending in similar letter
combinations, like the aforesaid Franclonde and weorde. The program will therefore have a
number of premade files, one of which will list possible morphs for the
morphtruncating script, and others will list actual lemmas and their orthographic varieties.
That being said, the program should be capable of:
searching for isolated (tokenized) words from Middle English samples in the
premade lexicon files;
truncating morphs at a length of 1 to 3 symbols from the last character of a word;
checking whether truncated morphs are found in the preset morph lists and whether
they correspond to parts of speech as per the lexicon;
correct association of inflected forms and lemmas and further registration of such
associations in the output file, which is to be imported in the Oxford Concordance
Program for further analysis of the Helsinki Corpus of English Texts;
dealing with the strong verbs and overall inflectional peculiarities of verbs as a
grammatical class;
returning an error log containing all the words that could not be lemmatized for the
subsequent manual lemmatization thereof.
      </p>
      <p>The next paragraph describes the implementation of these functions.
5</p>
    </sec>
    <sec id="sec-8">
      <title>Computational Implementation. Experiments</title>
      <p>For the purpose of automatic lemmatization of Middle English texts, we have
developed our own program titled MiddleEnglishLem using Python 2.7.9 as the
programming language. We have chosen this language because it provides well-developed
high-level data structures as well as a simple and efficient approach to object-oriented
programming. Besides, Python is perfect for script-making and fast development of
multi-platform applications for various purposes.</p>
      <p>The application we have developed uses a set of input files, one of which is a plain
text file that contains the corpus to process; another one is an .xslx file that
enumerates morphs and their respective PoS properties; the rest files are PoS-specific tabular
dictionaries, i.e. noun.xlsx, verb.xslx, etc., containing pre-associated lemmas and their
orthographic varieties as listed in Mayhew and Skeat’s (collectively, “the lexicon”).</p>
      <p>The output of the application is recorded in two separate files, one for lemmatized
words and their forms/varieties and one for errors (words the algorithm cannot
process).</p>
      <p>The algorithm functions by simple morph truncation. It truncates a sequence of one
symbol, adding one more if a single end symbol is not sufficient for processing.
Truncated morphs are searched for in the morph set, the remainder of the token is searched
for in the lexicon. As soon as the morph-associated part-of-speech tag matches that of
the stem as indicated in the dictionary, the application records the token from the
corpus in the tabulated output file under the lemma (if the latter is not present in the
output file, it is copied from the lexicon). If the PoS tag is V (verb), the algorithm
checks whether the participle II prefix is present and removes it. Verbal ablauts are
not dealt with specifically as the lexicons list past tense forms for strong verbs. If the
token cannot be processed after all steps are taken, it is written in the error log and
skipped; the application then proceeds to the next token.</p>
      <p>To test the efficiency of our script, we decided to use Zipf’s law as an indicator of
statistical representativeness:
f (k; s, N )</p>
      <p>
        1
k s H N ,s
where f is the relative frequency of a word in a corpus,
k is the rank of the word,
H is nth generalized harmonic number,
N is the number of words in the corpus,
s is the exponent value characterizing the Zipfian distribution of the text [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>The exponent was found by least squares calculations in the R programming
environment, where we used the lzipf package. It is believed that the value should be as
closed as possible to 1 for natural languages. For the raw (unprocessed) text of the
MEI subcorpus, it was 0.8546697, which does not meet the requirement above.
Whether the word frequency distribution in the sample matched or did not match
Zipf’s law was verified by Pearson’s chi-square test:
2
n (Oi
where Oi is the total of all actual frequencies in the ith interval,</p>
      <p>
        Ei is the total of projected (Zipfian) frequencies in the same interval [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. For
calculations, we divided the entire frequency table into 20 intervals with each subsequent
interval being shorter under the harmonic principle; thus we had 18 degrees of
freedom; at a 0.95 confidence interval, the chi-square distribution quantile for 18 degrees
of freedom equals 28.8693. The chi-square value we obtained per formula (2) for the
non-lemmatized sample was 68.5504 which meant that the sample did not match
Zipf’s law.
      </p>
      <p>The sample was then lemmatized using MiddleEnglishLem; exclusive of Latin
citations and proper names, the application returned 550 words, which amounted to
circa 3% of the number of ranks in the pre-lemmatization frequency table.
Lemmatization also reduced the number of such ranks by 34%, i.e. a third of the entire volume
became associated with other lexemes in the table. The s value for Zipf’s law was
recalculated and found equal to 0.9625564. The chi-square was recalculated as well
and equaled 26.52005. As this value was less than the quantile in our case,
postlemmatization word frequency distribution in the sample was confirmed to be in line
with Zipf’s law, a fact we believe proves that MiddleEnglishLem can help improve
the statistical representativeness of MiddleEnglishTexts.
6</p>
    </sec>
    <sec id="sec-9">
      <title>Conclusion: Unresolved Issues and Further</title>
    </sec>
    <sec id="sec-10">
      <title>Development</title>
      <p>
        We have so far developed an algorithm allowing to lemmatize Middle English texts at
a relatively low error rate; the built-in stemmer of our own making is considerably
strong due to the natural morphological complexity and relatively poor vocabulary of
Middle English. However, it will take more time and effort to prepare a full-fledged
lexicon and apply the algorithm to the Helsinki Corpus of English Texts Middle
English sections in their entirety. Besides, the algorithm still does not deal with some
orthographic ambiguities of this language, i.e. it is not capable of recognizing
character clusters with graphical varieties like c/k or u/ou. This may result in significant
understemming, if such varieties are not included in the input lexicon files. On the
other hand, some grammatical forms of different lexemes can be homographic, e.g. fet
as a 3SGPresInd form of a verb, and fet as a noun. This issue, which in rare cases may
lead to overstemming, can be solved by implementing syntagmatic analysis at the
part-of-speech identification step, as non-functional parts of speech naturally tend to
occur in certain syntactic structures [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The morphological complexity of verbs,
especially strong verbs, is also a problem to be solved; we can further address the way
it is dealt with in lemmatization algorithms for Standard German, where similar
complexity exists. These three issues will be our main priorities when attempting to
enhance the algorithm further.
      </p>
      <p>One more challenge we are facing that will require very thorough analysis is the
non-codified orthography of Middle English. While adding multiple orthographic
varieties to the lexicon is a suitable solution, it means our program is only
semiautomatic and still requires a lot of manual preparations. Use of finite-state
transducers, a completely different approach, could be a solution to this problem if we had
larger text samples for proper supervised machine learning. However, the approach
will be discussed in further works.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. Helsinki Corpus of English Texts, http://www.helsinki.fi/varieng/CoRD/corpora/HelsinkiCorpus/</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Arapov</surname>
            ,
            <given-names>M.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herz</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          : Mathematical Methods in Historical Linguistics. Nauka,
          <string-name>
            <surname>Мoscow</surname>
          </string-name>
          (
          <year>1974</year>
          ).
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Yatsko</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          :
          <article-title>Algorithms and programs for automatic text processing</article-title>
          . Bulletin of Irkutsk State Linguistic University, vol.
          <volume>1</volume>
          (
          <issue>17</issue>
          ), pp.
          <fpage>150</fpage>
          -
          <lpage>161</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hull</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          :
          <article-title>Stemming algorithms: a case study for detailed evaluation</article-title>
          .
          <source>Journal of the American Society for Information Science</source>
          , vol.
          <volume>47</volume>
          , N 1, pp.
          <fpage>70</fpage>
          -
          <lpage>84</lpage>
          (
          <year>1996</year>
          ). doi:
          <volume>10</volume>
          .1002/(SICI)
          <fpage>1097</fpage>
          -
          <lpage>4571</lpage>
          (
          <issue>199601</issue>
          )47:
          <article-title>1%3C70::AID-ASI7%3E3.0</article-title>
          .CO;
          <fpage>2</fpage>
          -%
          <fpage>23</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jivani</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>A Comparative Study of Stemming Algorithms</article-title>
          .
          <source>International Journal of Computer Technology and Applications</source>
          , vol.
          <volume>2</volume>
          (
          <issue>6</issue>
          ), pp.
          <fpage>1930</fpage>
          -
          <lpage>1938</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gries</surname>
            ,
            <given-names>S.Th.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berez</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Linguistic annotation in/for corpus linguistics</article-title>
          .
          <source>In: Nancy Ide &amp; James Pustejovsky (eds.)</source>
          ,
          <source>Handbook of Linguistic Annotation</source>
          . Berlin &amp; New York: Springer (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Karimov</surname>
          </string-name>
          , R.D.:
          <article-title>Predictive Modelling of the Development of Middle English vocabulary. Linguistics and Translation Issues Studied by Young Scientists, Nizhny Novgorod</article-title>
          , vol.
          <volume>1</volume>
          , pp.
          <fpage>189</fpage>
          -
          <lpage>198</lpage>
          (
          <year>2013</year>
          ).
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hogg</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>An Introduction to Old English</article-title>
          . Edinburgh University Press, Edinburgh (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ward</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waller</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Changes in the Language to the Days of Chaucer: Middle English Spelling</article-title>
          .
          <source>In: The Cambridge History of English and American Literature in 18 Volumes</source>
          , vol.
          <volume>1</volume>
          .
          <article-title>From the Beginnings to the Cycles of Romance</article-title>
          . Cambridge University Press, Cambridge (
          <year>1907</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ilyish</surname>
            ,
            <given-names>B.A.</given-names>
          </string-name>
          :
          <article-title>History of the English Language</article-title>
          .
          <source>Vysshaya Shkola</source>
          , Moscow (
          <year>1968</year>
          ).
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mayhew</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skeat</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>A concise dictionary of Middle English from A.D. 1150 to 1580</article-title>
          . Clarendon Press, Oxford (
          <year>1888</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Moreno-Sánchez</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Font-Clos</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corral</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Large-Scale Analysis of Zipf's Law in English Texts</article-title>
          .
          <source>PLoS ONE</source>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1371/journal.pone.0147073
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Preacher</surname>
            ,
            <given-names>K. J.</given-names>
          </string-name>
          :
          <article-title>Calculation for the chi-square test: An interactive calculation tool for chisquare tests of goodness of fit and independence</article-title>
          [Computer software] (
          <year>2001</year>
          ). Available from http://www.quantpsy.org/chisq/chisq.htm
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Sidorov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velasquez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chanona-Hernández</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Syntactic Dependency-Based N-grams: More Evidence of Usefulness in Classification</article-title>
          .
          <source>CICLing 2013. Part I. LNCS 7816</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>24</lpage>
          (
          <year>2013</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>642</fpage>
          -37247-
          <issue>6</issue>
          _
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>