<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Challenge of Vernacular and Classical Chinese Cross-Register Authorship Attribution</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haining Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xin Xie</string-name>
          <email>rwxiexin@shnu.edu.cn</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Allen Riddell</string-name>
          <email>riddella@indiana.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indiana University Bloomington</institution>
          ,
          <addr-line>Bloomington, Indiana</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Shanghai Normal University</institution>
          ,
          <addr-line>Shanghai</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <fpage>299</fpage>
      <lpage>309</lpage>
      <abstract>
        <p>Ming-Qing fiction is widely regarded as the pinnacle of classical Chinese literature, but over threequarters of vernacular fictional works were anonymously or pseudonymously composed, frustrating literary-historical research. To begin to address the problem, we propose a cross-register authorship attribution task: recover the authorship of a vernacular Chinese text given classical Chinese writing samples of known authorship. A corpus of eight authors known to have written in both registers was assembled to serve as a testbed. We describe the performance of models using diferent sets of function character/word frequencies as input features. This standard approach to authorship attribution performs well in the same-register setting but poorly in the cross-register setting. We discuss the degree of vernacularization and the amount of dialog in texts as key factors contributing to the low cross-register accuracy.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;authorship attribution</kwd>
        <kwd>Ming-Qing fiction</kwd>
        <kwd>classical Chinese</kwd>
        <kwd>vernacular Chinese</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>1.1. Backgrounds</title>
        <p>
          the previously mentioned factors. In practice, candidate authors are usually already provided
to researchers thanks to the labor of literary and cultural historians. The challenge lies in
matching writing styles. Numerous factors reportedly influence writing style, including genre
[
          <xref ref-type="bibr" rid="ref14">17, 8, 14</xref>
          ], topic [
          <xref ref-type="bibr" rid="ref15 ref9">15, 16, 9</xref>
          ], gender [
          <xref ref-type="bibr" rid="ref13 ref6">6, 13</xref>
          ], period [
          <xref ref-type="bibr" rid="ref1 ref5">5, 1</xref>
          ], and culture [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Typically researchers
begin by finding, for each candidate author, writing samples which resemble—in terms of
the previously mentioned factors—the disputed text. Ming-Qing vernacular fiction poses a
particular challenge here: candidate authors tended not to sign any vernacular works. In most
cases, the texts we have available were written in classical Chinese.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2. Classical Chinese, Vernacular Chinese, and Cross-Register authorship attribution</title>
        <p>
          Classical Chinese can be understood as preserving the grammar and semantics of Chinese
as it was used before the Qin period (i.e., before 221 BCE). Classical Chinese predominates
in official texts and texts written by members of the educated class throughout the imperial
period (221 BCE - 1912 CE) [
          <xref ref-type="bibr" rid="ref18">20</xref>
          ]. For example, most of official documents were composed
using classical Chinese.
        </p>
        <p>
          Written vernacular often emerged from written dialog in classical works. This way of writing
developed into various genres. For example, Bianwen ( ) paraphrases canonical Buddhist
texts using speech-like writing. And Huaben ( ) describes actors’ movements and scripts
when performing. It was not until the middle sixteenth century before vernacular written
Chinese became a recognized literal register [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>The diferences between the two versions of written Chinese are considerable. First, the
classical lexicon tends to use single characters, while vernacular words often use pairs of
characters. Take an excerpt from the Mencius as an example (Figure 1). The Mencius is a classic
of Confucianism composed in classical Chinese. In the example, “ ” (being hungry) is used
in isolation, but in vernacular is expected to be collocated with “ ” (“ ”) to express the
same meaning. Second, classical has more frequent part-of-speech ambiguity. “ ” (clothes) is
a noun in vernacular most of the time, but it functions as a verb when used before “ ” (silk),
meaning “wear.” Third, word order in the classical register is more variable. In most cases,
Chinese uses subject-verb-object order. In the classical clause “ ”, “ ” is the object
and appears before the linking verb “ ”. This order is unconventional in vernacular Chinese.</p>
        <p>Frequently, especially during the Ming and Qing periods, the boundary between vernacular
and classical Chinese is not clear. Many texts mix the two registers in various ways. For
example, dialog in classical texts often resembles the vernacular equivalent. Vernacular fiction
also has a tradition of opening and closing a chapter with classical verse. The boundary
blurred further when classical grammar was mixed with the vernacular lexicon at the end
of the imperial period. In addition, written Chinese shares the same convention of having
no obvious break markers (corresponding to punctuation) between “sentences” and delimiting
paragraphs with empty spaces after the ending of a previous paragraph.1</p>
        <p>In this study, we consider the task of cross-register authorship attribution using texts of
known authorship. We aim to infer the authorship of a vernacular text given classical texts by
the same author. Developing reliable cross-register authorship attribution techniques will be
required to resolve the long-standing debates about disputed authorship of vernacular fictions,
such as the Golden Plum Vase and the Marriage Destinies to Awaken the World, and roughly
four hundred works as we found in our survey.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Corpus</title>
      <sec id="sec-2-1">
        <title>2.1. Description</title>
        <p>We assemble a corpus of eight authors known to have written in both registers. All authors
lived between 1570 and 1870. All but one are from southern China, and all authors are men.2
The imbalance of gender and region reflects relevant social and economic circumstances of the
period.3</p>
        <p>The corpus contains 4.2 million characters of fictional and non-fictional prose, although
these genres are not always distinct.4 Topics are diverse, including jokes, the care of pregnant
women, history, opera commentary, war diaries, and personal reflections. All works address a
general audience.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Collecting and Preprocessing</title>
        <p>In practice, only authors who have written in each register and whose surviving works have
machine-readable editions are considered.5 We refrain from picking texts that are disputed
or use rhyme. We also avoid texts which are revisions of pre-existing texts or mixtures of
vernacular glosses with classical grammar. Table 1 shows the texts in the corpus.</p>
        <p>After downloading all texts, we performed the following preprocessing:
1. We checked all texts for flaws and missing parts by consulting other digital editions and
print editions.6 If a character in a text lacks a Unicode code point, we used the modern
1Readers familiar with the evolution of Latin may gain some appreciation of how the registers difered by
considering the lexical and syntactic diferences between Classical Latin (75 BCE to 300 CE) and Modern Latin
(ca. 1500 - 1900 CE). The analogy is not exact, of course.</p>
        <p>
          2We spent 20 hours searching for woman authors to include, but we were unable to find an author with
available texts. We gathered the candidate authors by consulting two bibliographies of Ming-Qing novels[
          <xref ref-type="bibr" rid="ref21">23, 7</xref>
          ].
We welcome suggestions for candidates to include in a future, expanded version of the corpus.
        </p>
        <p>3At the time, education for women received less attention. And, southern China, such as Zhejiang and
Jiangsu, was relatively wealthier and developed.</p>
        <p>4Distinguishing between fictional and non-fictional historical narratives is difficult or, in some cases,
impossible.</p>
        <p>5The existence of an edition in a machine-readable format usually indicates a canonical author. This
introduces a bias towards authors who were well known at the time or who subsequently became well known.
6If multiple versions exist, we choose the one which has fewer characters missing or obvious errors.
variant. We refer to zdic.net and Unicodepedia to check whether a character falls outside
of the Chinese Unicode set (between \u4e00 and \u9ff). We keep outside characters if
they are valid Chinese characters or punctuation, deleting them otherwise.7 Characters
are rarely deleted.
2. We remove title, heading, table of contents, preface, postscript, and editorial comments,
as well as rhymed verse and prose if they are not part of the main body. Blank lines,
redundant spaces, and indentation marks are removed too.
3. All the texts are automatically converted into UTF-8 encoded simplified Chinese using
the Python package “hanziconv” (v.0.3.2).
4. Texts are then segmented into roughly 1,000-character chunks without breaking
clauselevel structures. Modern publishers punctuate ancient Chinese texts, which originally
did not have punctuation. We leverage these delimiters introduced by editors to avoid
breaking clauses when segmenting.
5. We eliminate all punctuation in the next step to restore the original formatting.
6. For authors who have only one work in a register, we split the work into two parts.8
After performing these steps, we organize all chunks by register under each candidate’s
directory with informative file names.</p>
        <p>The corpus consists entirely of texts in the public domain and is available at https://zenodo.
org/record/5513043.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <sec id="sec-3-1">
        <title>3.1. Feature Set and Algorithm</title>
        <p>
          Given our research question—Is it possible to recover the authorship of a vernacular text using
classical texts as training data?— we choose function words/characters for features because
they have been shown to be useful stylistic markers [
          <xref ref-type="bibr" rid="ref20 ref22">22, 24</xref>
          ]. We transcribed two published
lists of function words—one for classical and the other for modern—as the feature sets.9 The
classical feature set contains 479 function characters [
          <xref ref-type="bibr" rid="ref19">21</xref>
          ]; the modern feature set has 819
function words (262 character unigrams, 545 bigrams, 10 trigrams, 2 tetragrams) [
          <xref ref-type="bibr" rid="ref16">18</xref>
          ].10 We
also use a feature set which is the union of the two feature sets (the “combined” feature set).
        </p>
        <p>
          A linear support vector machine (SVM) is chosen as the classifier. We use LIBSVM’s
implementation [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], wrapped by scikit-learn [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. We use the default cost parameter (C = 1.0).
Features are standardized by dividing by feature standard deviations after deducting the means.11
3.2. Setup
The goal of our cross-register task is to recover the authorship of a vernacular text based on
classical writing samples. To this end, we set up two experiments. In the first experiment, we
7Some valid Chinese characters fall outside of the aforementioned Unicode range.
        </p>
        <p>8We do this to prevent severe inflation in calculating same-register accuracy. See justification for this
treatment in the Appendix.</p>
        <p>
          9Chinese vernacular function words overlap heavily with modern Chinese’s. Also, there is no function word
dictionary built for vernacular Chinese specifically, to the best of our knowledge. The modern function word
list [
          <xref ref-type="bibr" rid="ref19">21</xref>
          ] is particularly comprehensive.
        </p>
        <p>
          10We released a Python package (“functionwords”) on PyPI to help others use these lists.
11A pilot study shows a standard logistic regression with L2 regularization (regularization parameter 1.0)
achieves similar accuracy as SVM [
          <xref ref-type="bibr" rid="ref17">19</xref>
          ].
Hard Gourd Collection classical
Romance of the Sui and Tang Dynas- vernacular
ties
Wenmu Collection
The Scholars
Romance of the Northern Dynasties
Romance of the Southern Dynasties
Amusing and Awakening
Pregnancy &amp; Childbirth, A Revision
Dream of the Red Chamber, An
unfinished Twenty-Chapter Complement
        </p>
        <p>Register
classical
classical
vernacular
vernacular
vernacular
vernacular
classical
vernacular
classical
classical
vernacular
classical
vernacular
vernacular
vernacular
classical
vernacular
classical
classical
vernacular
classical
vernacular</p>
        <p>Genre
consider authorship attribution in a scenario when a reasonable amount of classical training
text is available. In the second experiment, we consider a scenario in which extensive classical
training material is available.</p>
        <p>
          We do not have a golden rule for how many classical Chinese characters constitute a large
amount of training material. For English language authorship attribution, Rao, Rohatgi, et al.
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] recommends around 6,500 English words as adequate. We estimate the corresponding
character count in classical Chinese by counting English words used in the first ten stories
of Herbert Giles’ translation of Strange Stories from a Chinese Studio ( ) [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] (ca.
14,190 words) and Chinese characters of the corresponding plots in the original work (ca.
9,350 Chinese characters). The English-word-to-Chinese-character ratio is roughly 1.5. We
ifnally decide to fit the model with ca. 4,000 classical characters from each author for the
“limited data” scenario. In the other experiment—the “abundant data” scenario—we give the
model more training material., ca. 10,000 classical characters. For both settings, we evaluate
the model on its ability to predict the authorship of ca. 1,000-character vernacular texts.
Predictive success is measured using accuracy.
        </p>
        <p>The experimental procedure is described using the limited data scenario. For a given
candidate size from two to eight, we randomly choose four classical chunks from each author to
ift the classifier. With the same trained classifier, we make predictions on test samples from
diferent registers. The first prediction is made on ca. 1,000 vernacular characters, where
the samples are randomly chosen from each author’s vernacular writing. We care about the
classical-vernacular task the most because it shows how well a classifier trained with classical
texts can successfully infer vernacular texts’ authorship.</p>
        <p>The other prediction, the classical-classical task, uses 1,000 classical characters from each
author. The classical-classical task’s testing samples are chosen from documents or parts of
documents that are not in the training set to avoid inflating the accuracy (see the Appendix
for a discussion of this concern). This task works as a baseline by indicating the level of
accuracy the same classifier can perform with “vanilla” authorship attribution with classical
Chinese. We also use random chance— candida1te size ×100%—as another baseline. The accuracy
of the cross-register task should be bounded from above by the classical-classical accuracy and
bounded from below by chance. The experiment is performed 2,000 times for every candidate
size.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>We predict the authorship of 1,000-character vernacular texts and 1,000-character classical
Chinese texts with three feature sets under the limited data and abundant data scenarios. The
classical-vernacular and classical-classical experiments use a standard linear SVM trained on
the same training texts but evaluated on texts written in a diferent register (see Figure 2).</p>
      <p>The classical-vernacular accuracy barely deviates from chance in the limited data scenario.
Classical-vernacular accuracy is very slightly better than chance when trained on 10,000
character texts. In contrast, the classical-classical task, trained on the same data, performs far
better than chance. With the optimal setting (combined feature set and 10,000 characters
training data), the mean accuracy for classical-vernacular, classical-classical, and chance are
25.3%, 83.4%, and 24.5%, in turn.</p>
      <p>The function characters/words feature sets play a trivial role in determining the cross-register
accuracy but afect the classical-classical accuracy.</p>
      <p>Confusion Matrix Confusion matrices show a similar pattern. For brevity, we only draw the
confusion matrices under the abundant data scenario with the combined feature set assigning
authorship given eight candidates (see Figure 3). The confusion matrices are normalized by
rows (true labels).</p>
      <p>In the confusion matrices, each row represents the true author, and each column represents
the predicted author. Taking Feng Menglong (FML) as an example, the first row (labeled by
“FML”) indicates the empirical probability of the linear SVM predicting the author indicated
by the bottom labels when the true author is FML. In the classical-classical task (the left
panel), FML has a probability of 0.6 to be predicted correctly; Du Gang (DG), Ding Yaokang
(DYK), Wu Jingzi (WJZ), Chu Renhuo (CRH), and Li Yu (LY) have probabilities of 0.2, 0.1,
0.08, 0.05, and 0.006 to be misclassified as FML, respectively. However, CRH, DG, and LY
are more likely to be predicted to be FML than the true author FML when it comes to a
cross-register situation (the right panel).</p>
      <p>In the classical-classical task, the values on diagonal are higher than random guess ( 18 =
0.125). Rarely is the classifier confused by authors with a similar writing style in the classical
register. The most difficult case is predicting WJZ’s classical texts. Though CRH is a
competitive candidate with a probability of 0.3 being assigned, the probability of correct prediction
(0.4) is highest.</p>
      <p>In the right panel, the probability of correct classification (the diagonal values) are much
lower for classical-vernacular authorship attribution. Instead, CRH tends to be predicted as
the author, regardless of the true author. DG, LY, and FML also attract incorrect attributions.
Other authors receive little attention from the classifier.</p>
      <p>A follow-up experiment Since CRH and FML are among favorite candidates in the
crossregister task and both are known for their vernacular works, we investigate whether
vernacularization plays a part in the main experiment. We made a follow-up experiment by removing
the most favorite author from the candidate pool one by one. The queue follows the popularity
(CRH, DG, LY, and FML, in turn).</p>
      <p>The result shows that the next favored author keeps taking the lead by removing the most
preferred. For instance, after removing the most popular author (CRH), the second popular
author (DG) draws almost all the attention of the classifier when seven candidates are present.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>Can we successfully infer the authorship of an unsigned vernacular text from classical texts
with a standard authorship attribution technique? Our finding shows that inferring the author
of a vernacular text using classical text as training data is challenging based on function
characters/words frequency. Increasing training sample size is of limited help.</p>
      <p>We speculate that the difficulty of the cross-register task lies in the elusive degree of
“vernacularization” in classical texts. It is entirely possible that one’s classical style is “more
vernacular” than others because the transition of Chinese from classical to vernacular
unfolded gradually over time. For example, Chu Renhuo (CRH), the author most likely to be
predicted by the model in the cross-register settings, is well-known as a vernacular novelist.
In his classical writing, the Hard Gourd Collection, CRH documents many anecdotes about
the composition of doggerel verse and rhymed verse (Qu) with extensive use of the function
character “ ” (i.e., “ ” and “ ”).12 And “ ” as a modal particle rarely
appears in written classical Chinese, if at all.</p>
      <p>Also, the amount of dialog—another likely source of confusion—also varies across works.
Dialog parts make a classical text resemble a vernacular text. For instance, Du Gang’s the
Romance of the Northern Dynasties contains a large portion of dialog. This factor also relates
to genre. Fiction and diaries tend to have more speech-like prose relative to history and poetry.
The “signal” ordinarily picked up on by function words/characters is no longer detectable given
this variability.</p>
      <p>Future work might experiment with new, purpose-built feature sets which can mitigate the
problem created by dialog and vernacularization. A larger corpus, especially one with limited
dialog in classical texts, would also be valuable.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We thank the anonymous reviewers for their insightful comments on the manuscript of our
12Qu is one of the most colloquial form among rhymed verse forms in Chinese classical literature.
paper.
[16] E. Stamatatos. “Authorship attribution using text distortion”. In: Proceedings of the 15th
Conference of the European Chapter of the Association for Computational Linguistics:
Volume 1, Long Papers. 2017, pp. 1138–1149.
[17] E. Stamatatos. “Masking topic-related information to enhance authorship attribution”.</p>
      <p>In: Journal of the Association for Information Science and technology 69.3 (2018),
pp. 461–473.</p>
    </sec>
    <sec id="sec-7">
      <title>A. Measuring same-register accuracy inflation</title>
      <p>When processing the corpus, we performed a “create-two-from-one” strategy on the classical
works for authors who only have one work. We did that because training and testing with
texts from the same work can inflate the accuracy by overfitting a model to pick up
topical information. Although function words are nominally content-free, genre and content can
influence function word rates. Indeed, this problem is perhaps clearer in Chinese than it is
in authorship attribution work using other languages (where the problem also exists). For
example, (“very”) is a common classical function character, and it is part of the name of
a Buddha, Mañjuśrī ( ). This may increase the likelihood of assigning a novel in which
Mañjuśrī appears to an author who frequently uses as a function character.</p>
      <p>To probe if our strategy prevents accuracy inflation, we ran another experiment only applying
candidates who have at least two works in the classical register and compare it with the main
experiment’s same-register accuracy (already with the “create-two-from-one” strategy used).
Only three authors have at least two classical works. We calculated the same-register accuracy
with diferent feature sets under the same setting corresponding to the limited data scenario
of the main experiment (See Figure 4). By comparing the three-author settings, in one case
strictly using diferent works and in the other case applying “create-two-from-one” strategy, we
estimate that the accuracy inflation is low, less than 3%. Same-register accuracy is far better
than chance.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Baayen</surname>
          </string-name>
          ,
          <string-name>
            <surname>H. van Halteren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Neijt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Tweedie</surname>
          </string-name>
          . “
          <article-title>An experiment in authorship attribution”</article-title>
          .
          <source>In: 6th JADT</source>
          . Vol.
          <volume>1</volume>
          .
          <year>2002</year>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>75</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.-C.</given-names>
            <surname>Chang</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.-J.</given-names>
            <surname>Lin</surname>
          </string-name>
          . “
          <string-name>
            <surname>LIBSVM</surname>
          </string-name>
          :
          <article-title>A library for support vector machines”</article-title>
          .
          <source>In: ACM Transactions on Intelligent Systems and Technology 2 (3</source>
          <year>2011</year>
          ),
          <volume>27</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          :
          <fpage>27</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Chen</surname>
          </string-name>
          . “
          <article-title>Coexistence of Classical Chinese and Vernacular Chinese in Fiction Writing”</article-title>
          .
          <source>In: A Historical Study of Early Modern Chinese Fictions</source>
          <volume>(</volume>
          <fpage>1890</fpage>
          -
          <lpage>1920</lpage>
          ). Springer,
          <year>2021</year>
          , pp.
          <fpage>123</fpage>
          -
          <lpage>145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ge</surname>
          </string-name>
          .
          <article-title>Out of the margins: the rise of Chinese vernacular fiction</article-title>
          . University of Hawaii Press,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Glover</surname>
          </string-name>
          and
          <string-name>
            <surname>G. Hirst.</surname>
          </string-name>
          “
          <article-title>Detecting stylistic inconsistencies in collaborative writing”</article-title>
          .
          <source>In: The new writing environment</source>
          . Springer,
          <year>1996</year>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>168</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Herring</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Paolillo</surname>
          </string-name>
          . “
          <article-title>Gender and genre variation in weblogs”</article-title>
          .
          <source>In: Journal of Sociolinguistics 10.4</source>
          (
          <issue>2006</issue>
          ), pp.
          <fpage>439</fpage>
          -
          <lpage>459</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>W.</given-names>
            <surname>Hu</surname>
          </string-name>
          . Bibliography on Women in Antiquity. Ed. by
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          . 3rd. Shanghai Lexicographical Publishing House,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Koppel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schler</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Bonchek-Dokow.</surname>
          </string-name>
          “Measuring Diferentiability: Unmasking Pseudonymous Authors.”
          <source>In: Journal of Machine Learning Research 8.6</source>
          (
          <issue>2007</issue>
          ), pp.
          <fpage>1261</fpage>
          -
          <lpage>1276</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>I.</given-names>
            <surname>Markov</surname>
          </string-name>
          , E. Stamatatos, and
          <string-name>
            <surname>G. Sidorov.</surname>
          </string-name>
          “
          <article-title>Improving cross-topic authorship attribution: The role of pre-processing”</article-title>
          .
          <source>In: International Conference on Computational Linguistics and Intelligent Text Processing</source>
          . Springer.
          <year>2017</year>
          , pp.
          <fpage>289</fpage>
          -
          <lpage>302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Duchesnay.</surname>
          </string-name>
          “Scikit-learn:
          <source>Machine Learning in Python (Version</source>
          <volume>0</volume>
          .24.
          <article-title>1)”</article-title>
          .
          <source>In: Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          ), pp.
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pu</surname>
          </string-name>
          .
          <article-title>Strange Stories from a Chinese Studio (Volumes 1 and 2</article-title>
          ).
          <source>Kelly &amp; Walsh</source>
          ,
          <year>1880</year>
          . url: https://www.gutenberg.org/files/43629/43629-
          <fpage>0</fpage>
          .txt.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rohatgi</surname>
          </string-name>
          , et al. “
          <article-title>Can pseudonymity really guarantee privacy?”</article-title>
          <source>In: USENIX Security Symposium</source>
          .
          <year>2000</year>
          , pp.
          <fpage>85</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Rubin</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Greene</surname>
          </string-name>
          . “
          <article-title>Gender-typical style in written language”</article-title>
          .
          <source>In: Research in the Teaching of English</source>
          (
          <year>1992</year>
          ), pp.
          <fpage>7</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>U.</given-names>
            <surname>Sapkota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Solorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bethard</surname>
          </string-name>
          . “
          <article-title>Domain adaptation for authorship attribution: Improved structural correspondence learning”</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers).</given-names>
          </string-name>
          <year>2016</year>
          , pp.
          <fpage>2226</fpage>
          -
          <lpage>2235</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>U.</given-names>
            <surname>Sapkota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Solorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bethard</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          . “
          <article-title>Cross-topic authorship attribution: Will out-of-topic data help?”</article-title>
          <source>In: Proceedings of COLING</source>
          <year>2014</year>
          ,
          <source>the 25th International Conference on Computational Linguistics: Technical Papers</source>
          .
          <year>2014</year>
          , pp.
          <fpage>1228</fpage>
          -
          <lpage>1237</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Huang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Classical Chinese Dictionary of Function Characters</article-title>
          . Peking University Press,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Riddell</surname>
          </string-name>
          . “
          <article-title>Cross-Register Authorship Attribution using Vernacular and Classical Chinese Texts”</article-title>
          .
          <source>In: DH Benelux</source>
          <year>2021</year>
          . Zenodo,
          <year>2021</year>
          . doi:
          <volume>10</volume>
          .5281/ zenodo.4886596.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          . History of Chinese Linguistics. Zhonghua Book Company,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <source>Modern Chinese Dictionary of Function Words</source>
          . Shanghai Lexicographical Publishing House,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>B.</given-names>
            <surname>Yu</surname>
          </string-name>
          . “
          <article-title>Function words for Chinese authorship attribution”</article-title>
          .
          <source>In: Proceedings of the NAACL-HLT 2012 Workshop on Computational Linguistics for Literature</source>
          .
          <year>2012</year>
          , pp.
          <fpage>45</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <source>Five Hundred Kinds of Ming and Qing Novels</source>
          . Shanghai Lexicographical Publishing House,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <article-title>“A framework for authorship identification of online messages: Writing-style features and classification techniques”</article-title>
          .
          <source>In: Journal of the American society for information science and technology 57.3</source>
          (
          <issue>2006</issue>
          ), pp.
          <fpage>378</fpage>
          -
          <lpage>393</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>