<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Investigating the Application of Distributional Semantics to Stylometry</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giulia Benotto</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emiliano Giovannetti</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pisa - Italy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>name.surnameg@ilc.cnr.it</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>English. The inclusion of semantic features in the stylometric analysis of literary texts appears to be poorly investigated. In this work, we experiment with the application of Distributional Semantics to a corpus of Italian literature to test if words distribution can convey stylistic cues. To verify our hypothesis, we have set up an Authorship Attribution experiment. Indeed, the results we have obtained suggest that the style of an author can reveal itself through words distribution too.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Stylometry, that is the application of the study of
linguistic style, offers a means of capturing the
elusive character of an author’s style by
quantifying some of its features. The basic stylometric
assumption is that each writer has certain
stylistic idiosyncrasies (a “human stylome”
        <xref ref-type="bibr" rid="ref23">(Van
Halteren et al., 2005)</xref>
        ) that define their style.
Analysis based on stylometry are often used for
Authorship Attribution (AA) tasks, since the main idea
behind computationally supported AA is that by
measuring some textual features, we can
distinguish between texts written by different authors
        <xref ref-type="bibr" rid="ref20">(Stamatatos, 2009)</xref>
        .
      </p>
      <p>One of the less investigated stylistic feature is
the way in which authors use words from a
semantic point of view, e.g. if they tend to use more,
when dealing with polysemous words, a certain
sense over the others, or senses that differ (even
slightly) from the one that’s more commonly used
(as it happens, typically, in poetry).</p>
      <p>
        A possible approach to the analysis of this
characteristic is to consider the textual contexts in
which certain words appear. According to
Distributional Semantics (DS), certain aspects of the
meaning of lexical expressions depend on the
distributional properties of such expressions, or
better, on the contexts in which they are observed
        <xref ref-type="bibr" rid="ref13 ref15">(Lenci, 2008; Miller and Charles, 1991)</xref>
        . The
semantic properties of a word can then be defined by
inspecting a significant number of linguistic
contexts, representative of the distributional behavior
of such word.
      </p>
      <p>In this work we would like to investigate if the
analysis of the distribution of words in a text can
be exploited to provide a stylistic cue. In order to
inspect that, we have experimented with the
application of DS to the stylometric analysis of
literary texts belonging to a corpus constituted by texts
pertaining to the work of six Italian writers of the
late nineteenth century.</p>
      <p>In the following, Section 2 gives a short
insight on the state of the art of computational
stylistic analysis, Section 3 describes the approach
together with the corpus used to conduct our
investigation and Section 4 discuss about results.
Finally, Section 5 draws some conclusions and
outlines some possible future works.
2</p>
    </sec>
    <sec id="sec-2">
      <title>State of the Art</title>
      <p>
        The very first attempts to analyze the style of an
author were based on simple lexical features such
as sentence length counts and word length counts,
since they can be applied to any language and
any corpus with no additional requirements
        <xref ref-type="bibr" rid="ref1 ref11 ref19 ref22 ref24">(Koppel and Schler, 2004; Stamatatos, 2006; Zhao and
Zobel, 2005; Argamon et al., 2007)</xref>
        . Similarly,
character measures have been proven to be quite
useful to quantify the writing style
        <xref ref-type="bibr" rid="ref14 ref25 ref4 ref7">(Grieve, 2007;
De Vel et al., 2001; Zheng et al., 2006)</xref>
        .
Basically, a text can be viewed as a mere sequence of
characters, so that various measures can be defined
(including alphabetic, digit, uppercase and
lowercase characters count, etc.). A more elaborate
text representation method is to employ
syntactic information
        <xref ref-type="bibr" rid="ref10 ref17 ref18 ref22 ref24 ref6">(Gamon, 2004; Stamatatos et al.,
2000; Stamatatos et al., 2001; Hirst and Feiguina,
2007; Uzuner and Katz, 2005)</xref>
        . The idea is that
authors tend to use similar syntactic patterns
unconsciously. Therefore, syntactic information is
considered a more reliable authorial fingerprint in
comparison to lexical information.
      </p>
      <p>
        More complicated tasks such as full syntactic
parsing, semantic analysis, or pragmatic
analysis cannot yet be handled adequately by current
NLP technologies for unrestricted text. As a
result, very few attempts have been made to exploit
high-level features for stylometric purposes.
Perhaps the most important method of exploiting
semantic information so far was described in
        <xref ref-type="bibr" rid="ref1">(Argamon et al., 2007)</xref>
        . This work was based on the
theory of Systemic Functional Grammar (SFG)
        <xref ref-type="bibr" rid="ref8">(Halliday, 1994)</xref>
        and consisted on the definition of a set
of functional features that associate certain words
or phrases with semantic information.
      </p>
      <p>
        The previously described features are
application independent since they can be extracted from
any textual data. Beyond that, one can define
application-specific measures in order to better
represent the nuances of style in a given text
domain (such as e-mail messages, or online forum
messages)
        <xref ref-type="bibr" rid="ref14 ref21 ref25">(Li et al., 2006; Teng et al., 2004)</xref>
        .
      </p>
      <p>
        To the best of our knowledge, the application of
DS to the analysis of literary texts has been
documented in a rather small number of works
        <xref ref-type="bibr" rid="ref3 ref9">(Buitelaar et al., 2014; Herbelot, 2015)</xref>
        . In both these
works, DS is used as a theoretical basis in order
to verify some hypotheses on specific semantic
characteristics of poetic works. In more details,
in
        <xref ref-type="bibr" rid="ref3">(Buitelaar et al., 2014)</xref>
        the authors investigated
through DS the influence of Lord Byron’s work
on Thomas Moore trying to find a shared
vocabulary or specific formal textual characteristics. In
        <xref ref-type="bibr" rid="ref9">(Herbelot, 2015)</xref>
        it is argued how distributionalism
can support the notion that the meaning of poetry
comes from the meaning of ordinary language and
how distributional representations can model the
link between ordinary and poetic language.
However, the role of DS in the study of a style of an
author was not the aim of these works.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experimental Setup</title>
      <p>First, we want to specify that it is not our purpose
to propose new ways to improve state-of-the-art
AA algorithms. Indeed, our aim is just to verify
the hypothesis that the distribution of words can
provide an indication of a distributional stylistic
fingerprint of an author. To do this, we have set up
a simple classification task. Subsection 3.1 briefly
depicts the data set we used, while Section 3.2
describes the steps implemented in our experiment.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Set Construction</title>
        <p>In order to build the reference and test corpora, we
started from texts pertaining to the work of six
Italian writers working at the turn of the 20th century,
namely, Luigi Capuana, Federico De Roberto,
Luigi Pirandello, Italo Svevo, Federigo Tozzi and
Giovanni Verga. We chose contiguous authors in
chronological sense, whose texts are available in
digital format (in fact we could not do a similar
survey on the narrative of the 90s because it is still
under copyrights). Indeed, we used texts freely
available for download from the digital library of
the Manunzio project, via the LiberLiber website1.
Since they were encoded in various formats, such
as .epub, .odt and .txt, our pre-processing
consisted in converting them all in .txt format and
getting rid of all xml tags, together with footnotes and
editors’ notes and comments.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experiment Description</title>
        <p>According to Rudman (1997), a striking problem
in stylometry is due to the lack of homogeneity
of the examined corpora, in particular to the
improper selection or fragmentation of the texts, that
might cause alterations in the writers’ style. In
order to create balanced reference corpora, i.e.
covering all the authors’ different stylistic and
thematic phases, for each author, as shown in
Figure 1, we built a reference corpus as the
composition of the 70% of each single work (usually a
novel). The same technique was used to create the
1http://www.liberliber.it/
test corpus by using the remaining 30% of each
work. Typical AA approaches consist in analyzing
known authors and assigning authorship to
previously unseen text on the basis of various features.
Train and test sets should then contain different
texts. Contrary to the classical AA task, our train
and test sets contain different parts of the same
texts. Indeed, with this experiment, we wanted to
understand if the semantics that an author bestows
to a word, is peculiar to his writing. To prove this,
we wanted to cover all the different stylistic and
thematic phases an author can go through during
his activity, hence the partition of all his texts in a
reference and a test portion.</p>
        <p>
          We then analyzed each reference and test
corpora with a Part-of-Speech (PoS) tagger and a
lemmatizer for Italian
          <xref ref-type="bibr" rid="ref5">(Dell’Orletta et al., 2014)</xref>
          . For
every author, we built two lists of word pairs (with
their lemma and PoS), one relative to the tagged
reference corpus (reference pairs) and the other to
the tagged test set (test pairs), where each word
was paired with all the other words with the same
PoS. We also filtered the pairs to leave only nouns,
adjectives and verbs. Starting from the tagged
corpora, we built two words-by-words matrixes2 of
co-occurrence counts (co-occurrence matrixes) for
each author, using a context window of 43. The
chosen DS model
          <xref ref-type="bibr" rid="ref2">(Baroni and Lenci, 2010)</xref>
          was
applied to each matrix to calculate the cosine
be2Being the corpus relatively small and not having
particular computability issues, we chose not to apply
decomposition techniques to reduce the size of the matrixes (and thus
not losing any information).
        </p>
        <p>3We performed different empiric setup of the window’s
size and chose the one that showed more suitable results,
according to what is stated by Kruszewski and Baroni (2014).
tween the vectors representing the two words of
each pair. This allowed us to evaluate the
semantic relatedness between the words by assessing
their proximity in the distributional space as
represented by the cosine value: the more this value
tends to 1, the more the two words of the pair are
considered to be related. We then obtained two
related word pair (RWP) lists for each author A:
RWPrefA and RWPtestA. Figure 1 shows the
process described above.</p>
        <p>Since we wanted to focus on the analysis of the
semantic distribution of words, we decided to
exclude any possible “lexical bias”. For this reason,
we restricted the analysis on a common
vocabulary, i.e. a vocabulary constituted by the
intersection of the six authors’ vocabularies. In this
way, we prevent our classifier to exploit, as a
feature, the presence of words used by some (but not
all) of the authors. Moreover, we removed from
the RWPtest lists all those pairs of words occurring
frequently together in the same context, since they
might constitute a multiword expression that, once
again, could be pertaining with the signature
lexicon of each author. To remove them, we
computed the number of times (#co-occ in Table 1)
they appeared together in the context window, as
well as their total number of occurrences (#occa
and #occb) and we excluded from the analysis
those pairs for which the ratio between the
number of co-occurrences and the total occurrences of
the less frequent word was higher than the
empirically set threshold of 0.5. The first two pairs of
Table 1 would be removed as probable multiword
(PM column in Table 1): “scoppio” (burst) and
“risa” (laughter) could mostly co-occur in
“scoppio di risa” (meaning “burst of laughter”) and the
words “man” and “mano” (both meaning “hand”)
could mostly co-occur in “man mano” (meaning
“little by little”, or “progressively”).</p>
        <p>Wa Wb
scoppio–s risa–s
man–n mano–n
nausea–n disgusto–n
piccolo–a grande–a
#occa #occb #co-occ ratio PM
19 9 7 0.78 yes
50 1325 47 0.94 yes
27 26 0 0 no
248 237 14 0.06 no</p>
        <p>Finally, we reduced the size of the six RWPref
and RWPtest lists by sorting them in decreasing
order of the cosine value and then by keeping the
pairs with the highest cosine, selected using a
percentage parameter as threshold4. We chose to
introduce the parameter for two reasons: i) to
avoid the classification algorithm to be disturbed
by noisy (i.e. not significative) pairs which would
not hold any relevant stylistic cue, and ii) to ease
a literary scholar in the interpretation of the
results by having to analyze just a limited selection
of (potentially) semantically related word pairs.</p>
        <p>For the last phase of our experiment we defined
a classification algorithm to test the effective
presence of stylistic cues inside the obtained RWPtest
lists. We defined a classifier using a nearest-cosine
method to attribute each test list to an author.
The method consisted in searching for a pair of
words contained in the test list inside each
reference list and incrementing by 1 the score of the
author whose reference list included the pair with
the more similar cosine value (i.e. having the
minimum difference): the chosen author was the one
with the highest score. Table 2 shows the
classification results for = 5%.</p>
        <p>Capuana</p>
        <p>Pirandello Svevo Tozzi Verga
De Roberto De Roberto De Roberto De Roberto
De Roberto
Pirandello</p>
        <p>Pirandello</p>
        <p>Pirandello</p>
        <p>Pirandello</p>
        <p>Pirandello
Tozzi/Verga Tozzi
Svevo
Verga</p>
        <p>Svevo
Verga
As summarized in Table 3, a correct classification
of all RWPs in RWPtest lists has been obtained
with a value of 5%.</p>
        <p>To help in interpreting the failure of the
algorithm in classifying Tozzi’s test list for values
lower than 5% (as shown in Table 3) we calculated
the cardinality of the RWPtest lists for each author
with the change in value (Tables 4).</p>
        <p>It is possible to observe how the choice of
influences the correct classification of Tozzi’s test
list. Indeed, the use of a value below 5% has
the effect of remarkably reducing an already small
4At the following url we have uploaded an archive
containing all the data we have used and processed for our
experiment: https://goo.gl/nrTqWh
Svevo
Verga</p>
        <p>Verga
#RWPtestCapuana
#RWPtestDe Roberto
#RWPtestPirandello
#RWPtestSvevo
#RWPtestTozzi
#RWPtestVerga</p>
        <p>Svevo
Verga
Verga
488
425
246
test list (RWPtextTozzi) as shown in Table 4. It is
apparent that increasing the value of and
consequently the number of significant RW pairs that
are analysed, the system is able to correctly
classify RWPtestTozzi (see the values in Tozzi’s row of
Table 3).
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Next Steps</title>
      <p>In this paper we investigated the possibility that
an analysis of the semantic distribution of words
in a text can be potentially exploited to get cues
about the style of an author. In order to
validate our hypothesis, we conducted a first
experiment on six different Italian authors. The results
seem to suggest that the way words are distributed
across a text, can provide a valid stylistic cue to
distinguish an author’s work. Of course, it is not
our intent, with this paper, to define new methods
for enhancing state-of-the-art authorship
attribution algorithms. Our research will focus, in the
next steps, in detecting and providing useful
indications about the style of an author. This can be
done by highlighting, for example, atypical
distributions of words (e.g. with contrastive
methods) or by analysing their distributional variability.
Furthermore, it could be interesting to use a
different distributional measure, than the cosine, to test
our hypothesis.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Argamon</surname>
          </string-name>
          , Casey Whitelaw, Paul Chase, Sobhan Raj Hota, Navendu Garg, and
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Levitan</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Stylistic text classification using functional lexical features</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          ,
          <volume>58</volume>
          (
          <issue>6</issue>
          ):
          <fpage>802</fpage>
          -
          <lpage>822</lpage>
          , April.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Distributional memory: A general framework for corpus-based semantics</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>36</volume>
          (
          <issue>4</issue>
          ):
          <fpage>673</fpage>
          -
          <lpage>721</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Paul</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          , Nitish Aggarwal, and
          <string-name>
            <given-names>Justin</given-names>
            <surname>Tonra</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Using distributional semantics to trace influence and imitation in romantic orientalist poetry</article-title>
          . In AHA!-Workshop 2014 on Information Discovery in Text. ACL.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Olivier De Vel</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alison Anderson</surname>
            ,
            <given-names>Malcolm</given-names>
          </string-name>
          <string-name>
            <surname>Corney</surname>
            , and
            <given-names>George</given-names>
          </string-name>
          <string-name>
            <surname>Mohay</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Mining e-mail content for author identification forensics</article-title>
          .
          <source>ACM Sigmod Record</source>
          ,
          <volume>30</volume>
          (
          <issue>4</issue>
          ):
          <fpage>55</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Felice</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          , Giulia Venturi, Andrea Cimino, and
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>T2kˆ 2: a system for automatically extracting and organizing knowledge from texts</article-title>
          .
          <source>In LREC</source>
          , pages
          <fpage>2062</fpage>
          -
          <lpage>2070</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Gamon</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Linguistic correlates of style: authorship classification with deep linguistic analysis features</article-title>
          .
          <source>In Proceedings of the 20th international conference on Computational Linguistics</source>
          , page 611.
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Jack</given-names>
            <surname>Grieve</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Quantitative Authorship Attribution: An Evaluation of Techniques</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <fpage>251</fpage>
          -
          <lpage>270</lpage>
          , May.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>Michael AK Halliday</source>
          .
          <year>1994</year>
          .
          <article-title>Functional grammar</article-title>
          . London: Edward Arnold.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>Aure´lie Herbelot</source>
          .
          <year>2015</year>
          .
          <article-title>The semantics of poetry: A distributional reading</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          ,
          <volume>30</volume>
          (
          <issue>4</issue>
          ):
          <fpage>516</fpage>
          -
          <lpage>531</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Graeme</given-names>
            <surname>Hirst</surname>
          </string-name>
          and
          <string-name>
            <given-names>Olga</given-names>
            <surname>Feiguina</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Bigrams of Syntactic Labels for Authorship Discrimination of Short Texts</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          ,
          <volume>22</volume>
          (
          <issue>4</issue>
          ):
          <fpage>405</fpage>
          -
          <lpage>417</lpage>
          , September.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Moshe</given-names>
            <surname>Koppel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Schler</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Authorship verification as a one-class classification problem</article-title>
          .
          <source>In Proceedings of the twenty-first international conference on Machine learning, page 62</source>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>Germa´n Kruszewski and Marco Baroni</source>
          .
          <year>2014</year>
          .
          <article-title>Dead parrots make bad pets: Exploring modifier effects in noun phrases</article-title>
          .
          <source>Lexical and Computational Semantics (* SEM</source>
          <year>2014</year>
          ), page 171.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Distributional semantics in linguistic and cognitive research</article-title>
          .
          <source>Italian journal of linguistics</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Jiexun</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Rong</given-names>
            <surname>Zheng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Hsinchun</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>From fingerprint to writeprint</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>49</volume>
          (
          <issue>4</issue>
          ):
          <fpage>76</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>George A Miller and Walter G Charles</surname>
          </string-name>
          .
          <year>1991</year>
          .
          <article-title>Contextual correlates of semantic similarity</article-title>
          .
          <source>Language and cognitive processes</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Rudman</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>The state of authorship attribution studies: Some problems and solutions</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>31</volume>
          (
          <issue>4</issue>
          ):
          <fpage>351</fpage>
          -
          <lpage>365</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          , Nikos Fakotakis, and
          <string-name>
            <given-names>George</given-names>
            <surname>Kokkinakis</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Automatic text categorization in terms of genre and author</article-title>
          .
          <source>Computational linguistics</source>
          ,
          <volume>26</volume>
          (
          <issue>4</issue>
          ):
          <fpage>471</fpage>
          -
          <lpage>495</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          , Nikos Fakotakis, and
          <string-name>
            <given-names>Georgios</given-names>
            <surname>Kokkinakis</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Computer-based authorship attribution without lexical measures</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>35</volume>
          (
          <issue>2</issue>
          ):
          <fpage>193</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Authorship attribution based on feature set subspacing ensembles</article-title>
          .
          <source>International Journal on Artificial Intelligence Tools</source>
          ,
          <volume>15</volume>
          (
          <issue>05</issue>
          ):
          <fpage>823</fpage>
          -
          <lpage>838</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>A survey of modern authorship attribution methods</article-title>
          .
          <source>J. Am. Soc. Inf. Sci. Technol</source>
          .,
          <volume>60</volume>
          (
          <issue>3</issue>
          ):
          <fpage>538</fpage>
          -
          <lpage>556</lpage>
          ,
          <year>March</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Gui-Fa</surname>
            <given-names>Teng</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mao-Sheng</surname>
            <given-names>Lai</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jian-Bin Ma</surname>
            , and
            <given-names>Ying</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>E-mail authorship mining based on svm for computer forensic</article-title>
          .
          <source>In Machine Learning and Cybernetics</source>
          ,
          <year>2004</year>
          . Proceedings of 2004 International Conference on, volume
          <volume>2</volume>
          , pages
          <fpage>1204</fpage>
          -
          <lpage>1207</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>O¨ zlem Uzuner and Boris Katz</source>
          .
          <year>2005</year>
          .
          <article-title>A comparative study of language models for book and author recognition</article-title>
          .
          <source>In Natural Language Processing-IJCNLP</source>
          <year>2005</year>
          , pages
          <fpage>969</fpage>
          -
          <lpage>980</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Hans Van Halteren</surname>
          </string-name>
          ,
          <string-name>
            <surname>Harald Baayen</surname>
            , Fiona Tweedie, Marco Haverkort, and
            <given-names>Anneke</given-names>
          </string-name>
          <string-name>
            <surname>Neijt</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>New machine learning methods demonstrate the existence of a human stylome</article-title>
          .
          <source>Journal of Quantitative Linguistics</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):
          <fpage>65</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Ying</given-names>
            <surname>Zhao</surname>
          </string-name>
          and
          <string-name>
            <given-names>Justin</given-names>
            <surname>Zobel</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Effective and scalable authorship attribution using function words</article-title>
          .
          <source>In Information Retrieval Technology</source>
          , pages
          <fpage>174</fpage>
          -
          <lpage>189</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Rong</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jiexun</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Hsinchun</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Zan</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>A framework for authorship identification of online messages: Writing-style features and classification techniques</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          ,
          <volume>57</volume>
          (
          <issue>3</issue>
          ):
          <fpage>378</fpage>
          -
          <lpage>393</lpage>
          , February.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>