<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GLAD: Groningen Lightweight Authorship Detection</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Language and Cognition Groningen, University of Groningen</institution>
          ,
          <addr-line>Groningen</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Manuela Hürlimann</institution>
          ,
          <addr-line>Benno Weck, Esther van den Berg, Simon Šuster, and Malvina Nissim</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>We present a simple and effective approach to authorship verification for Dutch, English, Spanish and Greek, which can be easily ported to yet other languages. We train a binary linear classifier both on the features describing known and unknown documents individually, and on the joint features comparing these two types of documents. The list of feature types includes, among others, character n-grams, the lexical overlap, visual text properties and a compression measure. We obtain competitive results that outperform the baseline and position our system among the top PAN shared task participants.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In the authorship verification task as set in the PAN competition1, a system is given a
collection of problem sets containing one or more known documents written by author
Ak and a new, unknown document written by author Au, and is then required to
determine whether Ak = Au without access to a closed set of alternatives. In this form, the
task is generally interpreted as a one-class classification problem [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], in the sense that
the negative class is not homogeneously represented and systems are based on
recognition of a given class rather than discrimination among classes [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This is akin to outlier
or novelty detection [
        <xref ref-type="bibr" rid="ref5 ref7">5,7</xref>
        ] and different from standard authorship attribution problems,
where a system must choose among a set of candidate authors in a more standard,
multi-class text categorisation fashion.
      </p>
      <p>
        Since multi-class classification is a more natural and efficient way of performing
classification, researchers have experimented with the introduction of negative instances
to each problem set by selecting random external documents [18], and with the addition
of positive examples by splitting long known documents into chunks [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This way,
each problem set is turned into a true binary classification task as both positive and
negative instances are represented, the latter being mimed by the external documents.
Approaches using external documents are referred to as extrinsic approaches, while
methods that do not introduce any external documents are called intrinsic approaches [21].
Building on the impostor method used by [18] within the PAN 2013 competition,
which indeed used sets of external documents to produce negative instances, [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] also
      </p>
      <sec id="sec-1-1">
        <title>1 http://pan.webis.de</title>
        <p>use an extrinsic approach and achieve optimal results at PAN 2014. This result is
interesting as most systems participating in PAN 2014 are instead intrinsic in nature.
Additionally, only three out of 13 approaches exploit actual trained models [21], while
the rest pursue a “lazy” strategy.</p>
        <p>In the light of the above discussion, we decided to cast the problem as a binary
classification task where class values are Y (Ak = Au) and N (Ak ≠ Au). In order to
maximise speed and simplicity, we do not introduce any negative examples by means
of external documents, thus adhering to an intrinsic approach. To fit a binary
classification setting, we train a model on the whole dataset, which contains both positive
(Ak = Au) and negative (Ak ≠ Au) problem sets in equal number. In other words,
we treat each problem set as a training vector, exploiting both positive and negative
instances in learning. Each instance is represented as a feature vector which contains
feature values representing the known document, values representing the unknown
document and values comparing the known and unknown document, and is associated with
a class value of either Y(es) or N(o). Furthermore, all features that we include in our
final models are simple and fast to obtain from any text without requiring any
complex processing. An average train or test run of our system on one of the PAN datasets
takes only minutes to complete. Finally, to ensure portability, we do not develop any
language-specific feature. Rather, we tune the system towards a specific language by
means of selecting and combining features in (possibly) different ways.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Approach and Data Representation</title>
      <p>There are two reasons for casting an authorship verification problem as a binary
classification task: (i) because in the PAN 2015 training data the problem sets are equally
distributed for positive and negative evidence and (ii) because of the success of
extrinsic methods, which bear more similarities to our multi-class classification than to
one-class (detection) approaches (see discussion in Section 1 above). However,
theoretically, the problem still resembles one-class classification more than a true binary
classification task, especially because the negative class is not inherently homogeneous
(our approach is more alike to telling apples from other fruit than telling apples from
pears). It also needs to be noted that, in a realistic setting, evidence would not
necessarily be balanced, with negative examples outnumbering positive ones (there are many
more fruits that are not apples than there are apples).</p>
      <p>
        For one-class classification, Support Vector Machines (SVM) have been shown to
perform very well [
        <xref ref-type="bibr" rid="ref13">13,23</xref>
        ], and this is especially true when the data are highly
unbalanced, as demonstrated by [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. They also show that, in presence of a balanced set up
to a ratio of 1:3.5, binary classification outperforms one-class classification. Since we
are dealing with balanced datasets, we use a binary class SVM. In a different setting, a
one-class SVM could be used, without major variations in the algorithm.
      </p>
      <p>
        Our system is implemented using Python’s scikit-learn [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] machine learning
library as well as the Natural Language Toolkit (NLTK) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We used an SVM with
default parameter settings in all final models, with an implementation based on libsvm.
The features we experimented with and the tools used to extract them are described
below.
      </p>
      <p>Building on the existing literature and on observations that came from preliminary
exploration of the training data (both from PAN 2014 and from PAN 2015), we
developed a set of 29 features. Not all of them were used in the final configuration, as
ablation experiments showed little or no contribution of some groups of features.
Nevertheless, we describe them all in this section, and present results from ablation
experiments in the next.</p>
      <p>There are two major ways to divide up the features we experimented with. The
first one has to do with the kind of information they represent, and we clustered them
in seven different groups (see Sections 2.1–2.7). The second one has to do with how
the information is encoded regarding the known and unknown documents. Specifically,
features can describe the known and unknown documents separately, in which case we
talk about individual features, or they can describe them together, in which case we talk
about joint features. For example, when comparing the average sentence length of the
known and the unknown portion of a training instance, one can represent the average
sentence lengths as two individual values or compute the difference of the averages to
get a single joint feature.
2.1</p>
      <sec id="sec-2-1">
        <title>N-gram Features</title>
        <p>
          Information on correspondence of character sequences in known and unknown
documents has been shown to be a successful feature for this task [20,21]. All n-gram
features in our model are joint features. We included n-grams with n ranging from 1 to
5, and calculated n-gram similarity in two different ways. According to [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], an author
profile is defined as the set of the k most frequent n-grams with their normalised
frequencies, as collected from training data. In order to deal with sparseness when n &gt; 2,
the (dis)similarity measure that they use, and that we adopt, takes into account the
difference between the relative frequency of a given n-gram averaged over all
knownunknown document pairs of one problem instance. We call this feature group n-gram
norm. We also use a simple n-gram overlap measure called SPI (simplified profile
intersection [20, p. 548]). This measure is based on the number of common n-grams in
the most frequent k n-grams for each document. We calculate the SPI score separately
for each n between 1 and 5.
        </p>
        <p>Profile size k is another parameter of the n-gram norm and SPI measures. For
our system we fixed k at 100, thus taking into account the 100 most frequent n-grams
of each document.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Token Features</title>
        <p>A common way to measure the similarity of two given texts is to compute the cosine
similarity of their vector representations. We hypothesise that texts written by different
authors are lexically less similar than texts written by the same author. We measure
the similarity in each training instance by averaging over the L2-normalised dot
product of the raw term frequency vectors of a single known document and the unknown
document. The token feature is considered a joint feature.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Sentence Features</title>
        <p>
          We consider the number of tokens per sentence to be a simple yet effective feature for
detecting authorship [
          <xref ref-type="bibr" rid="ref12">12,17</xref>
          ]. The values for average sentence length are obtained and
represented both as joint and individual features. Sentence boundaries are determined
using language-specific models of the NLTK Punkt Tokeniser2.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Entropy Features</title>
        <p>The notion of the entropy of a text was first introduced by Shannon [19]. We hypothesise
that authors have distinct entropy profiles due to the varying lexical and
morphosyntactic patterns they use. We use the average entropy of each known and unknown
documents (as individual features), as well as two joint measures: (i) the average entropy of
each known concatenated with the unknown document, (ii) and the absolute difference
between entropies of known and unknown documents.
2.5</p>
      </sec>
      <sec id="sec-2-5">
        <title>Visual Features</title>
        <p>Upon inspecting the training data, it became apparent that while Dutch documents
contain no newline characters, and Greek and Spanish texts are visually very uniform,
many of the English training instances had been drawn from plays and poems. These
documents are characterised by marked usage of white space and punctuation. For
instance, the sample text in Figure 1 includes apostrophes to reflect the speakers’ accents,
many mid-sentence line breaks and the segmentation of the text into turns of dialogue.
Compared to a piece of prose, or to a play, we predict that this text has a higher
number of apostrophes and newlines and that this information is crucial for distinguishing
same-author from different-author instances.</p>
        <p>Taking into consideration our aim of designing a lightweight system, we opted for
straight-forward measures of layout properties: (i) punctuation, to capture differences
in use of typographical signs, (ii) line endings, to measure preferred ways of closing
lines, (iii) letter case, to capture e.g. the occurrence of names and the capitalisation of
the start of sentences, and finally (iv) line length measures and (v) text block size, to
capture the amount and distribution of blank space in a document.</p>
        <p>In detail, the visual features consist of:
(i) Punctuation:
frequencies of exclamation marks, question marks, semi-colons, colons, commas,
full stops, hyphens and quotation marks
(ii) Line endings:
frequencies of full stops, commas, question marks, exclamation marks, spaces,
hyphens, and semi-colons at the end of a line
(iii) Letter case
– ratio of uppercase characters to lowercase characters
– proportion of uppercase characters</p>
        <sec id="sec-2-5-1">
          <title>2 http://www.nltk.org/api/nltk.tokenize.html#module-nltk.</title>
          <p>tokenize.punkt
Aw, goin ’ t o s c h o o l didn ’ t do me no good . The t e a c h e r s was a l l
down on me . I couldn ’ t l e a r n n o t h i n ’ t h e r e .</p>
          <p>Nor any o t h e r p l a c e , I ’m t h i n k i n ’ , you ’ r e t h a t
t h i c k , Whisht ! I t ’ s
t h e d o c t o r comin ’ down from E i l e e n . What ’ l l he say , I wonder ? Aw, Doctor , ⤦
Ç and how ’ s E i l e e n now? Have you g o t h e r
c u r e d o f t h e weakness ?
Here a r e two p r e s c r i p t i o n s
t h a t ’ l l have t o be f i l l e d i m m e d i a t e l y .</p>
          <p>You t a k e them , B i l l y , and run round t o t h e drug
s t o r e .</p>
          <p>
            Give me t h e money , t h e n .
After obtaining vectors for documents, counts were averaged across all known
documents in an instance, generating a simple author profile to be compared to the
unknown/questioned document. We use the cosine similarity for the property vectors
punctuation, line endings and line length, and simple subtraction for letter case and
text block. All visual features are joint.
Compression features have been successfully used for authorship identification, and can
yield performance similar to n-gram features [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. We used the “Compression
Dissimilarity Measure (CDM)” [12, p.19], which, for two documents x and y is defined as
the sum of the compressed lengths of x and y divided by the compressed length of the
concatenation of the two documents:
          </p>
          <p>CDM (x; y) = C(x) + C(y) : (1)</p>
          <p>C(xy)
Our implementation of CDM normalises the compressed lengths by the number of
characters in each document and uses the zlib algorithm3 for compression. By definition this
is a joint feature.</p>
        </sec>
        <sec id="sec-2-5-2">
          <title>3 http://zlib.net/</title>
        </sec>
      </sec>
      <sec id="sec-2-6">
        <title>2.7 (Morpho)syntactic Features</title>
        <p>
          We also investigate the role of more elaborate feature types, namely part-of-speech
(POS) tags and syntactic functions obtained from dependency trees. Previous attempts
have mostly dealt with shallow information obtained from POS tags (see [20] for an
overview), whereas methods using syntactic features are less common, especially those
using dependency- rather than phrase-based syntax [
          <xref ref-type="bibr" rid="ref16 ref6">6,16</xref>
          ]. There are many possible
ways to include (morpho)syntactic features, however the underlying motivation for
including them is typically that the frequency and the patterns of (morpho)syntactic
categories are not under conscious control by authors. We run these experiments for English
only.
        </p>
        <p>
          We use the MST dependency parser [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] for English trained on sections 2–21 of
the Penn Treebank WSJ (PTB) prepared using the standard procedures of [22] and [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
The parser achieves a reasonable labelled accuracy of around 0.85 on the testing part of
the PTB. POS tags are obtained with the Citar tagger4, which achieves an accuracy of
around 96% on the PTB test sections. We take a simple approach of comparing POS and
dependency label distributions of known and unknown documents, and use two
measures for comparing the distributions: cosine similarity and entropy. For the former, we
calculate the L2-normalised cosine similarity between two frequency distributions,
producing a joint feature. For entropy, we calculate the Shannon entropy (see Section 2.4)
of the known and unknown texts separately (averaged in case of multiple known
documents), yielding two individual features. We also use the absolute difference between
the two as a joint feature. Entropy features are only used for POS tags.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Feature Selection</title>
      <p>
        In order to assess the contribution of the feature types discussed in Section 2, we ran
a series of feature ablation and single-feature experiments, where "feature" actually
stands for each of the groups defined in Sections 2.1–2.7. To these, we added four more
categories which simply group selected feature types:
– all individual features
– all joint features
– visual and compression features (vis+comp)
– visual features, n-gram features and token (cosine) feature (vis+n+tok)
All tests for evaluating the contribution of the various features were conducted in Weka
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], using the LibSVM classifier, which has the same default parameters as the
scikitlearn implementation (cf. Section 2). We ran the experiments on the PAN 2015
training data, which consist of 100 problem sets per language, with a balanced distribution
of positive and negative instances. In Figures 2 and 3 we report results using 10-fold
stratified cross-validation, averaged over five runs.
      </p>
      <sec id="sec-3-1">
        <title>4 http://github.com/danieldk/citar</title>
        <p>Figure 2 reports on our ablation study, in which we leave out a single feature group at a
time. If removing a feature group causes lower scores, it is beneficial, while removing
harmful feature groups increases performance. We can see that for Dutch, removing the
individual features leads to the greatest drop in performance. These are therefore the
most useful features. The joint features, on the other hand, appear to be harmful to the
performance of the system. Most other feature groups do not greatly affect the results.</p>
        <p>For English, the most useful combination includes visual features, n-gram features,
and the cosine feature. Joint features also perform well, while individual features are
harmful. Part-of-speech features also appear to be harmful. We experimented with
(morpho)syntactic features in English only. The ablation results indicate that these features
are not effective, thus, we do not pursue these further for the other three languages. A
possible explanation is that the quality of the dependency parser is not sufficient,
however this is less likely to be the case for the POS tagger. We leave the design of more
complex (morpho)syntactic feature types for future work.</p>
        <p>Greek and Spanish both experience the greatest drop in performance when joint
features are removed, further underlining the usefulness of this feature group. It is
interesting that only Dutch shows a preference for individual over joint features, though
we could not identify a specific reason as to why this is the case, and it will require
further investigation.
3.2</p>
        <sec id="sec-3-1-1">
          <title>Single-Feature Results</title>
          <p>Figure 3 shows the effects of using single groups of features only. Consistent with the
results in Figure 2, individual features are the most useful combination for Dutch, while
sentence and compression features score conspicuously low.</p>
          <p>Visual features score highly for English (with this single feature group
outperforming the full feature set), again underlining the importance of capturing layout properties
for this data set.</p>
          <p>For Greek, we find character n-grams, visual features and token features to be most
beneficial. This cannot be attributed to any of the single component groups but seems
to arise out of the specific combination and interaction of these feature groups.</p>
          <p>Finally, scores on Spanish are maximised by the full feature set, which is marginally
better than only the joint features. The combination of all available information
therefore has a positive effect on Spanish, while the other languages benefit from restricting
analysis to a subset of features.
3.3</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Final Feature Sets</title>
          <p>Based on the observations gleaned from the above tests, we grouped the features into
four combinations with which we experimented in order to define the final feature set
for each language. Table 1 presents the cross-validated results for the combinations
across the four languages.</p>
          <p>– combo1: n-grams, visual
– combo2: n-grams, visual, token
– combo3: n-grams, visual, token, entropy (joint)
– combo4: full set (excluding morpho(syntactic) features)
We submitted runs for all four languages. The models were trained on the corresponding
four training sets of PAN 2015. Based on the experiments reported in Section 3, we
ran our system with the full feature set for English, Dutch, and Spanish (combo4),
but with a different configuration for Greek. Indeed, we observed consistent gains in
the overall score for this language when using a subset of features (combo2), namely
n-grams, visual, and token features. For English, Dutch and Spanish, only marginal
gains were observed using different combinations of features; therefore, we did not use
a reduced feature set for these languages.
We report the combined AUC-c@1 scores for all languages in Table 2. On a balanced
test set, these results outperform a baseline system which assigns Y (or N)
throughout and thus achieves a combined score of 0.25 (c@1=0.5 and AUC=0.5). The
crossvalidation results on the training data show similar performance for all systems except
for Spanish, where the score is much higher. This can be explained by the fact that
some instances (documents) were repeated in the Spanish training set. On the test set,
we obtain a combined score of around 0.6 for both Dutch and Greek, with which we
achieve the third- and fourth-best result out of all PAN 2015 participants. For
English, the score on the test set (0.41) is lower than for the other languages, and contrasts
somewhat with the cross-validation results. A possible explanation is that there is a
considerable domain shift between the training and test sets for English.</p>
          <p>Our method also achieves fast testing times – with some variance per language –
with an average runtime of one minute.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>We have presented a simple, yet effective approach to the authorship verification task
of the PAN competition. The task was to determine whether Ak = Au for one or more
known documents written by author Ak and a “questioned” document written by author
Au. Our solution to this problem was to treat it as a binary classification task, training a
model across the whole dataset. Based on the prediction that texts written by different
authors are less similar than texts written by the same author due to author-specific
patterns in writing not under the conscious control of the author, we developed an initial
set of 29 features to model similarity or lack thereof.</p>
      <p>Ablation and single-feature experiments were used to assess the contribution of the
various features we originally selected, and led to the configuration of the final model
for each language. Indeed, to ensure portability, we did not develop any
languagespecific features (apart from NLTK language-specific sentence-splitters), and instead
tuned the system by feature selection, as we noticed some differences in performance
between features combinations when applied to different languages. One interesting
observation was that only Dutch showed a preference for individual over joint features,
but no aspect of the Dutch training data could be found which might cause this result.
This raises the question for further research of whether there are theoretical grounds
for preferring joint similarity measures over separate, individual measures for author
verification tasks. Also, results for Greek showed that a smaller set of specific features
was consistently outperforming the full set, but further investigation is required to
understand exactly why.</p>
      <p>With parsimony and speed in mind, all selected features in our final models were
straight-forward to implement and easy to obtain for any text. Indeed, our resulting
system is easy to run, adaptable to new languages through feature selection, and fast,
with runs taking one minute on average.</p>
      <p>The PAN 2015 test results position our system in the third- and fourth-best place
out of all participants for the Dutch and Greek datasets. We obtain a somewhat lower
score for English, which contrasts with promising cross-validation results. While
inspection of the actual test data, once available, might shed light on the causes of such a
drop, this result hints at the variability which is inherent to binary classification when
applied to moderate-size training and test sets. Further exploration of features,
feature parameters and feature combinations for individual languages is left as a
challenging avenue for future work. We make our system GLAD publicly available at
https://github.com/.</p>
      <p>Acknowledgements The first three authors were supported by the Erasmus Mundus
Master’s Program in Language and Communication Technologies (EM LCT). We would
also like to thank the anonymous reviewers for their helpful comments.
17. Sapkota, U., Solorio, T., Montes-y GÃs¸mez, M., Rosso, P.: The use of orthogonal similarity
relations in the prediction of authorship. In: Gelbukh, A. (ed.) Computational Linguistics
and Intelligent Text Processing, Lecture Notes in Computer Science, vol. 7817, pp.
463–475. Springer Berlin Heidelberg (2013)
18. Seidman, S.: Authorship verification using the impostors method – notebook for pan at clef
2013. In: Forner, P., Navigli, R., Tufis, D. (eds.) CLEF 2013 Evaluation Labs and Workshop
– Working Notes Papers (2013)
19. Shannon, C.E.: A mathematical theory of communication. ACM SIGMOBILE Mobile</p>
      <p>Computing and Communications Review 5(1), 3–55 (2001)
20. Stamatatos, E.: A survey of modern authorship attribution methods. J. Am. Soc. Inf. Sci.</p>
      <p>
        Technol. 60(3), 538–556 (2009)
21. Stamatatos, E., Daelemans, W., Verhoeven, B., Stein, B., Potthast, M., Juola, P.,
Sánchez-Pérez, M.A., Barrón-Cedeño, A.: Overview of the author identification task at
PAN 2014. In: Cappellato et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], pp. 877–897, http:
//ceur-ws.org/Vol-1180/CLEF2014wn-Pan-StamatosEt2014.pdf
22. Vadas, D., Curran, J.R.: Adding Noun Phrase Structure to the Penn Treebank. In: ACL
(2007)
23. Yu, H.: Svmc: Single-class classification with support vector machines. In: Proceedings of
the 18th International Joint Conference on Artificial Intelligence. pp. 567–572. IJCAI’03,
Morgan Kaufmann Publishers Inc., San Francisco, CA, USA (2003)
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bellinger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Japkowicz</surname>
          </string-name>
          , N.:
          <article-title>One-class versus binary classification: Which and when?</article-title>
          <source>In: Proceedings of the 2012 11th International Conference on Machine Learning and Applications -</source>
          Volume
          <volume>02</volume>
          . pp.
          <fpage>102</fpage>
          -
          <lpage>106</lpage>
          . ICMLA '12, IEEE Computer Society, Washington, DC, USA (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loper</surname>
          </string-name>
          , E.:
          <article-title>Natural language processing with Python. "</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          ,
          <source>Inc."</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halvey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kraaij</surname>
          </string-name>
          , W. (eds.): Working Notes for CLEF 2014 Conference, Sheffield, UK,
          <source>September 15-18</source>
          ,
          <year>2014</year>
          , CEUR Workshop Proceedings, vol.
          <volume>1180</volume>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-1180
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>The weka data mining software: An update</article-title>
          .
          <source>SIGKDD Explorations</source>
          Volume
          <volume>11</volume>
          (
          <issue>1</issue>
          ) (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hodge</surname>
            ,
            <given-names>V.J.</given-names>
          </string-name>
          , Austin,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>A survey of outlier detection methodologies</article-title>
          .
          <source>Artificial Intelligence Review</source>
          <volume>22</volume>
          (
          <issue>2</issue>
          ),
          <fpage>85</fpage>
          -
          <lpage>126</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hollingsworth</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Using dependency-based annotations for authorship identification</article-title>
          . In: Sojka,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Horák</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Kopecek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Pala</surname>
          </string-name>
          ,
          <string-name>
            <surname>K</surname>
          </string-name>
          . (eds.)
          <source>TSD. Lecture Notes in Computer Science</source>
          , vol.
          <volume>7499</volume>
          , pp.
          <fpage>314</fpage>
          -
          <lpage>319</lpage>
          . Springer (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Japkowicz</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Myers</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gluck</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>A novelty detection approach to classification</article-title>
          . In: IJCAI. pp.
          <fpage>518</fpage>
          -
          <lpage>523</lpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Johansson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nugues</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Extended constituent-to-dependency conversion for English</article-title>
          .
          <source>In: NODALIDA</source>
          . pp.
          <fpage>105</fpage>
          -
          <lpage>112</lpage>
          . Tartu,
          <string-name>
            <surname>Estonia</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Keselj</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cercone</surname>
          </string-name>
          , N.,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>N-gram-based author profiles for authorship attribution</article-title>
          .
          <source>In: Proceedings of the Conference of the Pacific Association for Computational Linguistics, PACLING</source>
          . vol.
          <volume>3</volume>
          , pp.
          <fpage>255</fpage>
          -
          <lpage>264</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Khonji</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iraqi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>A slightly-modified gi-based author-verifier with lots of features (ASGALF)</article-title>
          .
          <source>In: Cappellato et al. [3]</source>
          , pp.
          <fpage>977</fpage>
          -
          <lpage>983</lpage>
          , http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1180</volume>
          /
          <fpage>CLEF2014wn</fpage>
          -Pan-KonijEt2014.pdf
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
          </string-name>
          , J.:
          <article-title>Authorship verification as a one-class classification problem</article-title>
          .
          <source>In: Proceedings of the Twenty-first International Conference on Machine Learning</source>
          . pp.
          <fpage>62</fpage>
          -.
          <source>ICML '04</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2004</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1015330.1015448
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>An Exploratory Study on Authorship Verification Models for Forensic Purpose</article-title>
          .
          <source>Master's thesis</source>
          , Delft University of Technology (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Manevitz</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yousef</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>One-class SVMs for document classification</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>2</volume>
          ,
          <fpage>139</fpage>
          -
          <lpage>154</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>McDonald</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Online learning of approximate dependency parsing algorithms</article-title>
          . In: EACL (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cournapeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duchesnay</surname>
          </string-name>
          , E.:
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          ,
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rygl</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zemková</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovár</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Authorship verification based on syntax features</article-title>
          .
          <source>In: RASLAN</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>