<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Machine Learning-based Intrinsic Method for Cross-topic and Cross-genre Authorship Verification</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Sheffield Regent Court</institution>
          ,
          <addr-line>211 Portobello Sheffield S1 4DP</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>This paper presents our approach for the Author Identification task in the PAN CLEF Challenge 2015. We identified the challenges of this year's are the limited amount of training data and the problems in the sub-corpora are independent in terms of topic and genre. We adopted a machine learning based intrinsic method to verify whether a pair of documents have been written by same or different authors. Several content-independent features, such as function words and stylometric features, were used to capture the difference between documents. Evaluation results on the test corpora show our approach works best on the Spanish data set with 0.7238 and 0.67 for the AUC and C@1 scores respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>Machine Learning</kwd>
        <kwd>Intrinsic Method</kwd>
        <kwd>Authorship Verification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Given a pair of documents (X,Y), the task of author verification is to identify whether
the documents have been written by same or different authors. Compared to authorship
attribution, the authorship verification task is significantly more difficult. Verification
does not learn about the specific character of each author, but rather about the
differences between a pair of documents. The problem is complicated by the fact that an
author may consciously or unconsciously vary his/her writing style from text to text [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>This year’s PAN lab Author Identification task focuses on cross-genre and
crosstopic authorship verification. In this case, the genre and/or topic may differ significantly
between the known and unknown documents. This task is more representative of real
world applications where we could not control the genre/topic of the documents.</p>
      <p>
        The PAN Author idenfitication task is defined as follows: “Given a small set (no
more than 5, possibly as few as one) of known documents by a single person and a
questioned document, the task is to determine whether the questioned document was
written by the same person who wrote the known document set. The genre and/or topic
may differ significantly between the known and unknown docu-ments” [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
1.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data set</title>
      <p>
        The data set consists of author verification problems in four different languages. In each
problem, there are some known documents written by single person and only one
unknown document. The genre and/or topic between documents may differ significantly.
The document length varies from a few hundred to a few thousand words. Table 1 shows
the sub-corpora including their language and type (cross-genre or cross-topic)
The author verification systems are tested on a set of problems. The system must provide
a probability score for each unknown document. The performance of the system will
be evaluated using area under the ROC curve (AUC). In addition, the output will also
be measured based on c@1 score [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Probability score which is greater than 0.5 is
considered as positive answer, while a score lower than 0.5 is considered as negative. If
the score is 0.5, then it will be considered as an i don’t know answer. The c@1 measure
can be define as follows:
1
n
nc +
nu
nc
n
(1)
where:
– n = number of problems
– nc = number of correct answer
– nu = number of unanswered problems
      </p>
      <p>The overall performance will be evaluated on the product of AUC and c@1
2</p>
      <sec id="sec-2-1">
        <title>Methodological Approach</title>
        <p>
          We adopted a machine learning-based intrinsic method to address this verification
problem. Intrinsic methods use only the provided documents (in this case known and
unknown documents) to determine whether they are written by same author or not. A
machine learning algorithm then will be trained on the labeled document pairs to construct
a model which can be used to classify the unlabeled pairs. Note that in the
verification problems, the machine learning does not learn about the specific character of each
author, but rather about the differences between a pair of documents [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Texts are
represented by various types of features such as function word, character n-grams, word
n-grams and several stylometric features.
2.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Textual Representation</title>
      <p>As the genre and/or topic may differ significantly between the known and unknown
documents, we can not rely on the content based features to capture the differences
between documents. We therefore focused more on content-independent features such
as function words and stylometric features. In addition, those features can be applied to
any of the language used in the task. We used six types of features in total including:
stylometric features (10), function words, character 8-grams, character 3-grams, word
bigrams, and word unigrams.</p>
      <p>
        Given collection of problems P = fPi : 8i 2 Ig where I = f1; 2; 3; ::; ng is
the index of P . Pi contains exactly one unknown document U and a set of known
documents K = fKj : 8j 2 J g where J is the index of K and 1 J 5. Our
approach represented each problem Pi as vector Pi = fR1; R2; ::; Rng where n is the
maximum number of feature types (in our case are six). Ri is the distance of two similar
feature vector representation of a set of known documents K and unknown document U .
If K contains more than one document, then the generated feature vector is an average
vector of J documents. Table 2 shows details of the features vector representation and
comparison measures used.
Stylometric Features There are ten sytlometric features used in our experiment. Some
features were adapted from Guthrie’s work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] on anomalous text detection and were
among the most effective features to separate anomalous segments from normal
segments of the text. The complete list of stylometric features are:
1. Average number of non standard word1
2. Average number of words per sentence
3. Percentage of short sentences (less than 8 words)
4. Percentage of long sentences (greater than 15 words)
5. Percentage word with three syllables
6. Lexical diversity (ratio of total number of unique words to total number of words
in a document)
1 Enchant spell checking library (http://www.abisource.com/projects/enchant/) was used to
identify non-standard English words
      </p>
      <sec id="sec-3-1">
        <title>7. Total number of punctuations</title>
        <p>We also implemented three readibility measures:</p>
      </sec>
      <sec id="sec-3-2">
        <title>1. Flesch-Kincaid Reading Ease [4]</title>
        <p>ReadingEase = 206:835</p>
      </sec>
      <sec id="sec-3-3">
        <title>2. Flesch-Kincaid Grade Level [4]</title>
        <p>
          GradeLevel = 11:8
3. Gunning-Fog Index [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
        </p>
        <p>total_words
total_sentences
84:6
total_syllables</p>
        <p>total_words
total_syllables
total_words
We experimented with several different comparison measures for computing similarity
between a pair of vectors. We noticed particular comparison metric performs better in
certain type of features, thus we applied different measure for each features type. Three
different distance measures were used:</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Cosine similarity measure</title>
      <p>d(x; y) =
x:y
jjxjjjjyjj
=
p
P xiyi
i=1
s p</p>
      <p>P xi2
i=1
s p</p>
      <p>P yi2
i=1</p>
    </sec>
    <sec id="sec-5">
      <title>Minimum maximum similarity measure</title>
      <p>City block distance (also called Manhattan distance or L1 distance)
p</p>
      <p>P min(xi; yi)
minmax(x; y) = i=1
p
P max(xi; yi)
i=1
p
d(x; y) = X
jxi
yij
(2)</p>
      <p>100
(4)
(5)
(6)
(7)
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Feature selection and classifier</title>
      <p>Our authorship identification software was written in Python. We applied feature
selection using Extratreeclassifier and the SVM classifier. The classifier hyperparameters
were optimized using the GridSearchCV. Scikit-learn library2 was used for both feature
selection and classification.
3
3.1</p>
      <sec id="sec-6-1">
        <title>Evaluation and Result</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Training corpora</title>
      <p>We evaluated the approach on the training data using 10-fold cross validation. Since
there are some incompatibility issues, we did not perform the verification task on Greek
data. Table 3 shows the result of our approach on three of the sub-language corpora. The
best result was achieved on the Spanish data set with 0.846 and 0.807 for the AUC and
C@1 scores respectively. Compared to other sub-language corpora, the Spanish data set
contains more known documents; which may explain why the results on this data set
are better than the results on the other sub-languages data. We also observed that the use
of some NLP libraries which are mainly trained on English data did not perform well
on the non-English language data sets. Thus, feature selection was applied to remove
unhelpful features.</p>
      <sec id="sec-7-1">
        <title>2 http://scikit-learn.org</title>
        <p>4</p>
        <sec id="sec-7-1-1">
          <title>Conclusion</title>
          <p>This year’s author verification problem is considerably harder than last year’s since
the number of known documents is very limited and the genre/topic between known
and unknown documents differ significantly. In addition, for English, the data set was
derived from Project Gutenberg’s opera play scripts which are an unusual type of text.</p>
          <p>We identified that the most challenging part of this task was to find suitable features
which could capture the differences between documents. In addition, for certain data
set, not all features were helpful. Thus applying feature selection were beneficial and
greatly improved the accuracy of the classifier.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <source>PAN Authorship Identification Task</source>
          <year>2015</year>
          , https://www.uniweimar.de/medien/webis/events/pan-15/pan15-web/author-identification.html
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Gunning</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>The Technique of Clear Writing</article-title>
          .
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          (
          <year>1952</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Guthrie</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Unsupervised Detection of Anomalous Text</article-title>
          .
          <source>Ph.D. thesis</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kincaid</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fishburne</surname>
            ,
            <given-names>R.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chissom</surname>
            ,
            <given-names>B.S.:</given-names>
          </string-name>
          <article-title>Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel</article-title>
          .
          <source>Tech. Rep. February</source>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
          </string-name>
          , J.:
          <article-title>Authorship verification as a one-class classification problem</article-title>
          .
          <source>Twentyfirst international conference on Machine learning - ICML '04</source>
          p.
          <volume>62</volume>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Determining if two documents are written by the same author</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>65</volume>
          (
          <issue>1</issue>
          ),
          <fpage>178</fpage>
          -
          <lpage>187</lpage>
          (
          <year>Jan 2014</year>
          ), http://doi.wiley.
          <source>com/10</source>
          .1002/asi.22954
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Peñas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A simple measure to assess non-response</article-title>
          .
          <source>In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies</source>
          . pp.
          <fpage>1415</fpage>
          -
          <lpage>1424</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>