<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Adapting f or Subject-Specific Term Length using Topic Cost in Author Verification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anna Vartapetiance</string-name>
          <email>a.vartapetiance@surrey.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lee Gillam</string-name>
          <email>l.gillam@surrey.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computing, University of Surrey</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Previous PAN workshops have offered us the opportunity to explore three different approaches using basic statistics of stopword pairs for author verification. In this PAN, we were able to select our 'best' approach and explore the question of how authors writing about different subjects would necessarily adapt to term lengths specific to the subject. The adaptation required is, essentially, a redistribution of frequency: where longer terms occur. We introduce the notion of a 'topic cost' which increases the propensity for matching. Results show AUC and C1 scores of 0.51, 0.46 and 0.59 for Dutch, Greek and Spanish respectively. The English results are not yet available, as the evaluation system was unable to run the approach due to as yet unknown reasons.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the 6th International Workshop on Uncovering Plagiarism, Authorship,
        <xref ref-type="bibr" rid="ref1">and
Social Software Misuse (PAN2012</xref>
        ), we gave first test to our ideas on co-occurrence
patterns of stopwords [1]. At the 8th iter
        <xref ref-type="bibr" rid="ref2">ation (PAN 2014</xref>
        ), we presented 3 variations
to our approach, largely geared around evaluating use of similarity/distance over
vector spaces [2].
      </p>
      <p>
        In this paper, we suggest extension to our
        <xref ref-type="bibr" rid="ref2">approaches to the PAN2014</xref>
        by
accounting for a ‘topic cost’. Simply, there are several reasons why specific
stopword-pair separation may be less able to indicate similarity, and accounting for
term length and term count offers potential for addressing this. In section 2, we briefly
discuss the previous approaches we have used for author verification. Section 3
explains how we determine and use topic cost. Section 4 offers results and evaluation,
and Section 5 concludes the paper.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Previous methods applied</title>
      <p>
        As discussed in [1], for P
        <xref ref-type="bibr" rid="ref1">AN2012</xref>
        , we approached author ‘attribution’ using a
mean-variance framework on patterns of stopwords with a specified maximum
window size for pairs of the 10 most common English stopwords to identify
positional frequencies, and allocated an author based on nearest
frequency-meanvariance match.
      </p>
      <p>For PAN2013, the core approach remained the same with output adapted to the
Boolean output required. The task introduced Greek and Spanish texts, of which the
authors have no real knowledge, and so lists of 10 frequent stopwords were sought for
each.</p>
      <p>For PAN2014, we reused these stoplists along with a stoplist for Dutch – with
Dutch as yet another language of which the authors have no real knowledge. We also
evaluated 3 approaches based on:</p>
      <p>Frequency-Mean-Variance: We follow the approach detailed at length in
Vartapetiance and Gillam 2013, generating frequency information for stopword pairs,
determining mean and variance for separation, then applying cosine distance to
compare the resulting feature vectors.</p>
      <p>Positioning: This approach is based on FMV, above, but omits step 4 and so acts
as a cosine comparison on positional frequencies for each pattern. This would tend to
require comparable frequencies for each feature to ensure a good match.</p>
      <p>Cosine: We modify the Positioning approach to consider the frequency
information for all patterns as a single vector, then apply cosine distances between
resulting vectors. Here we also consider how to determine a match: a single cosine
distance between one known and one unknown; a difference in distance within a
threshold when two known texts can be compared; and distances between the
unknown and many known texts to be at a suitable point on the distribution of
distances amongst knowns. Acceptability, according to thresholds, and cosine
distance can then be used together to determine match confidence.
3</p>
      <p>PAN 2015</p>
      <p>For this year’s task, we wanted to explore the ability to match where the same
author may necessarily vary their writing according to the topic. This would account
for, say, simple temporal modification– discussing for example ‘the former Prime
Minister of’ rather than ‘the Prime Minister of’ – but is principally geared to account
for differences in term lengths as relate to topics. In the ‘Prime Minister’ example
given, the same stopword pair of the-of is present, but with a positional mismatch.
Since position, and variability in position, is core to our approaches, we require a
simple way to address the pattern-specific positional mis-alignment that occurs.</p>
      <p>To approach this, we introduce the notion of a ‘topic cost’ and distribute positional
frequencies according to this topic cost. To determine topic cost, we simply count the
number of terms and the length of these terms, and use the difference between these
values for redistribution. The only additional resource employed is a
languagespecific stoplist as exposes the terms.</p>
      <p>As an example, consider the following passage of text:</p>
      <p>UK interest rates have been kept unchanged again by the Bank of England,
meaning they have now been at their record low of 0.5% for six years. Rates
were first cut to 0.5% in March 2009 as the Bank sought to lift economic
growth amid the credit crunch.</p>
      <p>Take stopword pairs as formed from [the, of, in, for, to]. If we ignore the sentence
break, the first pair of interest offers us: “for six years. Rates were first cut to”. The
distance covered by the pair is 6 (the number of words between “for” and “to”).
Collecting all multi-word terms, using all stopwords (not just those listed) as
delimiters (and, here, the full-stop also), results in 3 terms comprising 5 words – six
years, rates, first cut. The topic cost, then, is 2. Instead of counting once at position 6,
we uniformly distribute – other weightings possible but unexplored - across position 6
and the two preceding positions and so positions 4, 5 and 6 each receive 0.333. This
example, and further from the above passage, are shown in the table below.</p>
      <p>Results for each of the PAN 2015 collections are shown in the table below based
on 4 language categories.</p>
      <p>Due to yet unknown problem with English run, the system was unable to calculate
the outcomes of the test. Also, unfortunately, the results from the runs using last
year’s systems will not be available until after this paper is submitted, so the authors
are not able to provide a comparison between systems to see whether or not this
approach improves the outcome of detection. However, the results on runs on training
datasets using FMV, Positioning and Topic Cost systems (Table 3) show some
improvements in detection using the new system.</p>
      <p>
        In this paper, we suggested an extension to our
        <xref ref-type="bibr" rid="ref2">approaches to PAN2014</xref>
        for
authorship verification by accounting for a ‘topic cost’. For us, topic cost may account
for lower match values in our previous approaches, and our intention was to
determine whether a simple treatment of topic cost could improve our results. This
modification does require much more testing in respect to the test collections of
previous years to fully appreciate its effect. Unfortunately, other activities hindered
the authors’ abilities to allocate sufficient time to this testing during this round of
PAN.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgements</title>
      <p>The authors gratefully acknowledge prior funding from the UK’s Technology
Strategy Board (TSB, 169201), and also the efforts of the PAN organizers in crafting
and managing the tasks.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vartapetiance</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Gillam</surname>
          </string-name>
          , “
          <article-title>Quite Simple Approaches for Authorship Attribution , Intrinsic Plagiarism Detection</article-title>
          and
          <string-name>
            <surname>Sexual Predator</surname>
          </string-name>
          Identification - Notebook
          <source>for PAN at CLEF</source>
          <year>2012</year>
          ,” in
          <source>Working Notes Papers of the CLEF 2012 Evaluation Labs</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vartapetiance</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Gillam</surname>
          </string-name>
          , “
          <article-title>A Trinity of Trials : Surrey ' s 2014 Attempts at Author Verification Notebook for PAN at CLEF</article-title>
          <year>2014</year>
          ,” Work. Notes Pap.
          <source>CLEF 2014 Eval. Labs</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>