<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the Author Identification Task at PAN-2017: Style Breach Detection and Author Clustering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Tschuggnall</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Efstathios Stamatatos</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ben Verhoeven</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Walter Daelemans</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Günther Specht</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benno Stein</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Potthast</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bauhaus-Universität Weimar</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Antwerp</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Innsbruck</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of the Aegean</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Several authorship analysis tasks require the decomposition of a multiauthored text into its authorial components. In this regard two basic prerequisites need to be addressed: (1) style breach detection, i.e., the segmenting of a text into stylistically homogeneous parts, and (2) author clustering, i.e., the grouping of paragraph-length texts by authorship. In the current edition of PAN we focus on these two unsupervised authorship analysis tasks and provide both benchmark data and an evaluation framework to compare different approaches. We received three submissions for the style breach detection task and six submissions for the author clustering task; we analyze the submissions with different baselines while highlighting their strengths and weaknesses.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        An authorship analysis extracts information about the authors of given documents.
There are several related supervised tasks where a set of documents with known
information about its authors is available and which can be used to train a model that can
extract this information from other documents. Typical examples are authorship
attribution (extract the identity of authors) [
        <xref ref-type="bibr" rid="ref47">47</xref>
        ] and author profiling (extract demographics
such as age and gender of the authors) [
        <xref ref-type="bibr" rid="ref40">40</xref>
        ]. The vast majority of published work
focus on these two tasks. However, there are cases where authorship-related information
in a set of training documents is neither available nor reliable. Examples of
unsupervised tasks are intrinsic plagiarism detection (identification of plagiarized parts within
a given document without a reference collection of authentic documents) [
        <xref ref-type="bibr" rid="ref51">51</xref>
        ], author
clustering (grouping documents by authorship) [
        <xref ref-type="bibr" rid="ref32 ref44">32, 44</xref>
        ], and author diarization
(decomposing a multi-authored document into authorial components) [
        <xref ref-type="bibr" rid="ref12 ref2 ref29">2, 12, 29</xref>
        ]. Unsupervised
authorship analysis tasks are more challenging but can be applied to every authorship
analysis case since they do not require any training material.
      </p>
      <p>
        Previous editions of PAN focused on specific unsupervised tasks such as author
clustering or author diarization [
        <xref ref-type="bibr" rid="ref50">50</xref>
        ]. However, it has been observed that it was very
difficult for the submitted approaches to surpass even naive baseline methods. Given
the complexity of unsupervised tasks, it is essential to focus on fundamental problems
and to study them separately. In the current edition of PAN, we focus on two such
fundamental problems:
1. Segmentation of a multi-authored document into stylistically homogeneous parts.
      </p>
      <p>We call this task style breach detection.
2. Grouping of paragraph-length document parts by authorship. We call this task
author clustering.</p>
      <p>These two tasks are elementary processing steps for both author diarization and
intrinsic plagiarism detection. Style breach detection could also be useful in writing style
checkers, where it is required to ensure that homogeneous stylistic properties are found
within a document. Moreover, author clustering of short (paragraph-length) documents
could be useful in analysis of social media texts such as blog posts, comments, and
reviews. For example, author clustering could help to identify different user names that
correspond to the same person or user accounts that are used by multiple persons.</p>
      <p>In this paper we present an overview of the shared tasks in style breach detection
and author clustering at PAN-2017. We received three submissions for the former and
six submissions for the latter task. The evaluation framework including benchmark data,
evaluation measures, and baseline methods is described. In addition, we present an
analysis and a survey of the submitted methods.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Previous Work</title>
      <p>This section reviews related work on style breach detection and author clustering.
2.1</p>
      <sec id="sec-2-1">
        <title>Style Breach Detection</title>
        <p>
          The goal is to find positions within a document where the authorship changes, i.e.,
where the style changes. Thus, it is closely related to all fields within stylometry,
especially intrinsic plagiarism detection [
          <xref ref-type="bibr" rid="ref52">52</xref>
          ]. Several approaches exist that deal with the
latter, basically by creating stylistic fingerprints that include lexical features such as
character n-grams [
          <xref ref-type="bibr" rid="ref30 ref48">30, 48</xref>
          ], word frequencies [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] or average word/sentence lengths
[57], syntactic features such as Part-of-Speech (POS) tag frequencies/structures [
          <xref ref-type="bibr" rid="ref54">54</xref>
          ],
structural features such as average paragraph lengths, or indentation usages [57]. Using
these fingerprints, outliers are sought, either by applying different distance metrics on
sliding windows [
          <xref ref-type="bibr" rid="ref48">48</xref>
          ] or by storing distance matrices [
          <xref ref-type="bibr" rid="ref24 ref53">53, 24</xref>
          ].
        </p>
        <p>
          In contrast, related work targeting multi-author documents is rare. One of the first
approaches that uses stylometry to automatically detect boundaries of authors of
collaboratively written text was proposed by Glover and Hirst [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], with the aim to provide
hints in order to produce a homogeneously written text. Graham et al. [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] utilize
neural networks with several stylometric features, and Gianella [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] proposes a stochastic
model on the occurrences of words to split a document by authorship. An unsupervised
decomposition of multi-author documents based on grammar features has been
evaluated by Tschuggnall and Specht [56].
        </p>
        <p>
          The diarization task at PAN-2016 [
          <xref ref-type="bibr" rid="ref49">49</xref>
          ] dealt with building author clusters within
documents. The two submitted approaches use n-grams and other selected stylometric
features in combination with a classifier post-processed by a Hidden Markov Model
[
          <xref ref-type="bibr" rid="ref31">31</xref>
          ], as well as a sentence-based distance metric, computed from several features, that
is given to a k-means algorithm in order to build clusters [
          <xref ref-type="bibr" rid="ref46">46</xref>
          ].
        </p>
        <p>
          From a global point of view, style breach detection can also be seen as a text
segmentation problem that where a document is split into segments based on the writing
style. Common text segmentation approaches divide a text by different topics and/or
genres [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Compared to an intrinsic stylometric analysis, those approaches have the
advantage to be able to build dictionaries or other useful statistics for each targeted
topic or genre in advance. Thereby, a wide range of methods is used, often based on the
research by Hearst [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], in which the lexical cohesion of terms is analyzed. Other
approaches use Bayesian models [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], Hidden Markov Models [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], vocabulary analysis in
various forms such as word stem repetitions [
          <xref ref-type="bibr" rid="ref36">36</xref>
          ] or word frequency models [
          <xref ref-type="bibr" rid="ref41">41</xref>
          ]. While
some of the recent papers [
          <xref ref-type="bibr" rid="ref34 ref42">42, 34</xref>
          ] compare the segmentation approaches on the same
data sets, it is in general difficult to compare performances due to the heterogeneous
problems and data types.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Author Clustering</title>
        <p>
          Previous work on author clustering (also called author-based clustering, authorship
clustering, or authorial clustering), as it is defined in this paper, is limited. Iqbal et
al. [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] describe an approach based on k-means clustering which requires that the
number of authors is known, and apply it to a collection of e-mail messages. Layton et al.
[
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] propose a method that can automatically estimate the number of clusters (authors)
in a collection of documents using the iterative positive Silhouette method. The latter
has been demonstrated to be useful for clustering validation purposes [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]. These
techniques have been applied to literary texts (either books or book samples). Samdani et al.
[
          <xref ref-type="bibr" rid="ref44">44</xref>
          ] analyze postings in a discussion forum using an online clustering method. Daks &amp;
Clark [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] use POS n-grams and spectral clustering and tested their method in a variety
of corpora including newspaper articles, political speeches, and literary texts.
        </p>
        <p>
          A shared task on author clustering documents was included in PAN-2016 [
          <xref ref-type="bibr" rid="ref50">50</xref>
          ]. The
benchmark collections built for this task comprised texts in three languages (English,
Dutch, and Greek) and two genres (opinion articles and reviews) taken from various
sources. A total of eight submissions was received and the best-performing model was
based on a successful authorship verification method using a character-level
multiheaded recurrent neural network [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], closely followed by a simple approach based on
word and punctuation mark frequencies [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. Both of these methods first compute
pairwise distances between texts and then form clusters by joining texts that belong to a
path of small distances. In general, simple baseline methods, such as placing each text
in a distinct cluster were found very competitive since in the benchmark collections the
number of single-item clusters was high (&gt;50%) or very high (&gt;75%) [
          <xref ref-type="bibr" rid="ref50">50</xref>
          ].
As a specific type of author identification, the style breach detection task at PAN-2017
focuses on finding stylistic differences within a text document as a result of having
multiple authors collaborating on it. The main goal is to identify style breaches, i.e.,
exact positions in the text where the authorship changes. Thereby no training data is
available for the corresponding authors, nor can respective information be gained from
potential web searches. From this perspective, this year’s task attaches to a series of
subtasks of previous PAN events that focused on intrinsic characteristics of text
documents. Including intrinsic plagiarism detection [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ] and author diarization [
          <xref ref-type="bibr" rid="ref49">49</xref>
          ], the
main commonality is that the style of authors has to be extracted and quantified in some
way in order to tackle those problems. In a similar way, an intrinsic analysis of the
writing style is also key to approach the PAN-2017 style breach detection task, which can
be summarized as follows:
        </p>
        <sec id="sec-2-2-1">
          <title>Given a document determine whether it is multi-authored</title>
          <p>and, if yes, find the borders where authorship switches.</p>
          <p>In contrast to the author clustering task described in Section 4, the goal is to find
only borders, and thus it is irrelevant to identify or cluster authors of segments.</p>
          <p>
            The detection of style breaches, i.e., locating borders between different authors, can
be seen as an application of a general text segmentation problem. Nevertheless, a
significant difference to existing segmentation approaches is that while the latter usually
focus on detecting switches of topics or stories [
            <xref ref-type="bibr" rid="ref20 ref34 ref42">20, 42, 34</xref>
            ], the aim of this subtask is
to identify borders based on the writing style, disregarding the specific content. While
segmentation algorithms may include metrics built from precomputed dictionaries
comprising different topics or genres, an additional difficulty results from the fact that the
content type is not known and, more importantly, coherent throughout a document.
3.1
          </p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Approaches at PAN-2017</title>
        <p>
          This year, five teams registered to the style breach detection task, whereas three of them
actively submitted their software to TIRA [
          <xref ref-type="bibr" rid="ref16 ref38">16, 38</xref>
          ]. In the following, a short summary
of each approach is given. Moreover, the creation of two slightly different baseline
approaches for comparison is explained.
        </p>
        <p>
          – Karas´, S´piewak &amp; Sobecki [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The authors start by splitting the document into
paragraphs, either by detecting multiple newline characters, or, if there are no
newlines, by choosing a fixed number of sentences. In the following it is then assumed
that style breaches may occur only on paragraph endings. To quantify the style of
each paragraph, tf-idf matrices are computed using single words, word 3-grams,
stop words, POS tags, and punctuation characters. By concatenating all tf-idf
matrices, a paired samples Wilcoxon Signed Rank test [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is applied—a statistical
method to verify whether two given samples stem from the same distribution or
not. Computing this test for all pairs of consecutive paragraphs finally yields the
final prediction, where a style breach is predicted if the test suggests that the
paragraphs come from a different distribution.
– Khan [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. The author also segments the document into sentences within a
preprocessing step. Then, the sentences are traversed using two sliding windows that
share a sentence in the middle. For each window several statistics are computed,
including most frequent POS tags, non-/alphanumeric characters, or words.
Moreover, word statistics based on precomputed dictionaries are utilized which include
common English words and several sentiment dictionaries. Using all metrics, a
similarity score between the two adjacent sliding windows is calculated which is
finally compared to a predefined threshold in order to decide whether or not the
position before/after the overlapping sentence is predicted as style breach. In the latter
case the two sliding windows are merged and considered to be written by a single
author; a new second window is created, which is processed as described earlier.
– Safin &amp; Kuznetsova [
          <xref ref-type="bibr" rid="ref43">43</xref>
          ]. The authors approach the style breach detection task by
applying a sentence outlier detection, commonly used in intrinsic plagiarism
detection algorithms [
          <xref ref-type="bibr" rid="ref48 ref55">48, 55</xref>
          ]. After splitting the document into sentences, each one
is vectorized using two pretrained skip-thought models [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]. These models can be
seen as word embeddings operating on sentences as the atomic units, thereby
resulting in 2,400 dimensions for each sentence. The distance between each distinct
pair of sentences is stored in a distance matrix (similar to, e.g., [
          <xref ref-type="bibr" rid="ref24 ref53">53, 24</xref>
          ]) by
calculating the cosine distance of the corresponding vectors. Finally, an outlier detection
is performed using an optimized threshold, which is compared to the average
distance of each sentence. If the distance of a sentence s is larger than the threshold,
the beginning of s is marked as a style breach position.
– Baselines. To be able to compare the performances of the submitted approaches,
two simple baselines have been computed:
1. BASELINE-rnd randomly places between 0 and 10 borders at arbitrary
positions inside a document.
2. As a slight variant, BASELINE-eq also decides on a random basis how many
borders should be placed (also 0-10), but then places the borders uniformly,
i.e., such that all resulting segments are of equal size with respect to tokens
contained.
        </p>
        <p>Both baselines have been computed based on the average of 100 runs.
3.2</p>
      </sec>
      <sec id="sec-2-4">
        <title>Data Set</title>
        <p>
          To develop and optimize the respective algorithms, distinct training and test data sets
have been provided, which are based on the Webis-TRC-12 data set [
          <xref ref-type="bibr" rid="ref39">39</xref>
          ]. The original
corpus already served as origin for the PAN’16 diarization data set [
          <xref ref-type="bibr" rid="ref49">49</xref>
          ] and contains
documents on 150 topics used at the TREC Web Tracks from 2009-2011 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which was
created by hiring professional writers through crowdsourcing. Each writer was asked
to search for a given topic including assignments (e.g., “Barack Obama”, assignment:
include information about Obama’s family) and to compose a single document from the
search results. All sources of the resulting document are annotated respectively, so the
origin of each text fragment is known.
        </p>
        <p>Assuming that each distinct source represents a different author in the original data
set, a training and a test data set have been randomly created from these documents by
varying several parameters as shown in Table 1. Beside the number of style breaches
or collaborating authors, also authorship boundary types have been altered to be at
paragraph or sentence levels, i.e., authors may switch only at the end of paragraphs1
or also within paragraphs. Nevertheless, in order to not overcomplicate the task and to
build more realistic data sets, the atomic units were set to be sentences, i.e., borders
may not occur within sentences. With respect to the resulting segment lengths, it has
been varied whether they are equalized to be of similar lengths or of random lengths
within a document.</p>
        <p>As the original corpus has been partly used and published, the test documents have
been created from previously unpublished documents only. Overall, the number of
documents in the training data set is 187, whereas the test data set contains 99 documents.
The final statistics of the generated data sets are presented in Table 2.
3.3</p>
      </sec>
      <sec id="sec-2-5">
        <title>Experimental Setup</title>
        <p>
          Performance Measures The performance of the submitted algorithms have been
measured with two common metrics used in the field of text segmentation. The WindowDiff
metric [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ], which is proposed for general text segmentation evaluation, is computed
as it still is used widely for such problems. It calculates an error rate between 0 and 1
for predicting borders (whereby 0 indicates a perfect prediction), by penalizing
nearmisses less than other/complete misses or extra borders. Depending on the problem
types and data sets used, text segmentation approaches report near-perfect windowDiff
values of less than 0.01, while on the other side the error rate exceeds values of 0.6
and higher under certain circumstances [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. As an alternative, a more recent adaption
of the WindowDiff metric is the WinPR metric [
          <xref ref-type="bibr" rid="ref45">45</xref>
          ]. It enhances WindowDiff by
computing the common information retrieval measures precision (WinP) and recall (WinR)
and thus allows to give a more detailed, qualitative statement about the prediction.
Internally, WinP and WinR are computed based on the calculation of true and false
positives/negatives, respectively, also assigning higher values if predicted borders are closer
to the real border position.
        </p>
        <p>Both metrics are computed on a word-level, whereby the participants were asked
to provide character positions. This means that the tokenization was delegated to the</p>
        <sec id="sec-2-5-1">
          <title>1 to be identified by at least two consecutive line breaks</title>
          <p>evaluator script. For the final ranking of all participating teams, the F-score of WinPR
(WinF) is used.</p>
          <p>
            Workflow The participants designed and optimized their approaches with the given,
publicly available training data set described earlier. The performance could be
measured either locally using a provided evaluator script, or by uploading the respective
software to TIRA [
            <xref ref-type="bibr" rid="ref16 ref38">16, 38</xref>
            ] and running it against the training data set. Because the test
data set was not publicly available, it was necessary to use the latter option in this case.
I.e., the participants submitted their final software and ran it against the test data without
seeing performance results. It was manually ensured that no potentially helpful
information about the data set was publicly logged during the execution. Finally, participants
were allowed to submit three successful test data runs, whereby the latest submissions
are used for the final ranking and for all results presented in Section 3.4.
3.4
          </p>
        </sec>
      </sec>
      <sec id="sec-2-6">
        <title>Results</title>
        <p>
          The final results of the three submitting teams are shown in Table 3. In case of WinF, the
baseline equalizing the segment sizes could be exceeded by only one approach, whereas
the baseline using completely random positions could be outperformed by all
participants. With respect to WindowDiff, all approaches perform better than both baselines.
Interestingly, besides achieving the best WinF performance, Karas´ et al. also needed
the shortest runtime for the prediction, whereas Safin et al. required the significantly
longest runtime with over 20 minutes, probably by applying the cost-intensive neural
sentence embedding technique [
          <xref ref-type="bibr" rid="ref43">43</xref>
          ].
        </p>
        <p>Figure 1 depicts details about performances of all approaches including
BASELINE"=eq with respect to several parameters. In case of number of style breaches (a), it
can be seen that there is a significant difference between the approaches when analyzing
documents with no author switches. While the overall winning approach of Karas´ et al.
performs poorly, Safin et al. achieve their best score for these documents. The result
of the latter may be caused by the intrinsic plagiarism detection type of approach that
distinguishes between documents containing suspicious sentences and plagiarism-free
documents, i.e., containing style breaches or not. The other approaches assume style
borders to be existent, which accounts also for the baseline in over 90% of the cases as
it chooses a random number of borders between 0-10.</p>
        <p>While the number of authors (b) seems to have no significant impact on the
performances2, the document length (c) influences the results, especially for very short and
very long documents, respectively. The approach of Karas´ et al. basically gets better
with the length of the text, achieving best results for the majority of documents within
1,000-2,000 words. Khan achieves good results for both short and long documents,
while Safin et al. scores good only in the former case. With respect to the average
segment length, the performance of the winning approach of Karas´ et al. decreases
drastically for segment lengths of over 500 words. Nevertheless, it achieves good results for
documents with shorter segments, and, remarkably, the highest score for the documents
of very short segment lengths.</p>
        <p>Finally, the impact of the border position and the segment length distribution is
shown in subfigures (e) and (f) respectively. For the border position, only Karas´ et al.
indicate a significant improvement when style breaches appear only on paragraph ends.
This reflects also their approach, which treats paragraphs as potential natural border
positions, and if no paragraphs exist, creates artificial paragraphs using a fixed length
of sentences. Moreover, this may also be the reason why the approach performs better
for segments of similar lengths, as this scenario better matches the specified creation of
artificial paragraphs.
2 except for the distinction between one or more authors, which is already shown in subfigure (a),
where number of style breaches = 0 corresponds to number of authors = 1
iFn0,3
W
0,2
0,1
0
0,6
0,5
0,4
F
in0,3
W
0,2
0,1
0
0,6
0,5
0,4
F
in0,3
W
0,2
0,1
0
0 (20) 1 (15) 2 (13) 3 (16) 4 (12) 5 (9) 6 (4) 7 (10)
2 (26)
3 (18)
4 (24)</p>
        <p>5 (11)
(b) Number of Authors
0,4
iFn0,3
W
0,2
0,1
0
0,6
0,5
0,4
F
in0,3
W
0,2
0,1
0
0,6
0,5
0,4
F
in0,3
W
0,2
0,1
0
sentence(46)
paragraph (53)
equalized (55)
random (44)
(e) Border Position (f) Segment Length Distribution
Figure 1. Style breach detection results with respect to number of style breaches, number of
authors, document length, average segment length, border position and segment length distribution.
To highlight the potential of the approaches, their individual best results are listed in
Table 4. The upper part shows the best configurations for single-authored documents,
while the lower part presents the best performances for documents containing style
breaches. Again it can be seen that Karas` et al. assume style breaches to be existent and
thus reaches very poor results if a document contains no breaches. On the other side,
Khan as well as Safin et al. achieve perfect prediction rates, i.e., estimating correctly
that there are no style borders3. In case of documents containing style breaches, Karas´
et al. and Khan gain very good top results with WinF scores of over 80%.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Author Clustering</title>
      <sec id="sec-3-1">
        <title>Task Definition</title>
        <p>
          Given a collection D of short (paragraph-length) documents we approach the author
clustering task following two scenarios:
– Complete Clustering. All documents should be assigned to clusters whereas each
cluster corresponds to a distinct author. More specifically, each document d 2 D
should be assigned to exactly one of k clusters, while k is not given.
– Authorship-Link Ranking. Pairs of documents by the same author
(authorshiplinks) should be extracted. For each authorship-link (di; dj ) 2 D D, a
confidence score belonging to [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ] should be estimated and authorship-links are ranked
in decreasing order.
        </p>
        <p>
          All documents within a clustering problem are single-authored, in the same
language, and belong to the same genre; however, topic and text-length may vary. The
main difference with respect to the corresponding PAN-2016 [
          <xref ref-type="bibr" rid="ref50">50</xref>
          ] task is that the
documents are short including a few sentences (paragraph length). This makes the task
harder since text-length is crucial when attempting to extract stylometric information.
4.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Evaluation Datasets</title>
        <p>The datasets used for training and evaluation were extracted from the corresponding
PAN-2016 corpora that include three languages (English, Dutch, and Greek) and two
genres (articles and reviews). Each PAN-2016 text was segmented into paragraphs and
3 not shown in the Table, Khan and Safin et al. achieve perfect prediction for several of the
documents containing no style breaches</p>
      </sec>
      <sec id="sec-3-3">
        <title>Language</title>
        <p>English
g English
inn Dutch
ira Dutch
T Greek</p>
        <p>Greek
English</p>
        <p>English
ts Dutch
T Dutch
e</p>
        <p>Greek
Greek</p>
        <p>Genre
articles
reviews
articles
reviews
articles
reviews
articles
reviews
articles
reviews
articles
reviews
all paragraphs with less than 100 characters and more than 500 characters were
discarded. In each clustering problem, documents by the same authors were selected
randomly by all original documents. This means that paragraphs of the same original
document or other documents (by the same author) may be grouped. Certainly, when
paragraphs come from the same original document, there is much larger thematic similarity.
The only exception in this process was the Dutch reviews corpus because the texts were
already short (one paragraph each). In this case, the PAN-2017 datasets were built using
the PAN-2016 procedure.</p>
        <p>Table 5 shows details about the training and test datasets. Most of the clustering
problems include 20 documents (paragraphs) by an average of 6 authors. In each
clustering problem there is an average of about 50 authorship links and the largest cluster
contains about 8 documents. Each document has an average of about 50 words. Note
that in the case of Dutch reviews these figures deviate from the norm (documents are
longer and authorship links are less).</p>
        <p>
          An important factor to each clustering problem is the clusteriness ratio r = k=N ,
where N is the size of D. When r is high, most documents belong to single-item clusters
and there are few authorship links. When r is low, most documents belong to multi-item
clusters and there are plenty of authorship links. Estimating r (since N is known, k
should be estimated) is crucial for each clustering algorithm. In PAN-2016 three specific
values r=0.5, r=0.7, and r=0.9 were used focusing on relatively high values of r [
          <xref ref-type="bibr" rid="ref50">50</xref>
          ].
In the current edition, in both training and test datasets, r ranges between 0.1 and 0.5. as
can be seen in Figure 2. This means that the PAN-2017 corpus has far less single-item
clusters in comparison to PAN-2016.
4.3
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Performance Measures</title>
        <p>
          The same evaluation measures introduced in PAN-2016 are used. As a consequence,
the results are directly comparable to the ones from the corresponding PAN-2016 task.
In more detail, for the complete clustering scenario, Bcubed Recall, Bcubed Precision,
and Bcubed F-score are calculated. These are among the best extrinsic clustering
evaluation measures and were found to satisfy several important formal constraints including
cluster homogeneity, cluster completeness, etc. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] With respect to the authorship-link
ranking scenario, established measures are used to estimate the ability of systems to
rank high correct results. These are Mean Average Precision (MAP), R-precision, and
P@10.
To understand the complexity of the tasks and the effectiveness of participating systems
we used a set of baseline approaches and applied them to the evaluation datasets. The
baseline methods range from naive to strong and will allow to estimate weaknesses and
strengths of participant approaches. More specifically, the following baseline methods
were used:
– BASELINE-Random. Given a set of documents, the method randomly chooses the
number of authors and randomly assigns each document to one of the authors. It
extracts all authorship links from the produced clusters and assigns a random score
to each one of them. The average performance of this method over 30 repetitions is
reported. This naive approach can only serve as an indication of the lowest
performance.
– BASELINE-Singleton. This method sets k = N , that is all documents are by
different authors. It forms singleton clusters and no authorship links. As a result, it is
used only for the complete clustering scenario. This simple method was found very
effective in PAN-2016 datasets and its performance increases with r [
          <xref ref-type="bibr" rid="ref50">50</xref>
          ]. Since
the range or r is lower in PAN-2017 datasets, its performance should be negatively
affected.
– BASELINE-Cosine. Each document is represented by the normalized frequencies
of all words occurring at least 3 times in the given collection of documents. Then,
for each pair of documents the cosine similarity is calculated and it is used as an
authorship-link score. This simple method is only used in the authorship-link
ranking scenario and it was found hard-to-beat in PAN-2016 evaluation edition [
          <xref ref-type="bibr" rid="ref50">50</xref>
          ].
– BASELINE-PAN16. This is the top-performing method submitted to the
corresponding PAN-2016 task. It is based on a character-level recurrent neural network
and it is a modification of an effective authorship verification approach [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. There
was no attempt to modify this method to be more suitable for the PAN-2017 corpus.
Given that it follows a highly conservative approach to form multi-item clusters
(suitable for the PAN-2016 corpus) its Bcubed recall is expected to be very low in
the current corpus.
4.5
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Survey of Submissions</title>
        <p>
          We received six submissions from research teams in Cuba [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], Germany [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], the
Netherlands [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], Mexico [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], Poland [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], and Switzerland [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. All participants also
submitted a notebook paper describing their approach.
        </p>
        <p>
          In general, all submissions follow a bottom-up paradigm where first the pairwise
similarity between any pair of documents is estimated and then this information is used
to form clusters. Gómez-Adorno et al. use hierarchical agglomerative clustering [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]
while García et al. use -compact graph-based clustering. Kocher &amp; Savoy apply some
merging criteria in the pairwise similarities [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. Alberts [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] proposes a modification
of a similar method submitted to PAN-2016 [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. Halvani &amp; Graner use the k-medoids
clustering algorithm and Karas´ et al. are based on a variation of locality-sensitive
hashing [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ].
        </p>
        <p>To calculate the pairwise (dis)similarity between documents in a given collection
Alberts and Kocher &amp; Savoy propose simple formulas that compare two probability
distributions. García et al. use Dice index, Gómez-Adorno et al. use cosine similarity
while Halvani &amp; Graner are based on the Compression-based Cosine measure.</p>
        <p>
          A crucial issue is how to estimate the number of clusters k in a given collection
of documents. A common choice is the use of Silhouette coefficient to indicate the
most suitable number of clusters [
          <xref ref-type="bibr" rid="ref19 ref23">19, 23</xref>
          ] while Gómez-Adorno et al. use the
CalinskiHarabasz index [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Another idea is the use of graph-based clustering methods that can
be automatically adopted to a clustering problem [
          <xref ref-type="bibr" rid="ref1 ref11">1, 11</xref>
          ].
        </p>
        <p>
          As concerns the stylometric measures used to quantify the personal style of authors,
most of the submissions are based on low-level character or lexical features such as
word and character n-grams. García et al. was the only submission experimenting with
higher-level features requiring NLP tools such as lemmatizers and POS taggers, only for
the English datasets. Some submissions used a single type of features (e.g., character
ngrams [
          <xref ref-type="bibr" rid="ref1 ref28">1, 28</xref>
          ]) while others used a pool of different feature types and attempted to select
the most suitable type (or combination of types) for each language and genre [
          <xref ref-type="bibr" rid="ref11 ref17">17, 11</xref>
          ].
A feature-agnostic compression-based approach is proposed by Halvani &amp; Graner [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
All participant methods were submitted to the TIRA experimentation platform where
the participants were able to run their software on both training and test datasets [
          <xref ref-type="bibr" rid="ref16 ref38">16,
38</xref>
          ]. PAN organizers provided feedback in case a run produced errors or unexpected
output. The participants were given the opportunity to perform at most two runs on the
test dataset and have been informed about the evaluation results. However, the last run
was always considered for the final evaluation.
        </p>
        <p>
          Table 6 shows the overall evaluation results for both complete clustering and
authorship-link ranking on the entire test dataset. The elapsed runtime of each
submission is also reported. As can be seen, the method of Gómez-Adorno et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] achieves
the best results in both scenarios. Actually, this is the top-performing method taking
into account all but one evaluation measures, that is BCubed precision. By definition,
BASELINE-singleton achieves perfect Bcubed precision since it provides single-item
clusters exclusively. Moreover, BASELINE-PAN16 attempts to optimize precision by
following a very conservative strategy when multi-item clusters are considered. Within
the PAN-2017 submissions, the approaches of García et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and Kocher &amp; Savoy
[
          <xref ref-type="bibr" rid="ref28">28</xref>
          ] are the best ones in terms of Bcubed precision. However, the winning approach of
Gómez-Adorno et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] is the only one that achieves both Bcubed recall and precision
higher than 0.6. As concerns efficiency, almost all submitted approaches are quite fast.
The approaches of García et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and Halvani &amp; Graner are relatively slower than the
rest of submissions. However, both of them are much faster than the very demanding
approach of BASELINE-PAN16.
        </p>
        <p>
          Table 7 provides a more detailed view of performance (Bcubed F-score) in each
evaluation dataset separately for the complete clustering scenario. All submitted
methods are better than BASELINE-Singleton and BASELINE-Random in overall results.
Actually, the performance of these two baseline methods is quite similar, in contrast
to the results of PAN-2016 [
          <xref ref-type="bibr" rid="ref50">50</xref>
          ]. Moreover, all but one submission were better than
BASELINE-PAN16. These observations can be explained by the low values of
clusteriness ratio (r) used in PAN-2017 datasets. This means that single-item clusters are not
the majority in PAN-2017 datasets and approaches that attempt to optimize precision
over recall are not equally effective as in PAN-2016. Note that in the case of Dutch
reviews where r is higher, BASELINE-Singleton and BASELINE-PAN16 are improved.
        </p>
        <p>
          The approaches of Gómez-Adorno et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], Kocher &amp; Savoy [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], and Halvani
&amp; Graner seem to be more effective on articles rather than reviews, while the method
of García et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] is not affected significantly by genre. Moreover, the methods of
García et al. and Kocher &amp; Savoy seem to be better able to handle English and Dutch
texts rather than Greek.
        </p>
        <p>
          Figures 3 and 4 show Bcubed recall and precision for varying values of the
clustering ratio r in the entire test dataset. As can be seen, the tendency for Bcubed recall
is to improve, while Bcubed precision is decreased as r increases. BASELINE-PAN16
suffers in recall that increases almost linearly with r while it maintains almost perfect
precision. The approach of Gómez-Adorno et al. achieves the most balanced recall and
precision scores especially for relatively low r values. The rest of submissions either
favor recall (Alberts [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], Halvani &amp; Graner [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], Karas´ [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]) or precision (García et al.
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], Kocher &amp; Savoy [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]).
        </p>
        <p>
          Table 8 shows the evaluation results (MAP) per language and genre for the
authorship-link ranking scenario. Here, BASELINE-PAN16 is quite competitive and
only the method of Gómez-Adorno et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] is able to surpass it. Moreover,
BASELINE-Cosine is quite strong and outperforms half of submissions. Recall that
the winning approach of Gómez-Adorno et al. is also based on cosine similarity using a
richer set of features and a log-entropy weighting scheme. In general, almost all
submissions achieve their worst results in the Dutch reviews dataset. Recall from Table 5 that
this dataset has distinct characteristics. Despite the fact that it contains longer texts with
respect to the rest of datasets, Dutch reviews form the most difficult case. It seems that
the method of Gómez-Adorno et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], Kocher &amp; Savoy [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], and Karas´ et al. [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]
are better in handling articles than reviews. The same is true for BASELINE-Cosine,
indicating that thematic information is more useful in articles. In the authorship-link
scenario, the language factor seems not to be crucial since the evaluation results are
balanced over the three examined languages.
The author identification task at PAN-2017 focused on unsupervised author analysis
by decomposing text documents into their authorial components. To study different
aspects in detail, two separate subtasks were addressed: (1) style breach detection, aiming
to segment a document by stylistic characteristics, and (2) author clustering, aiming to
group paragraph-length texts by authorship as well as assigning confidence scores
between documents written by the same author. For both tasks, comprehensive data sets
have been provided, which allowed participants to train their approaches on the
respective training part prior to evaluating them against the inaccessible test part. Although
both subtasks seem not that different to approach, e.g., by computing similar stylometric
fingerprints, results indicate that intrinsically segmenting a text into distinct authorial
components is hard to be tackled. On the other hand, the gap for building clusters of
already segmented texts could be narrowed, in large parts due to the outcomes of similar
studies conducted in previous PAN events.
        </p>
        <p>For the style breach detection subtask, three approaches have been submitted,
utilizing common stylometric features and word dictionaries in combination with different
distance metrics, or by applying a neural network similar to the word embeddings
technique. Although all approaches achieved a better performance than the simple random
baseline, only one of them could exceed a slightly enhanced baseline, which is also
based on random guesses. Interestingly, this winning approach considers only the ends
of preformatted paragraphs as possible segment borders, and, if no paragraphs exist,
creates artificial paragraphs of predefined, fixed lengths. This fact underlines that there
is still room for significant improvements, e.g., by dynamically adjusting the borders.
Moreover, another approach basically used an intrinsic plagiarism detection method,
which aims at outlier detection over the whole document, marking them as borders.
Clearly, tackling the style breach detection task with this method is not optimal since the
intrinsic plagiarism detection algorithms assume a main author to be existent. Finally,
non of the approaches utilized standard machine learning techniques such as support
vector machines. Considering the findings of recent research using such techniques on
textual data, it can be assumed that—if optimized and used accordingly—the
performance of style breach detection algorithms can be improved significantly.</p>
        <p>
          For the author clustering task, in comparison to the evaluation results of the
corresponding task at PAN-2016, the submitted methods achieved lower Bcubed F-score.
However, this should not be explained by the fact that text-length in PAN-2017 datasets
is much lower. A more crucial factor is the much lower range of the clusteriness ratio
r which limits the number of single-item clusters and significantly increases the
number of true authorship-links. As a result, the performance of naive baseline methods,
like BASELINE-Singleton, is not so competitive as in the corresponding task at
PAN2016. Moreover, MAP scores are considerably increased in comparison to PAN-2016.
Given that the MAP scores of BASELINE-PAN16 are also improved with respect to its
performance on PAN-2016 datasets, this can be largely explained by the low values of
clusteriness ratio again. The success of the top-performing submission shows that very
good results can be obtained by using well-known clustering methods and similarity
functions given that a suitable feature set and feature weighting scheme is selected for
each dataset [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
[56] Tschuggnall, M., Specht, G.: Automatic decomposition of multi-author documents using
grammar analysis. In: Proceedings of the 26th GI-Workshop on Grundlagen von
Datenbanken. CEUR-WS, Bozen, Italy (October 2014)
[57] Zheng, R., Li, J., Chen, H., Huang, Z.: A framework for authorship identification of online
messages: Writing-style features and classification techniques. Journal of the American
Society for Information Science and Technology 57(3), 378–393 (2006)
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Alberts</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Author clustering with the aid of a simple distance measure</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          , T. (eds.)
          <source>CLEF 2017 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Aldebei</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Unsupervised multi-author document decomposition based on hidden markov model</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          , Volume
          <volume>1</volume>
          :
          <string-name>
            <given-names>Long</given-names>
            <surname>Papers</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Amigó</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Artiles</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A comparison of extrinsic clustering evaluation metrics based on formal constraints</article-title>
          .
          <source>Information Retrieval</source>
          <volume>12</volume>
          (
          <issue>4</issue>
          ),
          <fpage>461</fpage>
          -
          <lpage>486</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Bagnall</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Authorship Clustering Using Multi-headed Recurrent Neural Networks</article-title>
          .
          <source>In: CLEF 2016 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>Topic segmentation with an aspect hidden markov model</article-title>
          .
          <source>In: Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . pp.
          <fpage>343</fpage>
          -
          <lpage>348</lpage>
          . SIGIR '01,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2001</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/383952.384021
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>F.Y.</given-names>
          </string-name>
          :
          <article-title>Advances in Domain Independent Linear Text Segmentation</article-title>
          . In:
          <article-title>Proceedings of the 1st North American chapter of the Association for Computational Linguistics conference</article-title>
          . pp.
          <fpage>26</fpage>
          -
          <lpage>33</lpage>
          . Association for Computational Linguistics (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Clarke</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Craswell</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soboroff</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.:</given-names>
          </string-name>
          <article-title>Overview of the TREC 2009 web track</article-title>
          .
          <source>Tech. rep., DTIC Document</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Daks</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Unsupervised authorial clustering based on syntactic structure</article-title>
          .
          <source>In: Proceedings of the ACL 2016 Student Research Workshop</source>
          . pp.
          <fpage>114</fpage>
          -
          <lpage>118</lpage>
          . Association for Computational Linguistics (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Karas</surname>
          </string-name>
          ´,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Sobecki</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>OPI-JSA at CLEF 2017: Author Clustering and Style Breach Detection</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2017 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Eisenstein</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barzilay</surname>
          </string-name>
          , R.:
          <article-title>Bayesian Unsupervised Topic Segmentation</article-title>
          .
          <source>In: Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <fpage>334</fpage>
          -
          <lpage>343</lpage>
          . EMNLP '
          <volume>08</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>García</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castro</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavielle</surname>
          </string-name>
          , V., noz, R.M.:
          <article-title>Discovering Author Groups Using a -compact Graph-based Clustering</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          , T. (eds.)
          <source>CLEF 2017 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Giannella</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>An improved algorithm for unsupervised decomposition of a multi-author document</article-title>
          .
          <source>JASIST</source>
          <volume>67</volume>
          (
          <issue>2</issue>
          ),
          <fpage>400</fpage>
          -
          <lpage>411</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Gibbons</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chakraborti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Nonparametric Statistical Inference, pp.
          <fpage>977</fpage>
          -
          <lpage>979</lpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Glavaš</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nanni</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          :
          <article-title>Unsupervised text segmentation using semantic relatedness graphs</article-title>
          .
          <source>Association for Computational Linguistics</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Glover</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirst</surname>
          </string-name>
          , G.:
          <article-title>Detecting stylistic inconsistencies in collaborative writing</article-title>
          .
          <source>In: The New Writing Environment</source>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>168</lpage>
          . Springer (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burrows</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Ousting Ivory Tower Research:
          <article-title>Towards a Web Framework for Providing Experiments as a Service</article-title>
          . In: Hersh,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Callan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Maarek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Sanderson</surname>
          </string-name>
          , M. (eds.) 35th
          <source>International ACM Conference on Research and Development in Information Retrieval (SIGIR 12)</source>
          . pp.
          <fpage>1125</fpage>
          -
          <lpage>1126</lpage>
          . ACM (Aug
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Gómez-Adorno</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aleman</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , no,
          <string-name>
            <given-names>D.V.</given-names>
            ,
            <surname>Sanchez-Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.A.</given-names>
            ,
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Sidorov</surname>
          </string-name>
          , G.:
          <article-title>Author Clustering using Hierarchical Clustering Analysis</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          , T. (eds.)
          <source>CLEF 2017 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Graham</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirst</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marthi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Segmenting documents by stylistic character</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>11</volume>
          (
          <issue>04</issue>
          ),
          <fpage>397</fpage>
          -
          <lpage>415</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Halvani</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graner</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Author Clustering based on Compression-based Dissimilarity Scores</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          , T. (eds.)
          <source>CLEF 2017 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Texttiling: Segmenting text into multi-paragraph subtopic passages</article-title>
          .
          <source>Computational linguistics 23(1)</source>
          ,
          <fpage>33</fpage>
          -
          <lpage>64</lpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>D.I.</given-names>
          </string-name>
          :
          <article-title>The evolution of stylometry in humanities scholarship</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          <volume>13</volume>
          (
          <issue>3</issue>
          ),
          <fpage>111</fpage>
          -
          <lpage>117</lpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Iqbal</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Binsalleeh</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fung</surname>
            ,
            <given-names>B.C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Debbabi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Mining writeprints from anonymous e-mails for forensic investigation</article-title>
          .
          <source>Digital Investigation</source>
          <volume>7</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>56</fpage>
          -
          <lpage>64</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Karas</surname>
            <given-names>´</given-names>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
            , S´ piewak,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sobecki</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>OPI-JSA at CLEF 2017: Author Clustering and Style Breach Detection</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          , T. (eds.)
          <source>CLEF 2017 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Kestemont</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luyckx</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Intrinsic Plagiarism Detection Using Character Trigram Distance Scores</article-title>
          .
          <source>In: Notebook Papers of the 5th Evaluation Lab on Uncovering Plagiarism, Authorship and Social Software Misuse (PAN)</source>
          . Amsterdam, The Netherlands (
          <year>September 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Style Breach Detection: An Unsupervised Detection Model</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2017 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Kiros</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakhutdinov</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zemel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urtasun</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torralba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fidler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Skip-thought vectors</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <fpage>3294</fpage>
          -
          <lpage>3302</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Kocher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : UniNE at CLEF 2016:
          <article-title>Author Clustering</article-title>
          .
          <source>In: CLEF 2016 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Kocher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savoy</surname>
          </string-name>
          , J.: UniNE at CLEF 2017:
          <article-title>Author Clustering</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          , T. (eds.)
          <source>CLEF 2017 Working Notes. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akiva</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dershowitz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dershowitz</surname>
          </string-name>
          , N.:
          <article-title>Unsupervised decomposition of a document into authorial components</article-title>
          . In: Lin,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Matsumoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Mihalcea</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.)
          <article-title>Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics</article-title>
          . pp.
          <fpage>1356</fpage>
          -
          <lpage>1364</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Computational methods in authorship attribution</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          <volume>60</volume>
          (
          <issue>1</issue>
          ),
          <fpage>9</fpage>
          -
          <lpage>26</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motrenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuznetsova</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strijov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Methods for Intrinsic Plagiarism Detection and Author Diarization</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2016 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Layton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watters</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dazeley</surname>
          </string-name>
          , R.:
          <article-title>Automated unsupervised authorship analysis using evidence accumulation clustering</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>19</volume>
          ,
          <fpage>95</fpage>
          -
          <lpage>120</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Layton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watters</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dazeley</surname>
          </string-name>
          , R.:
          <article-title>Evaluating authorship distance methods using the positive silhouette coefficient</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>19</volume>
          ,
          <fpage>517</fpage>
          -
          <lpage>535</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Misra</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yvon</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jose</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cappe</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Text segmentation via topic modeling: an analytical study</article-title>
          .
          <source>In: Proceedings of the 18th ACM conference on Information and knowledge management</source>
          . pp.
          <fpage>1553</fpage>
          -
          <lpage>1556</lpage>
          . ACM (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Pevzner</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>A critique and improvement of an evaluation metric for text segmentation</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>28</volume>
          (
          <issue>1</issue>
          ),
          <fpage>19</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Ponte</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croft</surname>
          </string-name>
          , W.B.:
          <article-title>Text Segmentation by Topic</article-title>
          .
          <source>In: Research and Advanced Technology for Digital Libraries</source>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>125</lpage>
          . Springer (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eiselt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrón-Cedeño</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 3rd International Competition on Plagiarism Detection</article-title>
          .
          <source>In: Notebook Papers of the 5th Evaluation Lab on Uncovering Plagiarism, Authorship and Social Software Misuse (PAN)</source>
          . Amsterdam, The Netherlands (
          <year>September 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Improving the Reproducibility of PAN's Shared Tasks: Plagiarism Detection, Author Identification, and Author Profiling</article-title>
          . In: Kanoulas,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Hanbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Toms</surname>
          </string-name>
          , E. (eds.)
          <article-title>Information Access Evaluation meets Multilinguality, Multimodality, and Visualization</article-title>
          .
          <source>5th International Conference of the CLEF Initiative (CLEF 14)</source>
          . pp.
          <fpage>268</fpage>
          -
          <lpage>299</lpage>
          . Springer, Berlin Heidelberg New York (
          <year>Sep 2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Völske</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Crowdsourcing Interaction Logs to Understand Text Reuse from the Web</article-title>
          . In: Fung,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Poesio</surname>
          </string-name>
          , M. (eds.)
          <article-title>Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (ACL 13)</article-title>
          . pp.
          <fpage>1212</fpage>
          -
          <lpage>1221</lpage>
          . Association for Computational Linguistics (
          <year>Aug 2013</year>
          ), http://www.aclweb.org/anthology/P13-1119
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>Rangel</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Verhoeven</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Daelemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Overview of the 4th Author Profiling Task at PAN 2016: Cross-Genre Evaluations</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2016 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <surname>Reynar</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Statistical Models for Topic Segmentation</article-title>
          .
          <source>In: Proc. of the 37th annual meeting of the Association for Computational Linguistics on Computational Linguistics</source>
          . pp.
          <fpage>357</fpage>
          -
          <lpage>364</lpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <surname>Riedl</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biemann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Topictiling: a text segmentation algorithm based on lda</article-title>
          .
          <source>In: Proceedings of ACL 2012 Student Research Workshop</source>
          . pp.
          <fpage>37</fpage>
          -
          <lpage>42</lpage>
          . Association for Computational Linguistics (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <surname>Safin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuznetsova</surname>
          </string-name>
          , R.:
          <article-title>Style Breach Detection with Neural Sentence Embeddings</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2017 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <surname>Samdani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>K.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A discriminative latent variable model for online clustering</article-title>
          .
          <source>In: Proceedings of The 31st International Conference on Machine Learning</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <surname>Scaiano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inkpen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Getting more from segmentation evaluation. In: Proceedings of the 2012 conference of the north american chapter of the association for computational linguistics: Human language technologies</article-title>
          . pp.
          <fpage>362</fpage>
          -
          <lpage>366</lpage>
          . Association for Computational Linguistics (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <surname>Sittar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iqbal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawab</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Author Diarization Using Cluster-Distance Approach</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2016 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>A Survey of Modern Authorship Attribution Methods</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          <volume>60</volume>
          ,
          <fpage>538</fpage>
          -
          <lpage>556</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [48]
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Intrinsic Plagiarism Detection Using Character n-gram Profiles</article-title>
          .
          <source>In: Notebook Papers of the 5th Evaluation Lab on Uncovering Plagiarism, Authorship and Social Software Misuse (PAN)</source>
          . Amsterdam, The Netherlands (
          <year>September 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [49]
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Clustering by Authorship Within and Across Documents</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2016 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2016</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1609</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [50]
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Clustering by Authorship Within and Across Documents</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2016 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org (Sep</source>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [51]
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipka</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Intrinsic plagiarism analysis</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>45</volume>
          (
          <issue>1</issue>
          ),
          <fpage>63</fpage>
          -
          <lpage>82</lpage>
          (
          <year>Mar 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [52]
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipka</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Intrinsic plagiarism analysis</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>45</volume>
          (
          <issue>1</issue>
          ),
          <fpage>63</fpage>
          -
          <lpage>82</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [53]
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
          </string-name>
          , G.:
          <article-title>Plag-Inn: Intrinsic Plagiarism Detection Using Grammar Trees</article-title>
          .
          <source>In: Proceedings of the 18th International Conference on Applications of Natural Language to Information Systems (NLDB)</source>
          . pp.
          <fpage>284</fpage>
          -
          <lpage>289</lpage>
          . Springer, Groningen, The Netherlands (
          <year>June 2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          [54]
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
          </string-name>
          , G.:
          <article-title>Countering Plagiarism by Exposing Irregularities in Authors' Grammar</article-title>
          .
          <source>In: Proceedings of the European Intelligence and Security Informatics Conference (EISIC)</source>
          . pp.
          <fpage>15</fpage>
          -
          <lpage>22</lpage>
          . IEEE, Uppsala, Sweden (
          <year>August 2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          [55]
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Using Grammar-Profiles to Intrinsically Expose Plagiarism in Text Documents</article-title>
          .
          <source>In: Proc. of the 18th Conf. of Natural Language Processing and Information Systems (NLDB)</source>
          . pp.
          <fpage>297</fpage>
          -
          <lpage>302</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>