<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the Cross-Domain Authorship Attribution Task at PAN 2019</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mike Kestemont</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Efstathios Stamatatos</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrique Manjavacas</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Walter Daelemans</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Potthast</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benno Stein</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bauhaus-Universität Weimar</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Leipzig University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Antwerp</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of the Aegean</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Authorship identification remains a highly topical research problem in computational text analysis, with many relevant applications in contemporary society and industry. In this edition of PAN, we focus on authorship attribution, where the task is to attribute an unknown text to a previously seen candidate author. Like in the previous edition we continue to work with fanfiction texts (in four Indo-European languages), written by non-professional authors in a crossdomain setting: the unknown texts belong to a different domain than the training material that is available for the candidate authors. An important novelty of this year's setup is the focus on open-set attribution, meaning that the test texts contain writing samples by previously unseen authors. For these, systems must consequently refrain from an attribution. We received altogether 12 submissions for this task, which we critically assess in this paper. We provide a detailed comparison of these approaches, including three generic baselines.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Authorship attribution [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1,2,3</xref>
        ] continues to be an important problem in information
retrieval and computational linguistics, and also in applied areas such as law and
journalism where knowing the author of a document (such as a ransom note) may enable
law enforcement to save lives. The most common framework for testing candidate
algorithms is the closed-set attribution task: given a sample of reference documents from
a finite set of candidate authors, the task is to determine the most likely author of a
previously unseen document of unknown authorship. This task is quite challenging under
cross-domain conditions where documents of known and unknown authorship come
from different domains (such as a different thematic area or genre). In addition, it is
often more realistic to assume that the true author of a disputed document is not
necessarily included in the list of candidates [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        This year, we again focus on the attribution task in the context of transformative
literature, more colloquially know as ‘fanfiction’. Fanfiction refers to a rapidly
expanding body of fictional narratives, typically produced by non-professional authors who
self-identify as ‘fans’ of a particular oeuvre or individual work [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Usually, these
stories (or ‘fics’) are openly shared online with a larger fan community on platforms such
as fanfiction.net or archiveofourown.org. Interestingly, fanfiction is the
fastest growing form of writing in the world nowadays [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. When sharing their texts,
fanfiction writers explicitly acknowledge taking inspiration from one or more cultural
domains that are known as ‘fandoms’. The resulting borrowings take place on various
levels, such as themes, settings, characters, story world, and also style. Fanfiction
usually is of an unofficial and unauthorized nature [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], but because most fanfiction writers
do not have any commercial purposes, the genre falls under the principle of ‘Fair Use’
in many countries [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. From the perspective of writing style, fanfiction offers valuable
benchmark data: the writings are unmediated and unedited before publication, which
means that they should accurately reflect an individual author’s writing style.
Moreover, the rich metadata available for individual fics presents opportunities to quantify
the extent to which fanfiction writers have modeled their writing style after the original
author’s style [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>In the previous edition of PAN we also dealt with authorship attribution in fanfiction
and added extra difficulty with a cross-domain condition (i.e., different fandoms). This
year we have further increased the difficulty of the task by focusing on open-set
attribution conditions, meaning that the true author of a test text is not necessarily included
in the list of candidate authors. More formally, an open cross-domain authorship
attribution problem can be expressed as a tuple (A; K; U ), with A as the set of candidate
authors, K as the set of reference (known authorship) texts, and U as the set of unknown
authorship texts. For each candidate author a 2 A, we are given Ka K, a set of texts
unquestionably written by author a. Each text in U should either be assigned to exactly
one a 2 A, or the system should refrain from an attribution if the target author is not
supposed to be in A. From a text categorization point of view, K is the training corpus
and U is the test corpus. Let DK be (the names of) the fandoms of the texts in K. Then,
all texts in U belong to a single fandom dU , with dU 2= DK .
2</p>
    </sec>
    <sec id="sec-2">
      <title>Datasets</title>
      <p>This year’s shared tasks use datasets in four major Indo-European languages: English,
French, Italian, and Spanish. For each language, 10 “problems” were constructed on
the basis of a larger dataset obtained from archiveofourown.org in 2017. Per
language, five problems were released as a development set to the participants in order
to calibrate their systems. The final evaluation of the submitted systems was carried
out on the five remaining problems that were not publicly released before the final
results were communicated. Each problem had to be solved independently from the other
problems. It should be noted that the development material could not be used as mere
training material for supervised learning approaches since the candidate authors of the
development corpus and the evaluation corpus do not overlap. Therefore, approaches
should not be designed to particularly handle the candidate authors of the development
corpus but should focus on their generalizability to other author sets.</p>
      <p>One problem corresponds to a single open-set attribution task, where we distinguish
between the “source” and the “target” material. The source material in each problem
contains exactly 7 training texts for exactly 9 candidate authors. In the target material,
these 9 authors are represented by at least one test text (possibly more). Additionally,
the target material also contains so-called “adversaries”, which were not written by one
of the candidate authors (indicated by the author label “&lt;UNK&gt;”). The proportion of
the number of target texts written by the candidate authors in problems, as opposed
to &lt;UNK&gt; documents, was varied across the problems in the development dataset, in
order to discourage systems from opportunistic guessing.</p>
      <p>Let UK be the subset of U that includes all test documents actually written by the
candidate authors, and let UU be the subset of U containing the rest of the test
documents not written by any candidate author. Then, the adversary ratio r = jUU j=jUK j
determines the likelihood of a test document to belong to one of the candidates. If r = 0
or close to 0), the problem is essentially a closed-set attribution scenario since all test
documents belong to the candidate authors, or very few are actually written by
adversaries. If r = 1, then it is equally probable for a test document to be written by a
candidate author or by another author. For r &gt; 1 it is more likely for a test document to
be written by an adversary not included in the list of candidates.</p>
      <p>We examine cases where r ranges from 0.2 to 1.0. In greater detail, as can be seen
in Table 1, the development dataset comprises 5 problems per language that correspond
to r = [0:2; 0:4; 0:6; 0:8; 1:0]. This dataset was released for the participants in order
to develop and calibrate their submissions. The final evaluation dataset also includes
5 problems per language but with fixed r = 1. Thus, the participants are implicitly
encouraged to develop generic approaches, because of the varying likelihood that a
test document is written by a candidate or an adversary. In addition, it is possible to
estimate the effectiveness of submitted methods when r &lt; 1 by ignoring their answers
for specific subsets of UU in the evaluation dataset.</p>
      <p>
        Each of the individual texts belongs to a single fandom, i.e., a certain topical
domain. Fandoms were made available in the training material so that systems could
exploit this information, as done by Seroussi et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for instance. We only selected
works counting at least 500 tokens (according to the original database’s internal token
count), which is already a challenging text length for authorship analyses. Finally, we
normalized document length: for fics longer than 1 000 tokens, we only included the
middle 1 000 tokens of the text.
      </p>
      <p>
        Another novelty this year was the inclusion of a set of 5 000 problem-external
documents per language written by “imposter” authors (the authorship of these texts is
also encoded as &lt;UNK&gt;). These documents could be freely used by the participants
to develop their systems, for instance in the popular framework of imposter
verification [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The release of this additional corpus should be understood against the
backdrop of PAN’s benchmarking efforts to improve the reproducibility and comparability
of approaches. Many papers using imposter-based approaches do so on ad-hoc collected
corpora, making it hard to compare the effect of the composition of the imposter
collection. The imposter collection is given for the language as a whole and is thus not
problem-specific. We provide the information that these texts were not written by
authors who appear in the source or target sets for the problems in the language. When
selecting these texts from the base dataset, we have given preference to texts from the
fandoms covered in the problems, but when this selection was smaller than 5 000 texts,
we have completed it with a random selection of other texts.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Framework</title>
      <p>In the final evaluation phase, submitted systems were presented with 5 problems per
language: for each problem, given a set of documents (known fanfics) by candidate
authors, the systems had to identify the authors of another set of documents (unknown
fanfics) in a previously unencountered fandom (target domain). Systems could assume
that each candidate author had contributed at least one of the unknown fanfics to the
problem, which all belonged to the same target fandom. Some of the fanfics in the
target domain, however, were not written by any of the candidate authors. Like in the
calibration set, the known fanfics belonged to several fandoms (excluding the target
fandom), although not necessarily the same for all candidate authors. An equal number
of known fanfics per candidate author was provided: 7 fanfics for 9 authors. By contrast,
the unknown fanfics were not equally distributed over the authors.</p>
      <p>
        The submissions were separately evaluated in each attribution problem, based on
their open-set macro-averaged F1 score (calculated over the training classes, i.e., when
&lt;UNK&gt; is excluded) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Participants were ranked according to their average open-set
macro-F1 across all attribution problems of the evaluation corpus. A reference
implementation of the evaluation script was made available to the participants.
3.1
      </p>
      <p>
        Baseline Methods
As usual, we provide the implementation of three baseline methods that provide an
estimation of the overall difficulty of the problem given the state of the art in the field.
These implementations are in Python (2.7+) and rely on Scikit-learn and its base
packages [
        <xref ref-type="bibr" rid="ref12 ref13">12,13</xref>
        ] as well as NLTK [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Participants were free to base their approach on one
of these reference systems, or to develop their own approach from scratch. The provided
baseline are as follows:
1. BASELINE-SVM. A language-independent authorship attribution approach, framing
attribution as a conventional text classification problem [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. It is based on a
character 3-gram representation and a linear SVM classifier with a reject option. First,
it estimates the probabilities of output classes based on Platt’s method [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Then,
it assigns an unknown document to the &lt;UNK&gt; class when the difference of the
probabilities of the top two candidates is less than a predefined threshold. Let a1
and a2, a1; a2 2 A, be the two most likely authors of a certain test document while
P r1 and P r2 are the corresponding estimated probabilities (i.e., all other
candidates obtained lower probabilities). Then, if P r1 P r2 &lt; 0:1, the document is left
unattributed. Otherwise it is attributed to a1.
2. BASELINE-COMPRESSOR. A language-independent approach that uses text
compression to estimate the distance of an unknown document to each of the candidate
authors. This approach was originally proposed by [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and was later reproduced
by [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. It uses the Prediction by Partial Matching (PPM) compression algorithm
to build a model for each candidate author. Then, it calculates the cross-entropy
of each test document with respect to the model of each candidate and assigns the
document to the author with the lowest score. In order to adapt this method to the
open-set classification scenario, we introduced a reject option. In more detail, a
test document is left unattributed when the difference between the two most likely
candidates is lower than a predefined threshold. Let a1 and a2, a1; a2 2 A, be the
two most likely candidate authors for a certain test document while S1 and S2 are
their cross-entropy scores (i.e., all other candidate authors have higher scores). If
(S1 S2)=S1 &lt; 0:01, then the test document is left unattributed. Otherwise, it is
assigned to a1.
3. BASELINE-IMPOSTERS. Implementation of the language-independent imposters
approach for authorship verification [
        <xref ref-type="bibr" rid="ref19 ref4">4,19</xref>
        ], based on character tetragram features.
During a bootstrapped procedure, the technique iteratively compares an unknown
text to each candidate author’s training profile, as well as to a set of imposter
documents, on the basis of a randomly selected feature subset. Then, the number of
times the unknown document is found more similar to the candidate author’s
documents rather than to the imposters indicates how likely it is for that candidate to be
the true author of the document. Instead of performing this procedure separately for
each candidate author, we examine all candidate authors within each iteration (i.e.,
in each iteration, a maximum of one candidate author’s score is increased). If after
this repetitive process the highest score (corresponding to the most likely author)
does not pass a fixed similarity threshold (here: 10% of repetitions), the document
is assigned to the &lt;UNK&gt; class and is left unattributed. This baseline method is the
only one that uses additional, problem-external imposter documents. We provided a
collection of 5 000 imposter documents (fanfics on several fandoms) per language.
      </p>
      <p>Finally, we also compare the participating systems to a plain “majority” baseline:
through a simple voting procedure with random tie breaking, this baseline accepts a
candidate for a given unseen text if the majority of submitted methods agree on it;
otherwise, the &lt;UNK&gt; label is predicted. No meta-learning is applied to weigh the
importance of the votes of individual systems.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Survey of Submissions</title>
      <p>
        In total, 12 methods were submitted to the task and evaluated using the TIRA
experimentation framework. All but one (Kipnis) of the submissions are described in the
participants’ notebook papers. Table 5 presents an overview of the main
characteristics of the submitted methods as well as the baselines. We also record whether
approaches made use of the language-specific imposter material or language-specific NLP
resources, such as pretrained taggers and parsers. As can be seen, there is surprisingly
little variance in the approaches. The majority of submissions follow the paradigm of
the BASELINE-SVM or the winner approach [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] of PAN-2018 cross-domain
authorship attribution task [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], which is an ensemble of classifiers each of which based on a
different text modality.
      </p>
      <p>
        Compared to the baselines, most submitted methods attempt to exploit richer
information that corresponds to different text modalities as well as variable-length
ngrams [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] in contrast to fixed-length n-grams [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. The most popular features are
n-grams extracted from plain text modalities, such as character, word, token,
part-ofspeech tag, or syntactic level sequences. Given the cross-domain conditions of the task,
several participants attempted to use more abstract forms of textual information such
as punctuation sequences [
        <xref ref-type="bibr" rid="ref22 ref24 ref25">24,22,25</xref>
        ] or n-grams extracted from distorted versions [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]
of the original documents [
        <xref ref-type="bibr" rid="ref22 ref25 ref27">27,22,25</xref>
        ]. There is limited effort to enrich n-gram features
with alternative stylometric measures like word and sentence length distributions [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]
or features related to syntactic analysis of documents [
        <xref ref-type="bibr" rid="ref24 ref29">24,29</xref>
        ]. Only one participant
used word embeddings [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. Other participants report that they were discouraged to use
more demanding types of word and sentence embeddings due to hardware limitations
of TIRA [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ], which points to important infrastructural needs that may be addressed in
future editions. Those same teams, however, informally reported that the difference in
performance (in the development dataset) when using such additional features is
negligible.
      </p>
      <p>
        With respect to feature weighting, tf-idf is the most popular option while the
baseline methods are based on the simpler tf scheme. There is one attempt to use both of
these schemes [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. A quite different approach uses a normalization scheme based on
zscores [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. In addition, a few methods apply dimension reduction (PCA, SVD) to the
features [
        <xref ref-type="bibr" rid="ref22 ref32 ref33">32,33,22</xref>
        ]. Judging from the results, such methods for dimension reduction
have the potential to boost performance.
      </p>
      <p>
        As concerns the classifiers, the most popular choices are SVMs and ensembles of
classifiers, usually exploiting SVM base models followed by Logistic Regression (LR)
models. In a few cases, the participants informally report that they have experimented
with alternative classification algorithms (random forests, k-nn, naive Bayes) and found
that SVM and LR are the most effective classifiers for this kind of task [
        <xref ref-type="bibr" rid="ref27 ref30">27,30</xref>
        ]. None
of the participant’s methods is based on deep learning algorithms, most probably due
to hardware limitations of TIRA or because of the discouraging reported results in the
corresponding task of PAN-2018 [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        Given the fact that the focus of PAN-2019 edition of the task is on open-set
attribution, it can be noted that none of the participants attempted to build a pure open-set
classifier [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ]. By contrast, they just use closed-set classifiers with a reject option (the
classification prediction is dropped when the confidence of prediction is low), similar
to the baseline methods [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ].
      </p>
      <p>A crucial issue to improve the performance of authorship attribution is the
appropriate tuning of hyperparameters. Most of the participants tune the hyperparameters of
their approach globally based on the development dataset, that is, they estimate the most
suitable parameter values that are applied to any problem of the test dataset. In contrast,
a few participants attempt to tune the parameters of their method in a language-specific
way, estimating the most suitable values for each language separately. None of the
submitted methods attempts to tune parameter values for each individual attribution
problem.</p>
      <p>
        The submission of van Halteren focuses on the cross-domain difficulty of the task
and attempts to exploit the availability of multiple texts of unknown authorship in the
target domain within each attribution problem [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. This submission performs a
sophisticated strategy composed of different phases. Initially, a typical cross-domain classifier
is built and each unknown document is assigned to its most likely candidate author but
the prediction is kept only for the most confident cases. Then, a new in-domain
classifier is built using the target domain documents (for which the predictions were kept
in the previous phase) and the remaining target domain documents are classified
accordingly. However, this in-domain classifier can only be useful for certain candidate
authors, the ones with enough confident predictions in the initial phase. A final phase
combines the results of cross-domain and in-domain classifiers and leaves documents
with less confident predictions unattributed.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation Results</title>
      <p>
        to be the most difficult case. It is also remarkable that the baseline-compressor method
achieves the best baseline results for English, French, and Italian, but it is not as
competitive in Spanish. Furthermore, note that Muttenthaler et al.’s submission is the only
one to outperform the otherwise very competitive majority baseline, albeit by a very
small margin. The latter reaches a relatively high precision, but must sacrifice quite a
bit of recall in return. That the winner outperforms the majority baseline is surprising:
in previous editions of this shared task (e.g. [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]), similar meta-level approaches proved
very hard to beat. This result is probably an artifact of the lack of diversity among the
submissions in the top-scoring cohort, which seem to have produced very similar
predictions (see below), thus reducing the beneficial effects of a majority vote among those
systems.
      </p>
      <p>In order to examine the effectiveness of the submitted methods for a varying
adversary ratio, we performed the following additional evaluation process. As can be seen in
Table 1, all attribution problems of the evaluation dataset have a fixed adversary ratio
r = 1, meaning that an equal number of documents written either by the candidate
authors or adversary authors is included in the test set of each problem. Once the submitted
methods processed the whole evaluation dataset, we calculated the evaluation measures
with decreasing proportions of adversary documents at 100%, 80%, 60%, 40%, and
20%, resulting in an adversary ratio that ranges from 1 to 0.2. Table 4 presents the
evaluation results (averaged macro-F1) for such a varying adversary ratio. In general,
the performance of all methods increases when the adversary ratio decreases. Recall
that r = 0 corresponds to a closed-set attribution case. The performance of the two
top-performing approaches is very similar in the whole range of examined r-values.
However, the method of Muttenthaler et al. is slightly better for high r-values while
Bacciu et al. is slightly better for low r-values.</p>
      <p>Moreover, we have applied statistical significance tests to the systems’ output.
Especially since many systems have adopted a similar approach, it is worthwhile to discuss
the extent to which submissions show statistically meaningful differences. Like in
previous editions, we have applied approximate randomization testing, a non-parametric
procedure that accounts for the fact that we should not make too many assumptions as
to the underlying distributions for the classification labels. Table 6 lists the results for
pairwise tests, comparing all submitted approaches to each other, based on their
respective F1-scores for all labels in the problems. For 1 000 bootstrapped iterations, the test
returns probabilities which we can interpret as the conventional p-values of one-sided,
statistical tests—i.e., the probability of failing to reject the null hypothesis (H0) that
the classifiers do not output significantly different scores. The symbolic notation takes
into account the following straightforward thresholds: ‘=’ (not significantly different:
p &gt; 0:5), ‘*’ (significantly different: p &lt; 0:05), ‘**’ (very significantly different:
p &lt; 0:01), ‘***’ (highly significantly different: p &lt; 0:001). Interestingly, systems with
neighboring ranks often do not yield significantly different scores; this is also true for
the two top-performing systems. Note that almost all systems have produced an output
that is significantly different from the three baselines (which also display a high degree
of difference among one another). According to this test, the difference between
Muttenthaler et al. and Bacciu et al. is not statistically significant, although the former is
significantly different from the majority baseline.</p>
      <p>B e n
R in il
l e
e s
s a
a b
.la ity .la ion ise .l r n le n .la laa</p>
      <p>a te o i e m sro sr
lttreeah jraoM iteccauB raabP&amp; serdV&amp; ítezergu iIssb ssaoJnh saB lteraanH teyoguh agG il-sseevan -srecopm it-seoopm</p>
      <p>V a</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>
        The paper discussed the 12 submissions to the 2019 edition of the PAN shared task on
authorship identification. Like last year, we focused on cross-domain attribution in
fanfiction data. An important innovation this year was the focus on the open-set attribution
set-up, where participating systems had to be able to refrain from attributing unseen
texts as well. The analyses described above call for a number of considerations that are
not without relevance to future development in the field of computational authorship
identification. First of all, this year’s edition was characterized by a relative low
degree of diversity in approaches: especially the higher-scoring cohort almost exclusively
adopted a highly similar approach, involving a combination of SVMs as classifier
(potentially as part of an ensembles), character n-grams as features, and a rather simple
thresholding mechanism to refrain from attributions. It is not immediately clear which
directions future research might explore. Deep learning-based methods, which can be
pretrained on external corpora, have so far not led to a major breakthrough in the field,
despite the impressive improvements which have been reported for these methods in
other areas of NLP. Also, a more promising research direction might be to move away
from closed-set classifiers (with a naive reject-option), towards purely open-set
classifiers [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ]
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Koppel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Argamon</surname>
          </string-name>
          .
          <article-title>Computational methods in authorship attribution</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          ,
          <volume>60</volume>
          (
          <issue>1</issue>
          ):
          <fpage>9</fpage>
          -
          <lpage>26</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>P.</given-names>
            <surname>Juola</surname>
          </string-name>
          .
          <article-title>Authorship attribution</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          ,
          <volume>1</volume>
          (
          <issue>3</issue>
          ):
          <fpage>233</fpage>
          -
          <lpage>334</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          .
          <article-title>A survey of modern authorship attribution methods</article-title>
          .
          <source>JASIST</source>
          ,
          <volume>60</volume>
          (
          <issue>3</issue>
          ):
          <fpage>538</fpage>
          -
          <lpage>556</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Moshe</given-names>
            <surname>Koppel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yaron</given-names>
            <surname>Winter</surname>
          </string-name>
          .
          <article-title>Determining if two documents are written by the same author</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          ,
          <volume>65</volume>
          (
          <issue>1</issue>
          ):
          <fpage>178</fpage>
          -
          <lpage>187</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Karen</given-names>
            <surname>Hellekson</surname>
          </string-name>
          and Kristina Busse, editors.
          <source>The Fan Fiction Studies Reader</source>
          . University of Iowa Press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>K.</given-names>
            <surname>Mirmohamadi</surname>
          </string-name>
          .
          <source>The Digital Afterlives of Jane Austen. Janeites at the Keyboard. Palgrave MacMillan</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J.</given-names>
            <surname>Fathallah</surname>
          </string-name>
          .
          <article-title>Fanfiction and the Author</article-title>
          .
          <source>How FanFic Changes Popular Cultural Texts</source>
          . Amsterdam University Press,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>R.</given-names>
            <surname>Tushnet</surname>
          </string-name>
          .
          <article-title>Legal fictions: Copyright, fan fiction, and a new common law</article-title>
          . Loyola of Los Angeles Entertainment Law Review,
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          , Francisco M. Rangel Pardo, Michael Tschuggnall, Benno Stein, Mike Kestemont, Paolo Rosso, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Potthast</surname>
          </string-name>
          .
          <source>Overview of PAN</source>
          <year>2018</year>
          <article-title>- author identification, author profiling, and author obfuscation</article-title>
          .
          <source>In Experimental IR Meets Multilinguality</source>
          , Multimodality, and Interaction - 9th
          <source>International Conference of the CLEF Association, CLEF</source>
          <year>2018</year>
          , Avignon, France,
          <source>September 10-14</source>
          ,
          <year>2018</year>
          , Proceedings, pages
          <fpage>267</fpage>
          -
          <lpage>285</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Yanir</surname>
            <given-names>Seroussi</given-names>
          </string-name>
          , Ingrid Zukerman, and
          <string-name>
            <given-names>Fabian</given-names>
            <surname>Bohnert</surname>
          </string-name>
          .
          <article-title>Authorship attribution with topic models</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>40</volume>
          (
          <issue>2</issue>
          ):
          <fpage>269</fpage>
          -
          <lpage>310</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pedro R. Mendes</surname>
            <given-names>Júnior</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberto M. de Souza</surname>
            , Rafael de
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Werneck</surname>
          </string-name>
          , Bernardo V. Stein, Daniel V. Pazinato, Waldir R. de Almeida,
          <string-name>
            <surname>Otávio A. B. Penatti</surname>
            , Ricardo da
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Torres</surname>
            , and
            <given-names>Anderson</given-names>
          </string-name>
          <string-name>
            <surname>Rocha</surname>
          </string-name>
          .
          <article-title>Nearest neighbors distance ratio open-set classifier</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>106</volume>
          (
          <issue>3</issue>
          ):
          <fpage>359</fpage>
          -
          <lpage>386</lpage>
          ,
          <year>Mar 2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Duchesnay</surname>
          </string-name>
          .
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>12</volume>
          :
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Travis Oliphant.
          <article-title>NumPy: A guide to NumPy</article-title>
          . USA: Trelgol Publishing,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Steven</surname>
            <given-names>Bird</given-names>
          </string-name>
          , Ewan Klein, and
          <string-name>
            <given-names>Edward</given-names>
            <surname>Loper. Natural Language Processing with Python. O'Reilly Media</surname>
          </string-name>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          .
          <source>Machine Learning in automated text categorization. ACM Computing Surveys</source>
          ,
          <volume>34</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>47</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>J.</given-names>
            <surname>Platt</surname>
          </string-name>
          .
          <article-title>Probabilistic outputs for support vector machines and comparison to regularize likelihood methods</article-title>
          . In A.J.
          <string-name>
            <surname>Smola</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Bartlett</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Schoelkopf</surname>
          </string-name>
          , and D. Schuurmans, editors,
          <source>Advances in Large Margin Classifiers</source>
          , pages
          <fpage>61</fpage>
          -
          <lpage>74</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>William</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Teahan</surname>
            and
            <given-names>David J.</given-names>
          </string-name>
          <string-name>
            <surname>Harper</surname>
          </string-name>
          .
          <article-title>Using Compression-Based Language Models for Text Categorization</article-title>
          , pages
          <fpage>141</fpage>
          -
          <lpage>165</lpage>
          . Springer Netherlands, Dordrecht,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Martin</surname>
            <given-names>Potthast</given-names>
          </string-name>
          , Sarah Braun, Tolga Buz, Fabian Duffhauss, Florian Friedrich, Jörg Marvin Gülzow, Jakob Köhler, Winfried Lötzsch, Fabian Müller, Maike Elisa Müller, Robert Paßmann, Bernhard Reinke, Lucas Rettenmeier, Thomas Rometsch, Timo Sommer, Michael Träger,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Wilhelm</surname>
          </string-name>
          , Benno Stein, Efstathios Stamatatos, and
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Hagen</surname>
          </string-name>
          .
          <article-title>Who wrote the web? Revisiting influential author identification research applicable to information retrieval</article-title>
          . In Nicola Ferro, Fabio Crestani,
          <string-name>
            <surname>Marie-Francine</surname>
            <given-names>Moens</given-names>
          </string-name>
          , Josiane Mothe, Fabrizio Silvestri, Giorgio Maria Di Nunzio, Claudia Hauff, and Gianmaria Silvello, editors,
          <source>Proc. of the European Conference on Information Retrieval</source>
          , pages
          <fpage>393</fpage>
          -
          <lpage>407</lpage>
          . Springer International Publishing,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Mike</surname>
            <given-names>Kestemont</given-names>
          </string-name>
          , Justin Anthony Stover, Moshe Koppel, Folgert Karsdorp, and
          <string-name>
            <given-names>Walter</given-names>
            <surname>Daelemans</surname>
          </string-name>
          .
          <article-title>Authenticating the writings of julius caesar</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>63</volume>
          :
          <fpage>86</fpage>
          -
          <lpage>96</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>José</surname>
          </string-name>
          <article-title>Eleandro Custódio and Ivandré Paraboni</article-title>
          .
          <article-title>EACH-USP Ensemble cross-domain authorship attribution</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <surname>Jian-Yun Nie</surname>
          </string-name>
          , and Laure Soulier, editors,
          <source>Working Notes of CLEF 2018 - Conference and Labs of the Evaluation Forum, CEUR Workshop Proceedings. CLEF and CEUR-WS.org</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Mike</surname>
            <given-names>Kestemont</given-names>
          </string-name>
          , Michael Tschuggnall, Efstathios Stamatatos, Walter Daelemans, Günther Specht, Benno Stein, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Potthast</surname>
          </string-name>
          .
          <article-title>Overview of the author identification task at PAN-2018: cross-domain authorship attribution and style change detection</article-title>
          .
          <source>In Working Notes of CLEF 2018 - Conference and Labs of the Evaluation Forum. CEUR-WS.org</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Lukas</surname>
            <given-names>Muttenthaler</given-names>
          </string-name>
          , Gordon Lucas, and
          <string-name>
            <given-names>Janek</given-names>
            <surname>Amann</surname>
          </string-name>
          .
          <article-title>Authorship Attribution in Fan-fictional Texts Given Variable Length Character and Word n-grams</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Basile</surname>
          </string-name>
          .
          <article-title>An Open-Vocabulary Approach to Authorship Attribution</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>Martijn</given-names>
            <surname>Bartelds</surname>
          </string-name>
          and Wietse de Vries.
          <article-title>Improving Cross-Domain Authorship Attribution by Combining Lexical and Syntactic Features</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Carolina Martín del Campo</surname>
            <given-names>Rodríguez</given-names>
          </string-name>
          , Daniel Alejandro Pérez Álvarez, Christian Efraín Maldonado Sifuentes, Grigori Sidorov, Ildar Batyrshin, and
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          .
          <article-title>Authorship Attribution through Punctuation n-grams and Averaged Combination of SVM</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          .
          <article-title>Authorship attribution using text distortion</article-title>
          .
          <source>In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>1</volume>
          ,
          <string-name>
            <surname>Long</surname>
            <given-names>Papers</given-names>
          </string-name>
          , pages
          <fpage>1138</fpage>
          -
          <lpage>1149</lpage>
          . Association for Computational Linguistics,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Andrea</surname>
            <given-names>Bacciu</given-names>
          </string-name>
          , Massimo La Morgia, Alessandro Mei, Eugenio Nerio Nemmi, Valerio Neri, and
          <string-name>
            <given-names>Julinda</given-names>
            <surname>Stefa</surname>
          </string-name>
          .
          <article-title>Cross-domain Authorship Attribution Combining Instance Based and Profile Based Features</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <article-title>Fredrik Johansson and Tim Isbister. FOI Cross-Domain Authorship Attribution for Criminal Investigations</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29. Hans van Halteren.
          <article-title>Cross-Domain Authorship Attribution with Federales</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Mostafa</surname>
            <given-names>Rahgouy</given-names>
          </string-name>
          , Hamed Babaei Giglou, Taher Rahgooy, Mohammad Karami Sheykhlan, and
          <string-name>
            <given-names>Erfan</given-names>
            <surname>Mohammadzadeh</surname>
          </string-name>
          .
          <article-title>Cross-domain Authorship Attribution: Author Identification using a Multi-Aspect Ensemble Approach</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Martin</surname>
            <given-names>Potthast</given-names>
          </string-name>
          , Tim Gollub, Matti Wiegmann, and
          <article-title>Benno Stein. TIRA Integrated Research Architecture</article-title>
          . In Nicola Ferro and Carol Peters, editors,
          <source>Information Retrieval Evaluation in a Changing World - Lessons Learned from 20 Years of CLEF</source>
          . Springer,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <article-title>José Custódio and Ivandre Paraboni. Multi-channel Open-set Cross-domain Authorship Attribution</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <given-names>Lukasz</given-names>
            <surname>Gagala</surname>
          </string-name>
          .
          <article-title>Authorship Attribution with Logistic Regression and Imposters</article-title>
          . In Linda Cappellato, Nicola Ferro,
          <string-name>
            <given-names>David E.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and Henning Müller, editors,
          <source>CLEF 2019 Labs and Workshops, Notebook Papers. CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Walter</surname>
            <given-names>Scheirer</given-names>
          </string-name>
          , Anderson Rocha, Archana Sapkota, and
          <string-name>
            <given-names>Terrance</given-names>
            <surname>Boult</surname>
          </string-name>
          .
          <article-title>Toward open set recognition</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>35</volume>
          (
          <issue>7</issue>
          ):
          <fpage>1757</fpage>
          -
          <lpage>1772</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Chuanxing</surname>
            <given-names>Geng</given-names>
          </string-name>
          , Sheng-jun
          <string-name>
            <surname>Huang</surname>
            , and
            <given-names>Songcan</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Recent advances in open set recognition: A survey</article-title>
          . arXiv preprint arXiv:
          <year>1811</year>
          .08581,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Efstathios</surname>
            <given-names>Stamatatos</given-names>
          </string-name>
          , Walter Daelemans, Ben Verhoeven, Benno Stein, Martin Potthast, Patrick Juola, Miguel A.
          <string-name>
            <surname>Sánchez-Pérez</surname>
          </string-name>
          , and
          <string-name>
            <surname>Alberto</surname>
          </string-name>
          Barrón-Cedeño.
          <article-title>Overview of the author identification task at pan 2014</article-title>
          .
          <article-title>In CLEF 2014 Evaluation Labs</article-title>
          and Workshop - Working Notes Papers, Sheffield, UK,
          <year>September 2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>