<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Multi-label Style Change Detection by Solving a Binary Classification Problem</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eivind Strøm</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Norwegian University of Science and Technology</institution>
          ,
          <addr-line>Høgskoleringen 1, 7491 Trondheim</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>2</volume>
      <fpage>1</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>Style change detection is an important area of research with many applications, including plagiarism detection, digital forensics, and authorship attribution. This paper proposes a solution to the three sub-tasks for the PAN 2021 shared task on style change detection, featuring a challenging multi-label multi-output classification problem. We devise a pragmatic approach to the problem, solving it by binary classification before obtaining final predictions. A custom stacking ensemble is developed and trained separately on previously successful text embeddings and features for increased performance. Our solution achieves a macro-averaged F1-score of 0.7954 for single- and multi-author classification and is the best performing solution submitted for task 1. On task 2, we obtain a score of 0.7069 when detecting author change between paragraphs. For multi-label author attribution, our solution achieves a score of 0.4240 and performs significantly better than the random baseline. Being a pragmatic solution to a novel problem in the series of PAN tasks on style change detection, our approach ofers several opportunities for further research.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Style change</kwd>
        <kwd>Multi-authorship</kwd>
        <kwd>Binary classification</kwd>
        <kwd>Stacking ensemble</kwd>
        <kwd>BERT</kwd>
        <kwd>NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The task of quantifying writing style has been a challenging task since the 19th century,
beginning with the pioneering study by Mendenhall [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] on Shakespeare plays. This area
of research, stylometry, has several modern-day applications, including forensics, plagiarism
detection and the linking of social media accounts [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. In this paper, we examine a fundamental
and challenging question within the field of stylometry: Given a document, can we find evidence
for the document being written by multiple authors, and are we able to attribute paragraphs to
respective authors based on their writing style?
      </p>
      <p>
        Style change detection was introduced as a shared task for the PAN 2017 event, with the
goal of identifying the border positions where authorship changes [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The results, however,
proved the task to be extremely challenging. As a result, the task was significantly relaxed
for the following years. First simplified as a binary classification task, the goal was detecting
whether a document is single- or multi-authored [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. As participants obtained encouraging
results and accuracies of up to 0.8993, the task was further expanded to detecting the exact
number of authors within a document [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and identifying style changes between two consecutive
paragraphs [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        This paper proposes a solution to the PAN 2021 shared task on style change detection, which
has been steered back towards its original goal: detecting the position of authorship change
and assigning each paragraph to its respective author [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. We propose a pragmatic solution by
which the multi-label multi-output classification problem is first solved by binary classification
before recursively converted to its original formulation. For classification, we employ several
custom-built stacking ensembles that are trained on textual features and embeddings that have
been proven efective in previous tasks. Our solution is the best performing submitted solution
to task 1 and the second best on task 2. On task 3, our model performs significantly better
than random guesses. Lastly, we release our project code for reproducibility and aid in further
research: https://github.com/eivistr/pan21-style-change-detection-stacking-ensemble
      </p>
      <p>The rest of the paper is structured as follows. Section 2 provides an overview of the sub-tasks,
dataset and evaluation. The methodology and details of our approach are presented in section 3.
Section 4 and section 5 present the experimental setting and final results. Lastly, section 6
summarizes our findings and proposes directions for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>
        The shared PAN 2021 style change detection task is split into three sub-tasks [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]:
1. Given a text, find out whether it is written by a single or multiple authors.
2. Given a multi-authored text that contains a number of style changes between paragraphs,
ifnd the position of the changes.
3. Given a multi-authored text, assign all paragraphs of the text uniquely to some author
out of the number of authors you assume for the document.
      </p>
      <p>
        Of these sub-tasks, task 1 and 2 are the same as the previous year’s edition of the style change
task [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. However, this year the dataset provided is slightly diferent and does not contain a
separation of wide and narrow topic collections. The overall structure of the task is illustrated
by Figure 1.
      </p>
      <p>The dataset consists of a single collection of documents, where each document is based on a
user post (or concatenation of several user posts) from the StackExchange1 network, covering a
wide range of topics. The dataset is split into a training and validation set comprising 11, 200
and 2, 400 documents, respectively. The documents in the training and validation set contain
anywhere between 1 and 4 unique authors and the dataset is balanced in terms of number of
authors per document.</p>
      <p>
        Evaluation of the solution is performed on a test set consisting of 2, 400 documents and
scored by the macro averaged F1-score. As we do not have the test set and test labels available
during development, our solution is evaluated on the available validation set. Final test results
are obtained through submission to the TIRA evaluation platform [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>This section describes the methodology of our proposed solution. First, we present the feature
extraction process fundamental to our approach. Secondly, we describe the stacking ensemble
classifier and our reasoning behind this choice of architecture. Lastly, we present our approach
to solving the tasks as binary classification problems.</p>
      <sec id="sec-3-1">
        <title>3.1. Feature extraction</title>
        <p>
          For feature extraction we rely on two methods that have been successfully applied in previous
tasks: generating textual embeddings using Google AI’s BERT transformer [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and extracting
textual features and statistics based on the winning submission of the PAN 2018 task [11].
Features are extracted in two rounds, firstly at the document level (for use in task 1) and secondly
on the paragraph level, i.e., for each individual paragraph (task 2 and 3). In the following, we
summarize the important points of the feature extraction approach, for further details we refer
to the work of Iyer and Vosoughi [12] and Zuo et al. [13].
        </p>
        <p>Text embeddings: We follow the approach of Iyer and Vosoughi [12] for obtaining embedded
text features. Firstly, tokenized sentences appropriate for generating embeddings using the
BERT model are generated by splitting each paragraph into sentences before tokenizing. Prior to
splitting, paragraphs are pre-processed by the removal of noisy ’.’, ’?’ and ’!’ characters to ensure
that prefixes, sufixes, website domains, acronyms, abbreviations, and digits do not introduce
incomplete sentences. After processing, sentences are tokenized and embedded using the BERT
Base Cased model,2 generating a 12 ×  × 768 tensor for each sentence, where  is the sentence
2The BERT Base Cased model is case sensitive and configured with default parameters: layers=12, hidden
size=768, self-attention heads=12, total parameters=110M. We use the transformers library for PyTorch provided
Documents</p>
        <p>Pre-process
and split into
sentences
Split into
paragraphs
BERT
model</p>
        <p>Sum over last 4 of 12 layers
Sum over token count L</p>
        <p>Sum sentence
vectors per paragraph
length. We follow the recommended approach and sum the embeddings of the last four layers
to obtain an  × 768 tensor, and further sum over the first dimension to obtain sentence-level
embeddings. While summing rather than averaging can lead to large discrepancies between
sentences of diferent lengths, sentence length can be an important factor for style detection
and is likely important to capture [12].</p>
        <p>Converting sentence embeddings to paragraphs is achieved by adding each sentence
embedding to produce a final paragraph feature vector of size 1 × 768. Although, Iyer and Vosoughi
[12] achieves best performance by normalizing paragraph embeddings by sentence count, we
achieve slightly higher performance by simply summing each vector. Document embeddings
are obtained by summation of the containing paragraph embeddings and normalization by the
document sentence count.</p>
        <p>Text features: To extract text features that characterize writing style, we extract the same
features as Zuo et al. [13] and Zlatkova et al. [11] on the paragraph level:
1. Character-based lexical features: The number of each distinct special character, spaces,
punctuation, parentheses and quotation marks as separate features.
2. Sentence- and word-based features: Distribution of POS-tags, token length, number of
sentences, sentence length, average word length, words in all-caps and counts of words
above and below 2-3 and 6 characters as separate features.
3. Contracted word forms: Count of preference towards one type of contraction, e.g. "I’m"
versus "I am". The total number of occurrences of contractions and fully written forms
are used as separate features.
4. Function words: The frequency of each function word is counted and used as a separate
feature. We use the same list as Zuo et al. [13] which is a combination of previously
by Hugging Face: https://huggingface.co/transformers/model_doc/bert.html</p>
        <p>defined words and the function word list from the NLTK 3 library.
5. Readability indexes obtained using the Python Textstat4 library: Flesch reading ease score,
Dale-Chall readability score, SMOG grade, Flesch-Kincaid grade, Coleman-Liau index,
Gunning-Fog index, automated readability index and the Linsear Write readability metric.
Additionally, we count the number of dificult words and keep all indexes as separate
features.</p>
        <p>Features are both extracted and saved per paragraph, before summed per document to obtain
both paragraph- and document-level features. A total of 478 features per document and per
paragraph.5</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Stacking ensemble classifier</title>
        <p>The idea behind using a stacking ensemble for classification is to combine the learning
capabilities of multiple classifiers and reduce overall bias. Preliminary experiments showed that
LightGBM, a state-of-the-art gradient boosting tree for classification [ 14], trained on the
extracted text features outperformed models trained on BERT-embeddings. Furthermore, training
on both BERT-embeddings and text features did not improve upon solely using text features. We
hypothesized that training two sets of classifiers separately on each feature vector would allow
the classifiers to obtain more discriminatory power on each vector. Thus, we train classifiers on
text features and BERT-embeddings separately and combine their predictions by a meta-learner,
i.e., a stacking ensemble architecture. Figure 3 presents an overview of the classification process
for each task, and how the ensemble for each task consists of two sets of classifiers. More details
on how each task is structured as binary classification is presented in the next subsection.</p>
        <p>Our architecture is that of stacking with -fold cross validation to avoid target leakage
from base level classifiers to the meta-learner [ 15]. The meta-learner trains on  out-of-fold
predictions produced by the base level classifiers and is evaluated on the validation set. We
use LightGBM as our state-of-the-art classifier in all ensembles and select the best three
Scikitlearn classifiers for each feature vector on each task. The exact configuration of the ensemble
architecture is dependent on the task and we describe it in further detail in section 4.</p>
        <p>Stacking ensembles have frequently been the winner of Kaggle competitions, in which
avoiding overfitting to the training set is crucial for winning chances. Our results on the test set
show a marginal decrease in performance compared to the validation set on task 2 and 3, and
even improved performance on task 1. This indicates that we have been successful in avoiding
overfitting to the training data.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Classification approach</title>
        <p>With a pragmatic mindset we solve each task as binary classification. This requires some
combination and processing of extracted text features and embeddings. An overview of the
3https://www.nltk.org/
4https://pypi.org/project/textstat/
5We kindly thank a reviewer for pointing out that addition of readability indexes does not make intuitive sense.
Thus, performance could potentially be further improved by more carefully combining the paragraph-level feature
vectors.
Classifiers 1 - 4</p>
        <p>Classifiers 1 - 4</p>
        <p>Meta-learner
Task 1 predictions</p>
        <p>Concatenate
consecutive
paragraph vectors
1 x 956 vector
per case</p>
        <p>Sum consecutive
paragraph vectors
1 x 768 vector
per case</p>
        <p>Concatenate
vectors of paragraph</p>
        <p>pairs
1 x 956 vector
per case</p>
        <p>Sum vectors of
paragraph pairs
1 x 768 vector
per case
Add binary labels</p>
        <p>Add binary labels</p>
        <p>Add binary labels</p>
        <p>Add binary labels
Classifiers 1 - 4</p>
        <p>Classifiers 1 - 4</p>
        <p>Classifiers 1 - 3</p>
        <p>Classifiers 1 - 3</p>
        <p>Meta-learner
Task 2 predictions</p>
        <p>Meta-learner
Task 3 binary
predictions
Recursive algorithm
Task 3 predictions
approach is provided by Figure 3.</p>
        <p>Task 1: Classifying single- or multi-author documents is achieved by classification on the
document-level features already obtained in the feature extraction process. In this case, we
have one feature vector per document and label and classify each document as being either
single- or multi-authored. Previous research finds that classification on document-level features
generally works very well when identifying the presence of style changes in documents [11]
and multi-authored documents [12]. This is further supported by our high performance (macro
F1-score of 0.7954) on task 1.</p>
        <p>Task 2: To classify author change between two consecutive paragraphs, we combine the features
of the two considered paragraphs to obtain the feature vector and classify an author change.
This approach is similar to that of Iyer and Vosoughi [12]. Experimentation shows that, when
considering text embeddings, addition of the two feature vectors yields the best result. For text
features, we find that concatenating the arrays produce significantly better results than addition.
Reasons for this could be that addition of the feature vectors simply average out the diferences
in the paragraphs, while concatenation keeps the features for each paragraph separate and,
thus, provide more discriminative power. Keep in mind that this doubles the text feature vector
to a size of 1 × 956. This formulation results in a total of  − 1 cases for each document where
 is the number of paragraphs.</p>
        <p>As a reviewer points out in retrospect, we could have used task 1 predictions to first predict
whether a document is single- or multi authored before obtaining task 2 predictions, as we do
on task 3. We find this would have resulted in a performance increase of roughly 0.02 on task 2.
Unfortunately, we did not implement this before the submission deadline as our focus was on
solving task 3, being the novel task on style change detection.</p>
        <p>Task 3 (binary): To assign paragraphs uniquely to all authors, we realize that we can pose
this as a binary classification problem by asking whether any pair of paragraphs in the same
document are written by the same author. In other words, we need to compare each paragraph
to every other paragraph in a document and convert the provided multi-author label into a
binary label. This implies (︀ 2)︀ = 12 ( − 1) cases for a document with  paragraphs. To
classify two paragraphs, we combine their features in the same way as for task 2, i.e., addition
of embeddings and concatenation of text features. The approach of classifying whether two
paragraphs have the same author turns out to be very efective and achieves similar performance
to that of task 2. However, the challenge becomes assigning each paragraph to its unique author
in the multi-author multi-label setting.</p>
        <p>Task 3 (multi-label): To obtain final multi-label predictions and assign each paragraph a
unique author per document we devise a recursive strategy, using our task 1 and task 3 binary
predictions. The process is detailed as an algorithm in Figure 4. Firstly, we use our task 1
classifier to determine whether we believe the document is single- or multi-authored. Documents
classified as single-authored are assigned author 1 for all paragraphs. Secondly, in any given
multi-authored document, the first paragraph is always assigned to author 1. Continuing, we
compare the second paragraph with the first and determine the probability that it is authored
by the same author, i.e., using our task 3 binary classification. If the probability is greater than
0.5, paragraph two is assigned to author 1, otherwise we assign it to author 2. Paragraph 3
is compared to both paragraph 1 and 2, and we assign its author to the most similar author
that we previously determined. If it is unsimilar to both, we assign paragraph 3 to a new
author. At some point we might have assigned paragraphs to all 4 authors, at which point
we continue by assigning paragraphs by the most similar paragraph even though all previous
paragraphs could have similarity scores below 0.5. This is a pragmatic approach, and there is
certainly some information loss in this process. Especially for cases where there are errors on
the first paragraphs as these will propagate and cause additional errors for later paragraphs.
We did experiment with tuning the probability thresholds, however, this resulted in very minor
performance improvements and we elected to keep similarity thresholds standard at 0.5. The
improvement of this method is a suggested direction for future work.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Setting</title>
      <p>In this section we present details on the experimental setting. For training the ensemble we
used 4-fold stratified cross validation. Processing the training dataset for binary classification
yields a total of 8,400 positive (2,800 negative) training cases for task one, 35,416 positive (30,636
negative) cases for task 2 and 125,679 positive (154,082 negative) cases for task 3.</p>
      <p>
        The stacking ensemble was trained with half of the base level classifiers on text embeddings
and the other half on text features. The selection of classifiers for each task and feature vector was
based on selecting the best performing classifiers with diferent underlying methods (not only
tree and boosting methods). We relied on the scikit-learn library [16] for all classifiers except
LightGBM and evaluated the following classifiers as candidates: Support Vector Machines,
AdaBoost, Decision Trees, Random Forest, Extra Trees, Multi-layer Perceptron, k-Nearest
Neighbors, Logistic Regression, Linear Discriminant Analysis and Bernoulli- and Gaussian
Naive Bayes. Table 1 shows the final ensemble configuration and which classifiers are trained
on each feature vector and each task. Note that we only used 6 classifiers for task 3 due to the
larger dataset. Furthermore, due to initial size and performance restrictions using the TIRA
submission platform [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we excluded the random forest classifier on task 3 which reduced
validation performance by roughly 0.01.6
      </p>
      <p>Hyperparameters was optimized for the LightGBM classiefirs only, being the most important
classifier in the ensemble. All other classifiers were trained with default parameters in the
scikit-learn library, version 0.24.2. Given the large dataset, the tuning of 6-8 individual classifiers
for each task was intractable given the scope and available resources. Furthermore, one could
argue that optimization of each base level classifier is unnecessary as diferences in bias is
optimized by the meta-learner. LightGBM was optimized using the Optuna hyperparameter
optimization framework [17]. We find that correcting for unbalanced datasets using LightGBMs
6Using random forest on task 3 resulted in model files upwards of 12GB and unreasonable memory usage due
to the high dimensional data.
"is_unbalance" parameter improves performance only on task 1 and 2. Preliminary experiments
showed that reduction of the number of features based on LightGBMs feature importance
reduced performance. Therefore, feature vectors were kept at their original size. Although
common practice, we also found data scaling to reduce overall performance and was, thus,
omitted.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In the following, we present the results obtained by our approach. In addition to the macro
F1-score we provide the accuracy score to allow comparison of our results with the PAN 2020
style change detection sub-tasks.7</p>
      <p>Table 2 shows the validation results for the best LightGBM and ensemble model on each task.
The best validation scores are marked in bold, achieving a macro F1-score of 0.7828, 0.7107
and 0.4261 on task 1, 2, and 3, respectively. Our results indicate that the ensemble model is
superior to the best LightGBM model for any combination of task and performance measure.
On task 3, we compare the results to that of using a uniform random guesser instead of our
binary classifier and find that the ensemble outperforms the random baseline by 0.1375 points.
However, as expected there is a significant reduction in performance going from binary to
multi-label predictions, especially in the macro F1-score. As for binary classification, the custom
ensemble outperforms LightGBM. Note that the random guesser has slightly higher than 0.5
accuracy, as task 1 predictions are still used to classify single-authored documents and assign
labels to these cases to obtain comparable results.</p>
      <p>Table 3 shows our final test results after submitting and evaluating our model on the TIRA
platform. We achieve final scores of 0.7954, 0.7069 and 0.4240 for tasks 1, 2, and 3, respectively.
Our solution is the best performing solution among submitted solutions for task 1, and among
top 3 for tasks 2 and 3. The test scores are similar to validation scores, and we observe a
performance increase on task 1, indicating that our models have generalized well to the test set.</p>
      <p>7The PAN 2020 style detection tasks was evaluated by micro-averaged F1, which reduces to accuracy for binary
classification problems.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>In this paper, we propose a solution to the PAN 2021 shared task on style change detection,
attempting to answer the question: Given a document, can we find evidence for the document
being written by multiple authors, and are we able to attribute paragraphs to respective authors
based on their writing style?</p>
      <p>Our solution provides encouraging results. We apply previously tested feature extraction
methods and develop a custom stacking ensemble framework trained separately on diferent
feature vectors. Furthermore, we propose a pragmatic solution to the dificult problem of
multi-label multi-output classification by first solving a relaxed problem by binary classification.</p>
      <p>We achieve macro F1-scores on the validation set of 0.7761 when classifying single- and
multiauthored documents (task 1), and 0.7107 when detecting an author change between consecutive
paragraphs (task 2). For multi-label author attribution (task 3), our solution achieves a score of
0.4261, significantly outperforming random guesses (0.2886). The relaxed formulation of task
3 achieves a macro F1-score of 0.7175. On the test set, our final submission scores are 0.7954,
0.7069, and 0.4240 for tasks 1, 2 and 3 respectively and our solution is the best performing
solution on task 1 among other submitted solutions. Lastly, our results indicate that the use
of stacking ensembles improves performance across reported metrics when compared to the
optimized LightGBM model. We suggest interesting directions for future work:
1. The use of stacking ensembles yields small performance improvements considering the
efort spent on building the models. Thus, we hypothesize that identifying additional key
features is paramount for further improvements.
2. Given the reduced multi-label performance on task 3, there are likely opportunities to
increase performance going from binary to multi-label predictions. Exploring the use of
hierarchical classification or comparison of document paragraphs in unison (as opposed to
recursive comparisons) would be interesting for further research and might significantly
increase performance on task 3.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was motivated by academic interests and undertaken as a project in the course
TDT4310 - Intelligent Text Analytics and Language Understanding at the Norwegian University
of Science and Technology in Trondheim. I would like to thank professor Björn Gambäck for
organizing a complete and well rounded course on natural language processing.
transformers for language understanding, in: Proceedings of the 2019 Conference of the
North American Chapter of the Association for Computational Linguistics, Association
for Computational Linguistics, 2019, pp. 4171–4186. doi:10.18653/v1/N19-1423.
[11] D. Zlatkova, D. Kopev, K. Mitov, A. Atanasov, M. Hardalov, I. Koychev, P. Nakov, An
Ensemble-Rich Multi-Aspect Approach for Robust Style Change Detection—Notebook
for PAN at CLEF 2018, in: L. Cappellato, N. Ferro, J.-Y. Nie, L. Soulier (Eds.), CLEF 2018
Evaluation Labs and Workshop – Working Notes Papers, 10-14 September, Avignon, France,
CEUR-WS.org, 2018. URL: http://ceur-ws.org/Vol-2125/.
[12] A. Iyer, S. Vosoughi, Style Change Detection Using BERT—Notebook for PAN at CLEF 2020,
in: L. Cappellato, C. Eickhof, N. Ferro, A. Névéol (Eds.), CLEF 2020 Labs and Workshops,
Notebook Papers, CEUR-WS.org, 2020. URL: http://ceur-ws.org/Vol-2696/.
[13] C. Zuo, Y. Zhao, R. Banerjee, Style Change Detection with Feed-forward Neural Networks,
in: L. Cappellato, N. Ferro, D. Losada, H. Müller (Eds.), CLEF 2019 Labs and Workshops,
Notebook Papers, CEUR-WS.org, 2019. URL: http://ceur-ws.org/Vol-2380/.
[14] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, T.-Y. Liu, Lightgbm: A highly
eficient gradient boosting decision tree, in: I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach,
R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing
Systems, volume 30, Curran Associates, Inc., 2017.
[15] C. C. Aggarwal, Data Classification: Algorithms and Applications, 1st ed., Chapman &amp;</p>
      <p>Hall/CRC, 2014.
[16] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel,
P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher,
M. Perrot, Édouard Duchesnay, Scikit-learn: Machine learning in python, Journal of
Machine Learning Research 12 (2011) 2825–2830.
[17] T. Akiba, S. Sano, T. Yanase, T. Ohta, M. Koyama, Optuna: A next-generation
hyperparameter optimization framework, in: Proceedings of the 25th ACM SIGKDD International
Conference on Knowledge Discovery &amp; Data Mining, KDD ’19, Association for Computing
Machinery, New York, NY, USA, 2019, p. 2623–2631. doi:10.1145/3292500.3330701.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T. C.</given-names>
            <surname>Mendenhall</surname>
          </string-name>
          ,
          <article-title>The characteristic curves of composition</article-title>
          ,
          <source>Science</source>
          (
          <year>1887</year>
          )
          <fpage>237</fpage>
          -
          <lpage>246</lpage>
          . doi:
          <volume>10</volume>
          .1126/science.ns-
          <volume>9</volume>
          .
          <year>214S</year>
          .
          <fpage>237</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Iqbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Binsalleeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Fung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Debbabi</surname>
          </string-name>
          ,
          <article-title>Mining writeprints from anonymous e-mails for forensic investigation</article-title>
          ,
          <source>Digital Investigation</source>
          <volume>7</volume>
          (
          <year>2010</year>
          )
          <fpage>56</fpage>
          -
          <lpage>64</lpage>
          . doi:
          <volume>10</volume>
          .1016/j. diin.
          <year>2010</year>
          .
          <volume>03</volume>
          .003.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vosoughi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <article-title>Digital stylometry: Linking profiles across social networks</article-title>
          ,
          <source>in: Social Informatics</source>
          , Springer International Publishing,
          <year>2015</year>
          , pp.
          <fpage>164</fpage>
          -
          <lpage>177</lpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>319</fpage>
          -27433-1\_
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tschuggnall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Verhoeven</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Daelemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Specht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Potthast, Overview of the Author Identification Task at PAN 2017: Style Breach Detection and Author Clustering</article-title>
          , in: L.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Goeuriot</surname>
          </string-name>
          , T. Mandl (Eds.),
          <source>Working Notes Papers of the CLEF 2017 Evaluation Labs</source>
          , volume
          <volume>1866</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2017</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-1866/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tschuggnall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Daelemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Specht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Potthast, Overview of the Author Identification Task at PAN-2018: Cross-domain Authorship Attribution and Style Change Detection</article-title>
          , in: L.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>J.-Y.</given-names>
          </string-name>
          <string-name>
            <surname>Nie</surname>
          </string-name>
          , L. Soulier (Eds.),
          <source>Working Notes Papers of the CLEF 2018 Evaluation Labs</source>
          , volume
          <volume>2125</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2018</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2125</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Zangerle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tschuggnall</surname>
          </string-name>
          , G. Specht,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <source>Overview of the Style Change Detection Task at PAN</source>
          <year>2019</year>
          , in: L.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Losada</surname>
          </string-name>
          , H. Müller (Eds.),
          <article-title>CLEF 2019 Labs and Workshops, Notebook Papers, CEUR-WS</article-title>
          .org,
          <year>2019</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2380</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Zangerle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          , G. Specht,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <source>Overview of the Style Change Detection Task at PAN</source>
          <year>2020</year>
          , in: L.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickhof</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Névéol (Eds.),
          <article-title>CLEF 2020 Labs and Workshops, Notebook Papers, CEUR-WS</article-title>
          .org,
          <year>2020</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2696</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Zangerle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <article-title>Overview of the Style Change Detection Task at PAN 2021, in: CLEF 2021 Labs and Workshops, Notebook Papers, CEUR-WS</article-title>
          .org,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gollub</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , TIRA Integrated Research Architecture, in: N.
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Peters (Eds.),
          <source>Information Retrieval Evaluation in a Changing World, The Information Retrieval Series</source>
          , Springer, Berlin Heidelberg New York,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>030</fpage>
          -22948-1\_5.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of deep bidirectional</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>