<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring the Potential of Feature Density in Estimating Machine Learning Classifier Performance with Application to Cyberbullying Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Juuso Eronen</string-name>
          <email>eronen.juuso@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michal Ptaszynski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fumito Masui</string-name>
          <email>f-masuig@cs.kitami-it.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gniewosz Leliwa</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michal Wroczynski</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kitami Institute of Technology</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Samurai Labs</institution>
          ,
          <country country="PL">Poland</country>
        </aff>
      </contrib-group>
      <fpage>5</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>In this research, we analyze the potential of Feature Density (FD) as a way to comparatively estimate machine learning (ML) classifier performance prior to training. The goal of the study is to aid in solving the problem of resource-intensive training of ML models which is becoming a serious issue due to continuously increasing dataset sizes and the ever rising popularity of Deep Neural Networks (DNN). The issue of constantly increasing demands for more powerful computational resources is also affecting the environment, as training large-scale ML models are causing alarmingly-growing amounts of CO2 emissions. Our approach is to optimize the resource-intensive training of ML models for Natural Language Processing to reduce the number of required experiments iterations. We expand on previous attempts on improving classifier training efficiency with FD while also providing an insight to the effectiveness of various linguistically-backed feature preprocessing methods for dialog classification, specifically cyberbullying detection.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>One of the challenges in machine learning (ML) has always
been estimating how well different classification algorithms
will perform with a given dataset. Although there are
classifiers that tend to be highly effective on a variety of different
problems, they might be easily outperformed by others on a
dataset specific scale. As it is difficult to identify a classifier
that would perform best with every kind of dataset [Michie et
al., 1995], it comes down to the user (researcher, or ML
practitioner) to determine experimentally, which classifier could
be appropriate based on their knowledge of the field and
previous experiences.</p>
      <p>A common way when estimating the performance of
different classifiers is to select a variety of possible classifiers
to experiment on and train them using cross-validation to
aid in getting the best possible average estimations of their
performances. With a sufficiently small dataset and using a
computationally efficient algorithm, this approach works very
well. Even though it is possible to get accurate estimations of
the classifier performance this way, it is multiple times more
costly.</p>
      <p>Previously, there have been some attempts to estimate the
performance of a ML model before any training. One
proposal to this problem is using meta-learning and training a
model using dataset characteristics to estimate classifier
performance [Gama and Brazdil, 1995]. Another approach is
extrapolating results from small datasets to simulate the
performance using larger datasets [Basavanhally et al., 2010].</p>
      <p>The importance of resolving this issue comes not only from
the increased computational requirements, but also from its
environmental effect. This is directly caused by the increased
popularity of the fields of Artificial Intelligence (AI) and ML.
Training classifiers on large datasets is both time consuming
and computationally intensive while leaving behind a
noticeable carbon footprint [Strubell et al., 2019]. To move towards
greener AI [Schwartz et al., 2019], it is necessary to inspect
the core of ML methods and find potential points of
improvement. In order to save computational power and reduce
emissions, it would be useful to roughly estimate classifier
performance prior to training.</p>
      <p>The ability to estimate classifier performance before the
training would also have important practical implications. In
dialog agent applications, one of the areas where the need
for this is becoming more urgent is in forum moderation,
specifically the detection of harmful and abusive behaviour
observed online, known as cyberbullying (CB). The number
of CB cases has been constantly growing since the increase
of the popularity of Social Networking Services (SNS)
[Hinduja and Patchin, 2010; Ptaszynski and Masui, 2018]. The
consequences of unattended cases of online abuse are known
to be serious, leading the victims to self mutilation, or even
suicides, or on the opposite, to attacking their offenders in
revenge. Being able to roughly estimate which classifier
settings can be rejected, would make the process of
implementation of automatic cyberbullying detection for various
languages and social networking platforms more efficient.</p>
      <p>To contribute to that, we conduct an in-depth analysis of
the effectiveness of FD proposed previously by [Ptaszynski et
al., 2017] to comparatively estimate the performance of
different classifiers before training. We also analyze the
effectiveness of various linguistically-backed feature
preprocessing methods, including lemmas, Named Entity Recognition
(NER) and dependency information-based features, with an
application to automatic cyberbullying detection.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Previous Research</title>
      <p>Classifier Performance Estimation
[Gama and Brazdil, 1995] proposed that classifier
performance could be estimated by training a regression model
based on meta-level characteristics of a dataset. The
characteristics used included simple measures like number of
examples and number of attributes, statistical measures like
standard deviation ratio and various information based measures
like class entropy. These measures are defined in the
STATLOG project [King et al., 1995].</p>
      <p>This meta-learning approach was taken further by
[Bensusan and Kalousis, 2001] who introduced the
Landmarking method, using learners themselves to characterize the
datasets. This means using computationally non-demanding
classifiers, like Naive Bayes (NB), to obtain important
insights about the datasets. The method outperformed the
previous characterization method and had moderate success in
ranking learners.</p>
      <p>Later, [Blachnik, 2017] improved on the Landmarking
method by proposing the use of information from instance
selection methods as landmarks. These instance selection
methods are most commonly used for cleaning the dataset
reducing it size by removing redundant information. They
discovered that the relation between the original and reduced
datasets can be used as a landmark to lower the error rates
when predicting classifier performance.</p>
      <p>Another approach to predicting classifier performance is to
extrapolate results from a smaller dataset to simulate the
performance of a larger dataset. [Basavanhally et al., 2010]
attempted to predict classifier performance in the field of
computer aided diagnostics, where data is very often limited in
quantity. Their experiments showed that using a repeated
random sampling method on small datasets to make predictions
on a larger set tended to have high error rates and should not
be generalized as holding true when large amounts of data
become available. Later, [Basavanhally et al., 2015] improved
this method by utilizing it together with cross-validation
sampling strategy, which resulted in lower error rates.</p>
      <p>In the field of NLP, [Johnson et al., 2018] applied the
extrapolation method to document classification using the
fastText classifier. They discovered that biased power law model
with binomial weights works as a good baseline extrapolation
model for NLP tasks.</p>
      <p>Instead of concentrating on meta information of the dataset
or performance simulation, our research directly targets
feature engineering and the relation between the available
feature space and classifier performance. This novel method that
can be utilized together with the existing methods to better
estimate the performance of different classifiers.
2.2</p>
      <sec id="sec-2-1">
        <title>Feature Density</title>
        <p>The concept of Feature Density (FD) was introduced by
[Ptaszynski et al., 2017] based on the notion of Lexical
Density [Ure, 1971] from linguistics. It is a score representing
an estimated measure of content per lexical units for a given
corpus, calculated as the number of all unique words divided
by the number of all words in the corpus. The score is called
Feature Density as it also includes other features, like
partsof-speech or dependency information, in addition to words.</p>
        <p>In this research, after calculating FD for all applied dataset
preprocessing methods we calculated Pearson’s correlation
coefficient ( -value) between dataset generalization (FD) and
classifier results (F-scores). If ideal ranges of FD can be
identified, or FD has a positive or negative correlation with
classifier performance, it could be useful in comparatively
estimating the performance of various classifiers. For
example, [Ptaszynski et al., 2017] showed that CNNs benefit
from higher FD while other classifiers’ score was usually
higher when using lower FD datasets. This suggests that it
could be possible to improve the performance of CNNs by
increasing the FD of the applied dataset, while other
classifiers could achieve higher scores by lowering FD [Ptaszynski
et al., 2017].</p>
        <p>In practice, we attempt to estimate what feature
engineering methods can achieve the highest performance for
different models in different languages. The method lets us ignore
redundant feature sets for a particular classifier or language
and only keep the ones with the highest performance
potential without actually training any models.
2.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Linguistically-backed Preprocessing</title>
        <p>Almost without exception, the word embeddings are learned
from pure tokens (words) or lemmas (unconjugated forms of
words). This is also the case with the recently popularized
pre-trained language models like BERT [Devlin et al., 2018].
To the best of our knowledge, embeddings backed with
linguistic information have not yet been researched extensively,
with only a handful of related work attempting to explore the
subject [Levy and Goldberg, 2014; Komninos and
Manandhar, 2016; Cotterell and Schu¨tze, 2019].</p>
        <p>To further investigate the potential of capturing deeper
relations between lexical items and structures and to filter out
redundant information, we propose to preserve the
morphological, syntactic and other types of information by adding
linguistic information to the pure tokens or lemmas. This
means, for example, including parts-of-speech or dependency
information within the used lexical features. These
combinations would then be used to train the word embeddings. The
method could be later applied to the pre-training of huge
language models to possibly improve their performance. The
preprocessing methods are described in-depth in section 3.2.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Dataset and Learners</title>
      <sec id="sec-3-1">
        <title>Dataset</title>
        <p>We tested the concept of FD on the Kaggle Formspring
Dataset for Cyberbullying Detection [Reynolds et al., 2011].
However, the original dataset had a problem of being
annotated by laypeople, whereas it has been pointed out before
that datasets for topics such as online harassment and
cyberbullying should be annotated by experts [Ptaszynski and
Masui, 2018]. Therefore in our research we applied the
version of the dataset after re-annotation with the help of highly
trained data annotators with sufficient psychological
background to assure high quality of annotations [Ptaszynski et</p>
      </sec>
      <sec id="sec-3-2">
        <title>Element type</title>
        <p>Number of samples
Number of CB samples
Number of non-CB samples
Number of all tokens
Number of unique tokens
Avg. length (chars) of a post (Q+A)
Avg. length (words) of a post (Q+A)
Avg. length (chars) of a question
Avg. length (words) of a question
Avg. length (chars) of a answer
Avg. length (words) of a answer
Avg. length (chars) of a CB post
Avg. length (words) of a CB post
Avg. length (chars) of a non-CB post
Avg. length (words) of a non-CB post
Value
al., 2018]. Cyberbullying is a phenomenon observed in many
SNS. It is defined as using online means of communication to
harass and/or humiliate individuals. This can include slurry
comments about someone’s looks or personality or spreading
sensitive or false information about individuals. This
problem has existed throughout the time of communication via
Internet between people but has grown extensively with the
advent of communication devices that can be used
on-thego like smartphones and tablets. Users’ realization of the
anonymity of online communications is one of the factors that
make this activity attractive for bullies since they rarely face
consequences of their improper behavior [Bull, 2010]. The
problem has been growing with the popularity of SNS.</p>
        <p>Table 1 reports some key statistics of the current
annotation of the dataset. The dataset contains approximately 300
thousand of tokens. There were no visible differences in
length between the posted questions and answers (approx. 12
words). On the other hand, the harmful (CB) samples were
usually slightly shorter than the non-harmful (non-CB)
samples (approx. 23 vs. 25 words). The number of harmful
samples was small, amounting to 7%, which roughly reflects the
amount of profanity on SNS [Ptaszynski and Masui, 2018].
3.2</p>
      </sec>
      <sec id="sec-3-3">
        <title>Preprocessing</title>
        <p>In order to train the linguistically-backed embeddings, we
first preprocessed the dataset in various ways, similarly to
[Ptaszynski et al., 2017]. This was done to verify the
correlation between the classification results and Feature
Density (FD) and to verify the performance of various versions of
the proposed linguistically-backed embeddings. The
preprocessing was done using spaCy NLP toolkit (https://spacy.io/).
After assembling combinations from the listed preprocessing
types, we ended up with a total of 68 possible preprocessing
methods for the experiments. The FDs for all separate
preprocessing types used in this research were shown in Table
2.</p>
        <p>• Tokenization: includes words, punctuation marks, etc.
separated by spaces (later: TOK).
• Lemmatization: like the above but with generic
(dictionary) forms of words (“lemmas”) (later: LEM).
• Parts of speech (separate): parts of speech information
is added in the form of separate features (later: POSS).
• Parts of speech (combined): parts of speech
information is merged with other applied features (later: POS).
• Named Entity Recognition (without replacement):
information on what named entities (private name of a
person, organization, numericals, etc.) appear in the
sentence are added to the applied word (later: NER).
• Named Entity Recognition (with replacement): same
as above but information replaces the applied word
(later: NERR).
• Dependency structure: noun- and verb-phrases with
syntactic relations between them (later: DEP).
• Chunking: like above but without dependency relations
(“chunks”, later: CHNK).
• Stopword filtering: redundant words are filtered out
using spaCy’s stopword list for English (later: STOP)
• Filtering of non-alphabetics: non-alphabetic
characters are filtered out (later: ALPHA)
3.3</p>
      </sec>
      <sec id="sec-3-4">
        <title>Feature Extraction</title>
        <p>We generated a Bag-of-Words language model from each
of the 68 processed dataset versions. This resulted in
separate models for each of the datasets (Bag-of-Words,
Bag-ofLemmas, Bag-of-POS, etc.). Next, we applied a weighting
scheme, term frequency with inverse document frequency or
tf idf .</p>
        <p>When training a Convolutional Neural Network model, the
embeddings were trained as a part of the network for all of the
described datasets. Similarly to other classifiers, we trained a
separate model for each of the 68 datasets (Word/token
Embeddings, Lemmas Embeddings, POS Embeddings, Chunks
Embeddings, etc.). The embeddings were trained as part of
the network using Keras’ embedding layer with random
initial weights, meaning no pretraining was used.
3.4</p>
      </sec>
      <sec id="sec-3-5">
        <title>Classification</title>
        <p>We used two variants of Support Vector Machine [Cortes and
Vapnik, 1995], linear-SVM and linear-SVM with SGD
optimizer. We also used two different solvers for Logistic
Regression (LR), Newton and L-BFGS. We also used both
AdaBoost [Freund and Schapire, 1997] and XGBoost [Chen and
Guestrin, 2016]. Other classifiers applied include Random
Forest [Breiman, 2001], kNN, Na¨ıveBayes, Multilayer
Perceptron (MLP) And Convolutional Neural Network (CNN).</p>
        <p>In this experiment MLP refers to a network using
regular dense layers. We applied an MLP implementation with
Rectified Linear Units (ReLU) as a neuron activation
function [Hinton et al., 2012] and one hidden layer with dropout
regularization which reduces overfitting and improves
generalization by randomly dropping out some of the hidden units
during training [Hinton et al., 2012].</p>
        <p>We applied a CNN implementation with Rectified Linear
Units (ReLU) as a neuron activation function, and max
pooling [Scherer et al., 2010], which applies a max filter to
nonoverlying sub-parts of the input to reduce dimensionality and
Preprocessing type</p>
        <p>Classifier
in effect correct overfitting. We also applied dropout
regularization on penultimate layer. We applied two versions of
CNN. First, with one hidden convolutional layer containing
128 units. The second version consisted of two hidden
convolutional layers, containing 128 feature maps each, with 4x4
size of patch and 2x2 max-pooling, and Adaptive Moment
Estimation (Adam), a variant of Stochastic Gradient Descent
[LeCun et al., 2012].
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <sec id="sec-4-1">
        <title>Setup</title>
        <p>The preprocessed dataset provides 68 separate datasets and
the experiment was performed once for each preprocessing
type. Each of the classifiers (sect. 3.4) were tested on each
version of the dataset in a 10-fold cross validation
procedure. This gives us an opportunity to evaluate how
effective different preprocessing methods are for each classifier.
As the dataset was not balanced, we oversampled the
minority class using Synthetic Minority Over-sampling Technique
(SMOTE) [Chawla et al., 2002]. The preprocessing methods
represent a wide range of Feature Densities, which can be
used to evaluate the correlation with classifier performance.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Effect of Feature Density</title>
        <p>We analyzed the correlation of Feature Density with each
of the classifiers using the proposed preprocessing methods.
The results are represented in Table 5. As the results for
using only parts-of-speech tags, which had the lowest FD by
far, were extremely low (close to a coinflip). Thus, we can
say that POS tags alone do not contain enough information to
successfully classify the entries.</p>
        <p>After excluding the preprocessing methods that only used
POS tags, we can see that all classifiers, except CNNs have
a strong negative correlation with Feature Density. So these
classifiers seem to have a weaker performance if a lot of
linguistic information is added, and the best results being
usually within the range of :05 to :15 FD depending on the
classifier. This range includes 38 of the 68 preprocessing
methods (Table 2), meaning that the total training time could be
reduced by around 40-50%. This can be seen from, for
example, the highest performing classifier, SVM with SGD
optimizer (Figure 1), where the maximum classifier performance
starts high at around :05 and slowly falls until :14 after which
there is a noticeable drop. The performance only falls further
as the FD rises.</p>
        <p>For CNNs however, there was a very weak positive or no
correlation between FD and the classifier performance, with
the higher FD datasets performing equally or even slightly
better when comparing to the low FD datasets. Taking a look
at one layer CNN’s performance, which was better than the
CNN with two layers, we can see from Figure 1 that the
maximum performance starts at a moderate level and stays more
stable throughout the whole range of feature densities. The
most potential ranges of FD are between :05 to :1 and after
:45. The potential training time reduction seems to be similar,
around 40-50% The reduction in training time could be
especially important when considering demanding models like
Neural Networks.</p>
        <p>The results suggest that for non-CNN classifiers there is
no need to consider preprocessings with a high FD, such as
chunking or dependencies, as they had a considerably lower
performance. The performance seems to start falling rapidly
at around F D = :15 with most of the classifiers. For CNNs,
high performances were recorded on both low and high FDs.
This means that there is potential in the higher FD
preprocessing types, namely, dependencies for CNNs.</p>
        <p>The reason for CNNs relatively low performance could be
explained by the relatively small size of the dataset, especially
when considering the amount of actual cyberbullying entries,
as adding even a second layer to the network already caused
a loss of the most valuable features and ended up
degrading performance. With such small amount of data, it doesn’t
seem useful to train deep learning models to solve the
classification problem. Still, the dependency based features are
showing potential with CNNs. With a considerably larger
dataset and more computational power, it could be possible
to outperform other classifiers and the usage of tokens with
dependency based features when using deep learning.</p>
        <p>The experiments show that changing Feature Density in
moderate amount can yield good results when using other
classifiers than CNNs. However, excessive changes to
either too low or too high always showed diminishing results.
The treshold was in all cases approximately between 50%
and 200% of the original density (TOK), most optimal FDs
only slightly varying with each classifier. The exception
being Random Forest [Breiman, 2001], which showed a clear
spike at around :12 FD. As the usage of high Feature Density
datasets showed potential with CNNs, their usage needs to
be confirmed in future research. Also, more exact ideal
feature densities need to be confirmed for each classifier using
datasets of different sizes and fields to make a more accurate
ranking of classifiers by FD possible.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Analysis of Linguistically-backed</title>
      </sec>
      <sec id="sec-4-4">
        <title>Preprocessing</title>
        <p>From the results it can be seen that most of the classifiers
scored highest on pure tokens. CNNs also performed quite
well on the dependency-based preprocessings. Using lemmas
usually got slightly lower scores than tokens probably due to
information loss. Chunking got low performance overall and
was clearly outperformed by dependency-based features in
CNNs. Using only POS tags achieved very low performance
and thus it should be only used as a supplement to other
methods.</p>
        <p>Stopword filtering seemed to be the one of the most
effective preprocessing techniques for traditional classifiers, which
can be seen from Table 3 as it was used in the majority of
the highest scores. The problem with stopwords was that
the scores fluctuated a lot, having both low and very high
scores and scoring high mostly with Logistic Regression and
all of the tree based classifiers. An important thing to note is
that the preprocessing method had extremely polarized
performance with CNNs, scoring either very high or low.
Overall, stopwords yielded the most top scores of any
preprocessing method considering all the classifiers.</p>
        <p>Another very effective preprocessing method was Parts of
Speech merging (POS), which achieved high performance
overall when added to TOK or LEM. The method also got
the highest scores with multiple classifiers, especially SVMs.
Adding parts-of-speech information to the respective words
achieved a higher score than using them as a separate feature.
This keeps the information directly connected to the word
itself, which seems a better option when preserving
information.</p>
        <p>Using Named Entity Recognition reduced the classifier
performance most of the time, only achieving a high score
with one classifier, Newton-LR. The performance of using
NER seemed clearly inferior compared to stopwords or POS
information. Replacing words with their NER information
seems to cause too much information loss and reduces the
performance when comparing to plain tokens. Attaching
NER information to the respective words did not improve the
performance in most cases but still performed better than
replacement. These results are different to [Ptaszynski et al.,
2017], who noticed that NER helped most of the times for
cyberbullying (CB) detection in Japanese. This could come
from the fact that CB is differently realized in those
languages. In Japan, revealing victim’s personal information,
or “doxxing” is known to be one of the most often used form
of bullying, thus NER, which can pin-point information such
as address or phone number often help in classification, while
this is not the case in English.</p>
        <p>Filtering out non-alphabetic characters also reduced the
classifier performance most of the time and also got a high
score with only one classifier, kNN, which was the weakest
classifier overall. Non-alphabetic tokens seem to carry useful
information, at least in the context of cyberbullying
detection, as removing them reduced the performance comparing
to plain tokens due to information loss.</p>
        <p>Trying to generalize the feature set ended up lowering the
results in most cases with the exception of the very high
scores of stopword filtering using traditional classifiers. This
would mean that the stopword filter sometimes succeeded
in removing noise and outliers from the dataset while other
generalization methods ended up cutting useful information.
Adding information to tokens could be useful in some
scenarios as was shown with parts-of-speech tags and using
dependency information with CNNs, although using NER was not
so successful. Any kind of generalization attempt resulted in
a lower performance with CNNs, which shows their ability
to assemble more complex patterns from tokens and relations
that are unusable by other classifiers.</p>
        <p>An interesting discovery is that using raw tokens only
rarely resulted in the best performance considering the
proposed feature sets. This can be seen from Tables 3 and 5. This
proves the effectiveness of using linguistics-based feature
engineering instead of directly using words as features. Also,
the performance of one-layer CNN increased significantly
when using linguistic embeddings, from 0.659 (TOK) F-score
to 0.741 (DEPSTOP). The high scores of dependency-based
feature sets indicate that structural information could be
important.</p>
        <p>In order to compare the usage of linguistic preprocessing
to modern text classifiers, we fine-tuned RoBERTa [Liu et al.,
2019] on the dataset. This showed an F-score of 0.797, which
is similar to the highest scores by other models using our
method. Actually, the best score by SGD SVM is 0.798 which
is slightly higher. It is fascinating that a simple method like
SVM can outperform a complex modern text classifier when
using the right feature set. This shows that traditional, more
simple models should not be underestimated as with correct
preparations, they can achieve a similar performance as
stateof-the art models and require much less computational power.
Possibly, the performance of pretrained language models like
RoBERTa could also be increased by feature engineering and
applying embeddings with linguistic information. This needs
to be explored further in future research.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.4 Environmental Effect</title>
        <p>If the weaker feature sets were to be left out, the power
savings are approximately 35Wh calculated from Table 4 for
training the SGD SVM classifier, which is not very much.
But the classifier was very power efficient to train to begin.
A more impressive result can be seen with CNN, where the
power savings are approximately 21kWh, which is
considerably more compared to SVM.</p>
        <p>In order demonstrate the environmental effect of the
method, we will look at the CNN model and its power savings
(21kWh). According to European Environmental Agency
(EEA) 1, the average CO2 emissions of electricity generation
was 275 g CO2e/kWh in 2019. Thus the greenhouse gases
emitted during the training of CNN could be estimated to be
5.8 kg CO2e. For comparison, the average new passenger
car in the European Union in 2019, according to EEA, emits
around 122 g CO2e per kilometer driven. So when training a
simple CNN model, if we calculate the feature densities and
leaving out the weaker feature sets before training, we could
save as much as driving a new car for almost 50 kilometers in
emissions.</p>
        <p>Instead of having to run all of the experiments, it could be
useful to first discard the FD ranges of the overall weakest
feature sets. Then running a small subset of the experiments
with a set interval between preprocessing type feature
densities, look for the FD range with a high performance and
iterate around it by running more experiments with similar
feature densities in order to find the maximum performance.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper we presented our research on Feature Density
and linguistically-backed preprocessing methods, applied in
dialog classification and cyberbullying detection. Both
concepts are relatively novel to the field. We studied the effect of
FD in reducing the number of required experiments iterations
and analyzed the usage of different linguistically-backed
preprocessing methods in the context of CB detection.</p>
      <p>The results indicate that for non-CNN classifiers, there is
an ideal Feature Density that slightly differs between each
classifier. This can be taken into account in future
experiments in order to save time and computational power when
running experiments. For CNNs however, there is almost no
correlation between FD and classifier and thus the higher FD
datasets should also be considered when trying to achieve the
best performance.</p>
      <p>Using plain tokens to keep the original words and their
forms and reducing noise with stopword filtering yielded the
best results in general. With some classifiers, adding extra
information in the form of POS tags also proved useful. For
convolutional neural networks, using dependency based
information showed potential and their effect needs to be
confirmed in future research.</p>
      <p>Although the environmental effect of the method does not
seem very significant here, one has to keep in mind that the
tested models were quite simple. Assuming that the method
would work with other datasets and more resource intensive
classifiers, the savings could be very significant. It could
be useful to only run a subset of the experiments and
iterate around the most probable performance peak in order to
find the maximum performance.</p>
      <p>In the near future we will also confirm the potential of
linguistically-backed preprocessing and Feature Density for
other applications and languages. The research further
suggests that adding linguistic preprocessing can improve the
performance of classifiers, which needs to be also confirmed
on current state of the art language models.
[Ure, 1971] J Ure. Lexical density and register
differentiation. Applications of Linguistics, page 443–452, 1971.
(a) Alphabetic filtering (red) vs others (blue)
(b) Alphabetic filtering (red) vs others (blue)
(c) NER (red) vs others (blue)
(d) NER (red) vs others (blue)
(e) POS (red) vs others (blue)
(f) POS (red) vs others (blue)
(g) Stopword filtering (red) vs others (blue)
(h) Stopword filtering (red) vs others (blue)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Basavanhally et al.,
          <year>2010</year>
          ]
          <string-name>
            <given-names>A.</given-names>
            <surname>Basavanhally</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Doyle</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Madabhushi</surname>
          </string-name>
          .
          <article-title>Predicting classifier performance with a small training set: Applications to computer-aided diagnosis and prognosis</article-title>
          .
          <source>In 2010 IEEE International Symposium on Biomedical Imaging: From Nano to Macro</source>
          , pages
          <fpage>229</fpage>
          -
          <lpage>232</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Basavanhally et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Ajay</given-names>
            <surname>Basavanhally</surname>
          </string-name>
          , Satish Viswanath, and
          <string-name>
            <given-names>Anant</given-names>
            <surname>Madabhushi</surname>
          </string-name>
          .
          <article-title>Predicting classifier performance with limited training data: Applications to computer-aided diagnosis in breast and prostate cancer</article-title>
          .
          <source>PLOS ONE</source>
          ,
          <volume>10</volume>
          (
          <issue>5</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          ,
          <year>05 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[Bensusan and Kalousis</source>
          , 2001]
          <string-name>
            <given-names>Hilan</given-names>
            <surname>Bensusan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alexandros</given-names>
            <surname>Kalousis</surname>
          </string-name>
          .
          <article-title>Estimating the predictive accuracy of a classifier</article-title>
          . In Luc De Raedt and Peter Flach, editors,
          <source>Machine Learning: ECML 2001</source>
          , pages
          <fpage>25</fpage>
          -
          <lpage>36</lpage>
          , Berlin, Heidelberg,
          <year>2001</year>
          . Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[Blachnik</source>
          , 2017]
          <string-name>
            <given-names>Marcin</given-names>
            <surname>Blachnik</surname>
          </string-name>
          .
          <article-title>Instance selection for classifier performance estimation in meta learning</article-title>
          .
          <source>Entropy</source>
          ,
          <volume>19</volume>
          :
          <fpage>583</fpage>
          , 11
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Breiman</source>
          , 2001]
          <string-name>
            <given-names>Leo</given-names>
            <surname>Breiman</surname>
          </string-name>
          .
          <article-title>Random forests</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>45</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          ,
          <year>Oct 2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[Bull</source>
          ,
          <year>2010</year>
          ]
          <string-name>
            <given-names>Glen</given-names>
            <surname>Bull</surname>
          </string-name>
          .
          <article-title>The always-connected generation</article-title>
          .
          <source>Learning and Leading with Technology</source>
          ,
          <volume>38</volume>
          :
          <fpage>28</fpage>
          -
          <lpage>29</lpage>
          ,
          <year>November 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Chawla et al.,
          <year>2002</year>
          ]
          <string-name>
            <given-names>Nitesh</given-names>
            <surname>Chawla</surname>
          </string-name>
          , Kevin Bowyer, Lawrence Hall, and
          <string-name>
            <given-names>Philip</given-names>
            <surname>Kegelmeyer</surname>
          </string-name>
          . Smote:
          <article-title>Synthetic minority over-sampling technique</article-title>
          .
          <source>J. Artif. Int. Res.</source>
          ,
          <volume>16</volume>
          (
          <issue>1</issue>
          ):
          <fpage>321</fpage>
          -
          <lpage>357</lpage>
          ,
          <year>June 2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Chen and Guestrin</source>
          , 2016]
          <string-name>
            <given-names>Tianqi</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Guestrin</surname>
          </string-name>
          .
          <article-title>Xgboost: A scalable tree boosting system</article-title>
          .
          <source>CoRR, abs/1603.02754</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Cortes and Vapnik</source>
          , 1995]
          <string-name>
            <given-names>Corinna</given-names>
            <surname>Cortes</surname>
          </string-name>
          and
          <string-name>
            <given-names>Vladimir</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <article-title>Support-vector networks</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>20</volume>
          (
          <issue>3</issue>
          ):
          <fpage>273</fpage>
          -
          <lpage>297</lpage>
          ,
          <year>Sep 1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Cotterell and Schu¨tze, 2019]
          <article-title>Ryan Cotterell and Hinrich Schu¨tze. Morphological word embeddings</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1907</year>
          .02423,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Devlin et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          . Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[Freund and Schapire</source>
          , 1997]
          <string-name>
            <given-names>Yoav</given-names>
            <surname>Freund</surname>
          </string-name>
          and
          <string-name>
            <given-names>Robert E</given-names>
            <surname>Schapire</surname>
          </string-name>
          .
          <article-title>A decision-theoretic generalization of on-line learning and an application to boosting</article-title>
          .
          <source>Journal of Computer and System Sciences</source>
          ,
          <volume>55</volume>
          (
          <issue>1</issue>
          ):
          <fpage>119</fpage>
          -
          <lpage>139</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[Gama and Brazdil</source>
          , 1995]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gama</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Brazdil</surname>
          </string-name>
          .
          <article-title>Characterization of classification algorithms</article-title>
          . In Carlos PintoFerreira and Nuno J. Mamede, editors,
          <source>Progress in Artificial Intelligence</source>
          , pages
          <fpage>189</fpage>
          -
          <lpage>200</lpage>
          , Berlin, Heidelberg,
          <year>1995</year>
          . Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <source>[Hinduja and Patchin</source>
          , 2010]
          <string-name>
            <given-names>Sameer</given-names>
            <surname>Hinduja</surname>
          </string-name>
          and
          <string-name>
            <given-names>Justin</given-names>
            <surname>Patchin</surname>
          </string-name>
          .
          <article-title>Bullying, cyberbullying, and suicide</article-title>
          .
          <source>Archives of suicide research : official journal of the International Academy for Suicide Research</source>
          ,
          <volume>14</volume>
          :
          <fpage>206</fpage>
          -
          <lpage>21</lpage>
          ,
          <year>07 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Hinton et al.,
          <year>2012</year>
          ]
          <string-name>
            <given-names>Geoffrey E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          , Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <article-title>Improving neural networks by preventing coadaptation of feature detectors</article-title>
          .
          <source>CoRR, abs/1207.0580</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Johnson et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter Anderson</surname>
            ,
            <given-names>Mark</given-names>
          </string-name>
          <string-name>
            <surname>Dras</surname>
            , and
            <given-names>Mark</given-names>
          </string-name>
          <string-name>
            <surname>Steedman</surname>
          </string-name>
          .
          <article-title>Predicting accuracy on large datasets from smaller pilot data</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pages
          <fpage>450</fpage>
          -
          <lpage>455</lpage>
          , Melbourne, Australia,
          <year>July 2018</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>[King</surname>
          </string-name>
          et al.,
          <year>1995</year>
          ]
          <string-name>
            <given-names>R. D.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Sutherland</surname>
          </string-name>
          . Statlog:
          <article-title>Comparison of classification algorithms on large real-world problems</article-title>
          .
          <source>Applied Artificial Intelligence</source>
          ,
          <volume>9</volume>
          (
          <issue>3</issue>
          ):
          <fpage>289</fpage>
          -
          <lpage>333</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>[Komninos and Manandhar</source>
          , 2016]
          <string-name>
            <given-names>Alexandros</given-names>
            <surname>Komninos</surname>
          </string-name>
          and
          <string-name>
            <given-names>Suresh</given-names>
            <surname>Manandhar</surname>
          </string-name>
          .
          <article-title>Dependency based embeddings for sentence classification tasks</article-title>
          .
          <source>In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pages
          <fpage>1490</fpage>
          -
          <lpage>1500</lpage>
          , San Diego, California, June 2016.
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [LeCun et al.,
          <year>2012</year>
          ] Yann LeCun, Leon Bottou, Genevieve Orr., and
          <string-name>
            <surname>Klaus-Robert Muller</surname>
          </string-name>
          .
          <source>Efficient BackProp</source>
          , pages
          <fpage>9</fpage>
          -
          <lpage>48</lpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[Levy and Goldberg</source>
          , 2014]
          <string-name>
            <given-names>Omer</given-names>
            <surname>Levy</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yoav</given-names>
            <surname>Goldberg</surname>
          </string-name>
          .
          <article-title>Dependency-based word embeddings</article-title>
          .
          <source>In ACL</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Liu et al.,
          <year>2019</year>
          ] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .11692,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [Michie et al.,
          <year>1995</year>
          ]
          <string-name>
            <given-names>Donald</given-names>
            <surname>Michie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Spiegelhalter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , and John Campbell, editors.
          <source>Machine Learning, Neural and Statistical Classification. Ellis Horwood</source>
          , USA,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>[Ptaszynski and Masui</source>
          , 2018]
          <string-name>
            <given-names>Michal</given-names>
            <surname>Ptaszynski</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fumito</given-names>
            <surname>Masui</surname>
          </string-name>
          .
          <source>Automatic Cyberbullying Detection: Emerging Research and Opportunities. IGI Global</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Ptaszynski et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Michal</given-names>
            <surname>Ptaszynski</surname>
          </string-name>
          , Juuso Kalevi Kristian Eronen, and
          <string-name>
            <given-names>Fumito</given-names>
            <surname>Masui</surname>
          </string-name>
          .
          <article-title>Learning deep on cyberbullying is always better than brute force</article-title>
          .
          <source>In LaCATODA 2017 CEUR Workshop Proceedings, page 3-10</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Ptaszynski et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Michał</given-names>
            <surname>Ptaszynski</surname>
          </string-name>
          , Gniewosz Leliwa, Mateusz Piech, and Aleksander Smywin´
          <fpage>ski</fpage>
          -Pohl.
          <source>Cyberbullying detection - technical report 2/</source>
          <year>2018</year>
          , department of computer science agh, university of science and technology,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [Reynolds et al.,
          <year>2011</year>
          ]
          <string-name>
            <given-names>Kelly</given-names>
            <surname>Reynolds</surname>
          </string-name>
          , April Edwards, and
          <string-name>
            <surname>Lynne Edwards.</surname>
          </string-name>
          <article-title>Using machine learning to detect cyberbullying</article-title>
          .
          <source>Proceedings - 10th International Conference on Machine Learning and Applications, ICMLA</source>
          <year>2011</year>
          ,
          <volume>2</volume>
          ,
          <string-name>
            <surname>12</surname>
          </string-name>
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [Scherer et al.,
          <year>2010</year>
          ]
          <string-name>
            <given-names>Dominik</given-names>
            <surname>Scherer</surname>
          </string-name>
          , Andread Mu¨ller, and
          <string-name>
            <given-names>Sven</given-names>
            <surname>Behnke</surname>
          </string-name>
          .
          <article-title>Evaluation of pooling operations in convolutional architectures for object recognition</article-title>
          .
          <source>In ICANN 2010 Proceedings, Part III</source>
          , pages
          <fpage>92</fpage>
          -
          <lpage>101</lpage>
          ,
          <year>01 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>[Schwartz</surname>
          </string-name>
          et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Roy</given-names>
            <surname>Schwartz</surname>
          </string-name>
          , Jesse Dodge, Noah A.
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>and Oren</given-names>
          </string-name>
          <string-name>
            <surname>Etzioni</surname>
          </string-name>
          .
          <source>Green AI</source>
          . CoRR, abs/
          <year>1907</year>
          .10597,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [Strubell et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Emma</given-names>
            <surname>Strubell</surname>
          </string-name>
          , Ananya Ganesh, and
          <string-name>
            <surname>Andrew McCallum</surname>
          </string-name>
          .
          <article-title>Energy and policy considerations for deep learning in NLP</article-title>
          . CoRR, abs/
          <year>1906</year>
          .02243,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <article-title>(i) TOK (red), LEM (green), CHNK (yellow), DEP (blue) (j) TOK (red), LEM (green), CHNK (yellow), DEP (blue) Figure 1: FD &amp; F1 score for SGD SVM (left) and CNN1 (right) 14</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>