<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Systems and Knowledge Discovery (FSKD 2007)</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">2374-8486</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1145/1143844.1143865</article-id>
      <title-group>
        <article-title>An Evaluation of Machine Learning Methods for Predicting Flaky Tests</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Azeem Ahmad</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ola Lei er</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kristian Sandahl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Linköping University</institution>
          ,
          <addr-line>581 83 Linköping</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>1</volume>
      <fpage>161</fpage>
      <lpage>168</lpage>
      <abstract>
        <p>In this paper we have investigated as a means of prevention the feasibility of using machine learning (ML) classi ers for aky test prediction in project written with Python. This study compares the predictive accuracy of the three machine learning classi ers (Naive Bayes, Support Vector Machines, and Random Forests) with each other. We compared our ndings with the earlier investigation of similar ML classi ers for projects written inJava. Authors in this study investigated if test smells are good predictors of test akiness. As developers need to trust the predictions of ML classi ers, they wish to know which types of input data or test smells cause more false negatives and false positives. We concluded that RF performed better when it comes to precision (&gt; 90%) but provided very low recall (&lt; 10%) as compared to NB (i.e., precision &lt; 70% and recall &gt;30%) and SVM (i.e., precision &lt; 70% and recall &gt;60%).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Improve Software Quality</kwd>
        <kwd>Flaky Test Detection</kwd>
        <kwd>Machine Learning Classi ers</kwd>
        <kwd>Experimentation</kwd>
        <kwd>Test Smells</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>8th International Workshop on Quantitative Approaches to Software
Quality in conjunction with the 27th Asia-Paci c Software
Engineering Conference (APSEC 2020) Singapore, 1st December 2020</p>
      <p>azeem.ahmad@liu.se (A. Ahmad); ola.lei er@liu.se(O. Lei er);
kristian.sandahl@liu.se (K. Sandahl)
0000-0003-3049-1261 (A. Ahmad)</p>
      <p>© 2020 Copyright for this paper by its authors. Use permitted under Creative
CPWErooUrckReshdoinpgs hIStpN:/c1e6u1r3-w-0s.o7r3g CCoEmUmoRns WLiceonrsekAsthtriobuptioPnr4o.0cIneteerdnaitniognasl ((CCC EBYU4R.0)-.WS.org)
contents of a test cases written in Python. We com- categories of classi cation and can nd the best
hyperpared our ndings with what was presented by Pinto plane to partition a sample space [15].
et al. [23]. We looked for evidence if machine learning RF is an ensemble classi cation method (a technique
classi ers are applicable in predicting aky tests and that combines several base models to produce an
optithe results can be generalized to test cases written in mal predictive model) suitable for handling problems
other languages. In addition to this, our unique contri- that involve grouping data into di erent classes. RF
bution is to investigate if test smells are good predic- predicts by using decision trees. Trees are constructed
tors of test akiness. Through manual investigation of during training which can later be used for class
prefalse positives and false negatives, we concluded a list diction. There is a vote associated with each tree and
of test smells that are strong and weak predictors of once the class vote has been produced for all
individtest akiness. We investigated the following research ual trees, the class with the highest vote is considered
questions in this study. to be the output.</p>
      <sec id="sec-1-1">
        <title>2.2. Performance Metrics and</title>
      </sec>
      <sec id="sec-1-2">
        <title>Parameters Tuning</title>
        <p>RQ1: What are the predictive accuracy of Naive Bayes,
Support Vector Machine and Random Forest concerning
aky test detection and prediction?</p>
        <p>RQ2: To what extent the predicting power of machine
learning classi ers vary when applied on software
written in other programming language?</p>
        <p>RQ3: What can we learn about the predictive power of
test smells using machine learning classi ers mentioned
in RQ1?
To evaluate the predictive accuracy of classi ers,
accuracy as the only performance indices is not su cient
[16]. We must consider precision, recall, F1-score, ROC
curve, false positives and false negatives [16]. There is
always some cost associated with false positives and
false negatives. When a non aky test wrongly
classi ed as aky, it gives rise to a some what insigni
cant problem, because an experienced user can bypass
2. Data Set Description and the warning by looking at test case code. In contrast,
Prepossessing when a aky test is wrongly classi ed as non aky test,
this is obnoxious, because it indicates the test suite still
We wrote a script to extract the contents of all test have test cases whose outcome cannot be trusted.
cases from open-source projects, mentioned in Table The experiment started with the implementation of
1. After the test case content’s extraction, we checked simple NB without Laplace smoothing. The results did
which of the test cases, in our database, has been men- not provide good accuracy or precision, because
withtioned in [24] as aky. After this mapping, we nalized out Laplace smoothing, the probability of appearing a
a database with the project name, test case name, test rare test smell (i.e., test smell that was not in the
traincase content and a label. There are many keywords in ing set) in the test set is set to 0, given the formula
the test case code that are irrelevant for the identi
cation of test akiness. We performed extensive data = /
cleaning such as removing punctuation marks, digits where the is the probability that an individual test
and speci c keywords (i.e., int, string, array, assert*) smell is present in a aky test, represents the
numas well as converting text to lower case. ber of times that particular test smell appeared in a
test case and represents the number of times that
2.1. Classifiers: test smell appeared in any test case. Laplace
smoothAn NBC, rst proposed in 1998, is a probabilistic model ing refers to the modi cation in the equation:
which can determine the outcome (i.e., aky or not</p>
        <p>aky) of an instance (i.e., test case) based on the con- = ( + )/( + )
tents of its features (i.e., test case code). In our case, where we set the = 1 so that classi er adds 1 to the
the outcome of NBC is binary. NBC is widely applied probability of rare test smells that were not present in
in classi cation and known to obtain excellent results. the training set. Another step is to identify the
thresh[25]. old (i.e., 0.0 - 1.0) which will increase the predictive
ac</p>
        <p>
          The attractive feature of SVM is that it eliminates curacy of the outcome. As far as SVM was concerned,
the need for feature selections, which makes spam clas- although the feature data set space was linear, we
desi cation easy and faster [14]. SVM deals with the dual cided to use both kernels (i.e., linear and poly) for the
sake of experiment. For random forest, we used ntree
Table 1 peared 1150 times in aky tests and 52 times in
nonOpen-source project names provided by [24] with number aky tests.
of total test cases and flaky tests Figure 1 (A) represents the ROC curve [26]
concerning NBC with Laplace smoothing denoted as NBL with
Project Name TToCtsal Number of FTelasktsy di erent threshold (i.e., from 0.0 to 1.0). We conducted
di erent experiments with di erent training and test
ahpibaecrhnea-tqep4id-0.18 32233517 227834 data sets such as 50/50, 60/40, 70/30, 80/20 and 90/10.
apache-wicket-1.4.20 1250 216 We found similar values for k-fold cross validation.
apache-karaf-2.3 163 102 ROC curve provides a comparison between sensitivity
apache-struts 2.5 2346 60 and speci city helping in organizing classi ers and
viaappaacchhee--lduecrebnye--1s0o.9lr-3.6 3786342 470 sualizing their performance [26]. Sensitivity also known
apache-cassandra-1.1 523 4 as the true positive rate represents a bene t of
predictapache-nutch-1.4 7 4 ing aky tests correctly and speci city also known as
apache-hbase-0.94 29 2 false positive rate represents the cost of predicting non
jafpreaecchhea-hrti-v1e.-01.108.9 222392 02 aky tests as aky tests. In the case of false positive,
developers need to spend e ort and time, just to nd
between 300 - 700 as well as restricting number of vari- out that this is a classi er mistake and the test case is
ables available for splitting at each tree node known as not aky. The optimal target, in the ROC curve, is to
mtry between 25 and 100. rise vertically from origin to the top left corner (higher
true positive rate) as soon as possible because then the
classi er can achieve all true positives with the cost
3. Results of committing a few false positive. The diagonal line,
in Figure 1 (A), represents the strategy of randomly
This section discusses the performance of NBL, SVM guessing the outcome. Any classi er that appears in
and RF with di erent parameters. We compared our the lower right triangle performs worse than a
ranresults with the ndings of Pinto et al. to discuss how dom guessing and we can see that NBL lies in the
upresults vary between Java and Python projects. We per left triangle. Looking at 1 (A), NBL with 70/30 data
also discussed why some classi ers do not perform as partition is suitable to proceed further with 0.4
probexpected and what can we learn about the predictive ability score. NBL, as shown in 1 (A), has stopped
ispower of test smells for test akiness detection and suing positive classi cation (i.e., aky test prediction)
prediction. around 0.76 - 0.87 threshold. After 0.87, it commits
more false positive rate.
3.1. RQ1: Performance of Naive Bayes We tuned di erent parameters in NBL, SVM and RF
Classifier, Support Vector Machine before conducting further experiments. We do not
inand Random Forest tend to provide the results of all experiments because
those experiments were only conducted to nd the
opTable 2 shows the 20 features with the highest infor- timal parameters. The rest (i.e., simple NB, SVM with
mation gain together with their frequency with respect radial and sigmiod kernels) were not included in
furto aky and non- aky tests. We assigned the features ther experiments and discarded. Figure 1 (A-E)
proto the categories presented by Luo et al. in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. We vides comparisons of NBL, SVM-Linear and SVM-Poly
manually traversed the code of aky and non- aky (i.e., di erent kernels) for accuracy, precision, recall
tests to understand the context and how features were and F1-score. All classi ers have achieved good
acused in the tests to assign categories. The top fea- curacies ranging from 93% - 96%. NBL outperformed
ture "conn" appeared in 1361 aky tests and only 15 SVM although the di erence between them is not
dranon- aky tests. This feature is associated with exter- matic. Looking only at the accuracy results of
classinal connection to input/output devices and lies under ers can be deceiving. The important factor for
classithe category of "IO", presented by Luo et. al in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. er selection is to ask the right question and motivate
The second top feature is "double" which appeared in the choice of using speci c classi er such asare we
1190 aky tests and 12 non- aky tests assigned to the interested in detecting aky tests correctly (i.e.,
category of "IO" followed by " oating points opera- precision) or marking a non aky test as aky
tions". The top 3rd feature "tabl" was related to table is not cost e ective (i.e, recall). It is important to
creation during runtime for databases queries and ap- look at precision, recall and accuracy all together for
classi er selection. We can assume that practitioners
A 1.00
0.75
y
iit
v
its0.50
n
e
s
0.25
        </p>
        <p>● ●</p>
        <p>Features
new
assertequ
null
from
string
sclose
true
select
for
fals
not
int
asserttru
tabl*
should
doubl
valu
expr
tcommit
expcolnam
are more interested in precision than recall because
the test suite size, in many organizations, is very large
and they cannot inspect all test cases. In this
particular case, any classi er that correctly ag aky tests
will be encouraged. Precision can answer the question;
"If the lter says this test case is aky, what’s the
probability that it’s aky?”. Figure 1 (C,D) provides
precision and recall values for NBL and SVM. It can be
noticed that NBL precision is increasing (in C) with the
gradual decrease in recall (in D). NBL precision of 65%
dictates that 35% of what was marked as aky was not</p>
        <p>aky. Recall is also lower in NBL as compared to
SVMLinear. SVM-Poly performs worst in terms of precision
and recall as expected due the fact that the input data
set is not polynomial and is well suited for image
processing whereas linear kernel performs better for text
classi cation.</p>
        <p>F1-score, as presented in Figure 1 (E), is the
harmonic mean of precision and recall. F1-score is
useful and informative because of prevalent phenomenon
of class imbalance in text classi cation [27]. NBL is a
suitable candidate although it has a lower F1-score as
compared to SVM-Linear because NBL performs bet- it requires high computation and are very sensitive to
ter with short documents as in our case, the training noisy data [29].
test case consists of 6-15 lines of code [28]. NBL
provides higher precision and lower recall as compared RF provides lesser classi cation error and better
F1to SVM-linear. Another disadvantage of SVM is that scores as compared to decision trees, NBL and SVM.
Accuracy</p>
        <p>F1−Score</p>
        <p>Precision</p>
        <p>Recall
ntree</p>
        <p>The precision, in which we are most interested, is usu- feature deletion 2) it calculates approximation of
imally better than that of SVM and NBL. Authors in [16] portant features for classi cation and 3) it is very
roalso concluded that RF performs better than NBL and bust to noise and outliers [30]. Caruana in [17]
comSVM. The class outcomes are based on "votes" which pared 10 di erent ML classi ers and concluded that
are calculated by each tree in the forest. The outcome decision trees and random forest outperform all other
(i.e., aky or not aky) is selected based on the higher classi ers for spam classi cation.
votes. Figure 2 presents the performance of RF with
respect to selected metrics. mtry represents the number
of variables randomly sampled as candidates at each 3.2. RQ2: Predicting Power of ML
split while ntree is the number of trees to grow. There Classifiers with Respect to Other
is no way to nd an optimalmtry and ntree, so we
experimented with di erent settings, as shown in Figure Languages
2. The mtry has a direct e ect on precision and recall In comparison of our ndings with what was presented
as shown in Figure 2. With an increase in mtry, the by Pinto et al. [23], we observed two di erences. First,
precision is decreasing and recall in increasing; an un- the top 20 frequented features are very di erent in
wanted situation. The optimal value of mtry is 5 where both studies. Only one feature such as "tabl" marked
precision is higher and recall is lower regardless of the as star (*) in Table 2 were similar in both the ndings.
number of trees. The change in mtry did not a ect the However, we observed more features were related to
accuracy but as we discussed earlier, we are not only "IO" output category, as presented in Table 2, which
interested in accuracy but precision too. complemented the ndings of Pinto et al. stating "that</p>
        <p>We performed several experiments to nd optimal all projects manifesting akiness are IO-intensive" [23].
parameters within a classi er before comparing it to Second, we have a very lower precision, recall and
f1other classi ers. After these experiments, we identi- score as compared to Pinto et al. except at a instance
ed three unique classi ers with unique and optimal where random forest provided 0.92 precision. Table
parameters. Since, we are most interested in higher 3 provides detail statistics of precision, recall, and
f1precision, we can see that RF with mtry = 5 and ntree=250 score of three algorithms for comparison. The
algooutperforms all other classi ers only for precision. RF rithms on Python language continuously performed
has achieved more than 90% precision with less than worst contrary to what pinto et al. claimed: "Although
10% recall. We did not achieve high precision (i.e., the studied projects are mostly written in Java, we do not
&gt;90%) in all classi ers. NBL provides unexpected re- expect major di erences in the results if another
objectsults although it holds a good reputation in terms of oriented programming language is used instead, since
detecting spam emails [29]. As compared to NBL and some keywords maybe shared among them" [23].
SVM, RF have distinct qualities such as 1) it can work We speculate that there could be several reasons
aswith thousands of di erent input features without any sociated with these performance reduction such as (1)
Algo.</p>
        <p>Random Forest
Naive Bayes
Support Vector</p>
        <p>Precision
A B
0.92 0.99
0.62 0.93
0.51 0.93</p>
        <p>A
0.4
0.15
0.61</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Lesson Learned</title>
      <p>We implemented the code ourselves using R libraries ML and AI algorithms in recent years have established
for aforementioned classi ers whereas pinto et al. used a good reputation for predicting diseases based on
sympWeka [31] which is an open source machine learn- toms, spam emails based on email contents and many
ing software that can be accessed through a graphi- more. We believe that given a proper input data set
cal user interface, standard terminal applications [32], which clearly distinguishes between aky and non aky
(2) Number of features were very high in the training tests, ML and AI can provide high prediction
capabilsamples and in these cases other models should be con- ities saving e ort, time and resources. We strongly
sidered (i.e., regularized linear regression) that might believe that practitioners, during training of data set,
performed better, (3) the versatility o ered by param- should not consider complete test cases as an input but
eter tunning can become problematic and require spe- only the test codes (i.e., only few lines) that reveal test
cial considerations that can impact the classi ers, etc. akiness.</p>
      <p>
        It is inconclusive that predicting power of machine
3.3. RQ3: Test Smells Analysis and their learning vary with respect to software written in
anPredictive Power for Test Flakiness other languages. Investigation on Java test cases [23]
Detection and Prediction revealed good results while ndings forPython test
cases performed unexpected, thus requiring more
inWe investigated manually di erent cases of true pos- vestigations whether lexical information can be traced
itives (i.e., correct aky test prediction), false positive to akiness.
(i.e., aky test cases marked as non aky) and false Async wait, precision, randomness and IO test smells
negative (i.e., non aky test cases marked as aky) and are string predictors can be predicted by machine
learntrue negatives (correct non aky test prediction) to an- ing classi ers with 100% precision because they only
swer RQ3. We observed that it is not only the fre- exist in test case code and do not require additional
inquency of test smell that makes a test case aky but formation from test class or operating system. Whereas
its co-existence with the class code or external factors all other test smells mentioned in Table 4 are weak
presuch as operating systems or speci c product. For ex- dictors of test akiness and require additional sources
ample, The test smell ’Conditional Test Logic’ as men- of information. We are only aware of test smells that
tioned in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] refers to nested and complex ’if-else’ struc- are investigated in open-source repositories and
literture in the test case. Depending on which branch of ature on test smells in closed-source software is scarce.
’if-else’ is executed, the system under test may require
speci c environment settings. Failing to set the
environment, during di erent executions, will ip the test 5. Discussion and Implication
case outcome, thus making it aky.
      </p>
      <p>After manual investigation of all true/false positives Valuable Indicators for Testers These classi ers can
and true/false negatives, we come up with a list of increase the awareness about aky test vocabulary among
test smells that are strong or weak predictors of test testers. When a new test is added to a test suite, it
akiness, as shown in Table4. Strong predictors refer will be easy to identify whether this test case contains
to those test smells that existed in true positives and speci c test smells that were known to increase test
true negatives cases whereas weak predictors only ex- akiness during previous executions. Testers can take
isted in false negatives and false positives. Test smells advantage of these types of information to reduce test
that are classi ed as weak predictors in this study are akiness. Testers can easily identify test smells that
still useful and can help in identi cation of test ak- are independent of their environment with the help of
iness, but they are not useful with machine learning Table 4.
classi ers because they require additional information Precision Depends on Data Set: In the literature
Test Smell Category
of ML, particularly with spam detection, it is acknowl- crease precision at the expense of recall. When
enedged that precision is a function of the combination countering ’false negative’, an experienced developer,
of the classi er and the data set under investigation. having su cient knowledge of the test smells, will
byClassi er’s precision, in isolation of data set, does not pass the outcome, however, with ’false positive’,
demake sense. The right question is "how precise a clas- velopers are unaware of the fact that test suite still
si er is for a given data set". Unfortunately, there contains aky tests. The motivation of employing ML
is no data available that provides test case contents classi ers (i.e., higher precision - low recall vs balances
and an associated label thus, limiting the use of ad- precision and recall) should be made clear before
provanced ML and AI algorithms. In addition to lack of ceeding with implementation.</p>
      <p>
        aky test data, all research has been conducted with Multi-Factor Input Criteria for Flaky Test
Detecopen-source software and we know a little about what tion: We observed that the ML algorithm should
intest smells are present in closed-source software. Ah- clude di erent sources of information to increase
premad et. al. concluded that there are speci c test smells dictive accuracy. These sources may include 1)
assignthat are associated with the nature of the product [
        <xref ref-type="bibr" rid="ref4">33</xref>
        ] ing speci c weight (i.e., in numbers) to speci c test
known as ’company-speci c’ test smells. The classi er smells or test code, 2) developer’s experience (i.e., new
which are trained on a speci c data set or a domain developer, unaware of the test design guidelines are
cannot be generalized to be used with another data set more likely to write aky tests), 3) company-speci c
or domain. There is a long road ahead to explore the test smells.
best classi er given di erent data sets.
      </p>
      <p>
        Beyond Static Analysis of Test Smells and their
Frequency: ML is capable of incorporating di erent 6. Related Work
sources of information to increase predictive accuracy
as compared to the limited experiment in this study Luo et al., in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], investigated 52 open-source projects
where we only utilized the frequency of test smells in and 201 commits and categorized the causes of test
the test case. During the investigation of the cases of case. Asynchronous wait (45%), concurrency (20%),
’false negative’ and ’false positive’, it has been observed and test order dependency (12%) were found to be the
that the frequency of test smells in the test case will not most common causes of TF. Palomba and Zaidman in
be su cient for prediction. Some test case code (i.e., [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] partially replicated the results presented by Luo et
seeds()) will cancel the e ect of test smell (i.e., ran- al. concluding that the most prominent causes of TF
dom()), no matter how frequent the random() function are asynchronous wait, concurrency, and input output
appears in the test case. Some test smell, even with and network issues. Authors investigated, in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the
single appearance, will weight more than a test smell relationship between smells and TF. Another
empirifor higher frequency. cal study of the root causes of TF in Android Apps was
Precision Vs Recall: When a test suite grows in size, conducted by Thorve et al. [4] by analyzing the
comdevelopers would like any indications of tests that are mits of 51 Apache open-source projects. Thorve et al.
more likely to be aky rather than adopting an ap- [4] complement the results of Luo et al. and Palomba
proach of re-run which of-course is not cost e ective and Zaidman, but they also report two additional test
in terms of time and resources. Developers like to in- smells (user interface and program logic) that are
related to TF in Android Apps. Bell et al. in [
        <xref ref-type="bibr" rid="ref5">34</xref>
        ] and pro- ings of this study for other data set.
posed a new technique called DeFlaker, which
monitors the latest code coverage and marks the test case
as aky if the test case does not execute any of the 8. Conclusion
changes. Another technique called PRADET [
        <xref ref-type="bibr" rid="ref6">35</xref>
        ] does
not detect aky tests directly, rather it uses a system- At the moment of writing this paper, literature is scarce
atic process to detect problematic test order dependen- on test akiness (i.e., root causes, challenges,
mitigacies. These test order dependencies can lead to ak- tion strategies, etc.) which requires signi cant
atteniness. King et al. in [
        <xref ref-type="bibr" rid="ref7">36</xref>
        ] present an approach that tion from researchers and practitioners. We extracted
leverages Bayesian networks for aky test classi ca- aky and non aky test case contents from open source
tion and prediction. This approach considers akiness repositories. We implemented three ML classi ers such
as a decease mitigated by analyzing the symptoms and as Naive Bayes, Support Vector Machine and Random
possible causes. Teams using this technique improved Forest to see if the predictive accuracy can be increased.
CI pipeline stability by as much as 60%. To best of our The authors concluded that only RF performs better
knowledge, no study has been conducted to evaluate when it comes to precision (i.e., &gt; 90%) but the recall
the predictive accuracy of machine learning classi ers is very low (&lt; 10%) as compared to NBL (i.e.,
precithat can help developers in aky test case prediction sion &lt; 70% and recall &gt;30%) and SVM (i.e., precision
and detection. &lt; 70% and recall &gt;60%). The authors concluded that
      </p>
      <p>
        Dutta et al. [
        <xref ref-type="bibr" rid="ref8">37</xref>
        ] and Sjobom [
        <xref ref-type="bibr" rid="ref9">38</xref>
        ] investigated projects predicting accuracy of ML classi ers are strongly
aswritten in Python language to classify test smells that sociated with the lexical information of test cases (i.e.,
increase test akiness. Their study is limited to list test cases written in Java or Python). The authors
inthe test smells and their e ect on test akiness. Our vestigated why other classi ers failed to produce
exstudy worked with the test smells identi ed in [
        <xref ref-type="bibr" rid="ref9">38</xref>
        ]. pected results and concluded that; 1) it is a
combinaPinto et al. evaluated ve machine learning classi ers tion of the test smell and an external environment that
(Random Forest, Decision Tree, Naive Bayes, Support makes a test case aky, and in this study, the
exterVector Machine, and Nearest Neighbour) to generate nal environment was not taken into consideration, 2)
aky test vocabulary written inJava [23]. The con- ML classi ers should not only consider the frequency
cluded that Random Forest and SVM performed very of test smells in the test case but other important test
well with high precision and recall. They concluded codes that have an ability to cancel the e ect of test
that features such as "job", "action", and "services" were smells.
commonly associated with aky tests. We replicated
the similar experiment with di erent programming lan- 9. Acknowledgment
guage and extended the current knowledge by
answering RQ2 and RQ3.
      </p>
      <sec id="sec-2-1">
        <title>We appreciate Linköping University students to provide their expertise to collect aky test data from online repositories.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>7. Validity Threats</title>
      <sec id="sec-3-1">
        <title>The authors in this study selected only those ML classi ers which have established a good reputation of high accuracy in spam detection thus reducing the selection bias.</title>
        <p>The authors in this study reduced the experimenter
bias by performing several experiments with di erent
thresholds (i.e., probability scores, kernels, number of
trees, etc.) before selecting a champion.</p>
        <p>External validity refers to the possibility of
generalizing the ndings, as well as the extent to which the
ndings are of interest to other researchers and
practitioners beyond those associated with the speci c case
being investigated. Since the precision strongly
depends on the data set under investigation, we have an
external validity threat. We cannot generalize the
nd</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hariri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Eloussi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Marinov</surname>
          </string-name>
          ,
          <article-title>An Empirical Analysis of Flaky Tests</article-title>
          ,
          <source>in: Proceedings of the 22Nd ACM SIGSOFT International Symposium on Foundations of Software Engineering, FSE</source>
          <year>2014</year>
          , ACM, New York, NY, USA,
          <year>2014</year>
          , pp.
          <fpage>643</fpage>
          -
          <lpage>653</lpage>
          . URL: http://doi.acm.
          <source>org/10</source>
          .1145/ 2635868.2635920. doi:
          <volume>10</volume>
          .1145/2635868.2635920, eventplace: Hong Kong, China.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Palomba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaidman</surname>
          </string-name>
          ,
          <article-title>Does Refactoring of Test Smells Induce Fixing Flaky Tests?</article-title>
          ,
          <source>in: 2017 IEEE International Conference on Software Maintenance and Evolution (ICSME)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICSME.
          <year>2017</year>
          .
          <volume>12</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Palomba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaidman</surname>
          </string-name>
          ,
          <article-title>The smell of fear: on the relation between test smells and aky tests</article-title>
          ,
          <source>Empirical Software Engineering</source>
          <volume>24</volume>
          (
          <year>2019</year>
          )
          <fpage>2907</fpage>
          -
          <lpage>2946</lpage>
          . URL: https://doi.org/10.1007/ s10664-019-09683-z. doi:
          <volume>10</volume>
          .1007/s10664-019-09683-z. Software in Java, ???? URL: https://www.cs.waikato.ac.nz/ml/ weka/index.html.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ahmad</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <article-title>Lei er, K. Sandahl, Empirical Analysis of Factors and their E ect on Test Flakiness - Practitioners' Perceptions</article-title>
          , arXiv:
          <year>1906</year>
          .00673 [cs] (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1906</year>
          .00673, arXiv:
          <year>1906</year>
          .00673.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Legunsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Eloussi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yung</surname>
          </string-name>
          , D. Marinov, DeFlaker: Automatically Detecting Flaky Tests, in: 2018
          <source>IEEE/ACM 40th International Conference on Software Engineering (ICSE)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>433</fpage>
          -
          <lpage>444</lpage>
          . doi:
          <volume>10</volume>
          .1145/3180155. 3180164.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gambi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zeller</surname>
          </string-name>
          , Practical Test Dependency Detection,
          <source>in: 2018 IEEE 11th International Conference on Software Testing, Veri cation and Validation (ICST)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICST.
          <year>2018</year>
          .
          <volume>00011</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [36]
          <string-name>
            <surname>T. M. King</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Santiago</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Phillips</surname>
            ,
            <given-names>P. J.</given-names>
          </string-name>
          <string-name>
            <surname>Clarke</surname>
          </string-name>
          ,
          <article-title>Towards a Bayesian Network Model for Predicting Flaky Automated Tests</article-title>
          , in: 2018
          <source>IEEE International Conference on Software Quality</source>
          , Reliability and Security
          <string-name>
            <surname>Companion (QRS-C)</surname>
          </string-name>
          ,
          <source>IEEE Comput. Soc</source>
          , Lisbon,
          <year>2018</year>
          , pp.
          <fpage>100</fpage>
          -
          <lpage>107</lpage>
          . doi:
          <volume>10</volume>
          .1109/ QRS-C.
          <year>2018</year>
          .
          <volume>00031</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dutta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Choudhary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Misailovic</surname>
          </string-name>
          ,
          <article-title>Detecting aky tests in probabilistic and machine learning applications</article-title>
          ,
          <source>in: Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis</source>
          ,
          <source>ISSTA</source>
          <year>2020</year>
          ,
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <year>2020</year>
          , pp.
          <fpage>211</fpage>
          -
          <lpage>224</lpage>
          . URL: https://doi.org/10.1145/3395363. 3397366. doi:
          <volume>10</volume>
          .1145/3395363.3397366.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sjöbom</surname>
          </string-name>
          ,
          <source>Studying Test Flakiness in Python Projects : Original Findings for Machine Learning</source>
          ,
          <year>2019</year>
          . URL: http://urn.kb. se/resolve?urn=urn:nbn:se:kth:
          <fpage>diva</fpage>
          -
          <lpage>264459</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>