<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Early Prediction of Test Case Verdict with Bag-of-Words vs. Word Embeddings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wilhelm Meding</string-name>
          <email>wilhelm.meding@ericsson.com</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computing Science, Poznan University of Technology</institution>
        </aff>
      </contrib-group>
      <fpage>61</fpage>
      <lpage>72</lpage>
      <abstract>
        <p>Regression testing is an important testing activity in continuous integration (CI) since it provides con dence that modi ed parts of the system have not adversely a ected its expected behavior. As test suites grow in size over the evolutionary cycle of software, executing large test suites becomes costly and precluding. Since CI provides a large volume of data, machine learning approaches can be used to allow test orchestrators make inferences about which subset of test cases to run at each CI cycle. MeBoTS is a machine-learning-based method that utilizes CI data to improve test case selection (TCS). In order to decide which extraction algorithm is more suitable, we designed and performed an experiment to investigate the e ect of using one of two widely used feature extraction algorithms|Bag of Words (BoW) and Word Embeddings (WE)|on the predictive performance of the MeBoTS classi er. We used strati ed cross-validation and precision and recall measures to evaluate the performance of two machine-learning models trained on the input generated by the feature-extraction algorithms. The results from this experiment show a signi cant di erence between the models' performance scores with a higher mean precision and recall scores for the BoW based classi er. We conclude that the use of BoW allows training a more accurate MeBoTS classi er than WE.</p>
      </abstract>
      <kwd-group>
        <kwd>Machine Learning</kwd>
        <kwd>Verdicts</kwd>
        <kwd>Code Churn</kwd>
        <kwd>Test Case Selection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Continuous integration is used increasingly often in software engineering projects.
Both large and small software companies use this technique to increase the
quality of their products, as continuous integration advocates small increments in
software code and frequent testing. However, one of the challenges in continuous
integration is the need for resources for testing, which needs to be done on 3
every single code commit. In our work, we addressed this problem by de ning a
method for predicting whether a given source code commit should be tested by
a speci c test case. Based on the observation that faults may occur in similar
patterns of source code, our method abstracts these patterns by tokenizing lines
of code and weighing them using their associated frequency. This means that
the numbers of syntax tokens in source code are regarded as predictors to test
case failures. Using these feature vectors (source code tokens) as predictors to
test failures provides a basis for test orchestrators to prioritize and select test
cases.</p>
      <p>
        The prediction algorithms used in our method (MeBoTS, [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) combines
feature extraction from the source code and analysis of test verdicts. It is a modular
method where we can select di erent algorithms for feature extraction and
prediction. However, the selections have e ect on the performance of the predictions
(precision, recall and F1-score). Therefore, in this paper, we study the e ects of
two di erent techniques for feature extraction - Bag-of-words (BoW) and Word
Embeddings (WE). The rst technique is based on statistical analysis of the
frequency of using software code statements. The second technique is a
semantic program code analysis based on neural networks. Both techniques can be
interchanged, but they have di erent complexity and work di erently.
      </p>
      <p>Henceforth, in this paper, we pose the following research question:
RQ: Is there a statistically signi cant di erence between the performance
of the test predictor based on the usage of BoW vs. WE?</p>
      <p>In order to address this question, we design an experiment, where we use
15 di erent sets of code commits as experiment factors. We use the statistical
performance measures of recall and precision as the dependent variables. The
results of the experiment show that the prediction from the WE-based
classier have a statistically signi cant lower precision and recall scores than that
produced by a BoW-based classi er. This means that the simpler method for
making predictions is better in this context.</p>
      <p>The paper is organized as follows: in Section 2 we introduce related work
on code feature extraction techniques in text classi cations; in Section 3 we
introduce background information; in Section 4 we describe all the steps in the
design of our experiment; in Section 5 we provide the results of our experiment;
in Section 6 we introduce some threats to validity and limitations; in Section 7
we describe our recommendations for future implementation of MeBoTS; nally,
in Section 8 we make our conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>Since the ultimate goal of this research is to improve TCS, we start by presenting
an overview of some proposed TCS approaches and explore their drawbacks.
3 Copyright 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
Then we review the literature on studies that have examined the e ectiveness
of BoW and WE in text classi cation.
2.1</p>
      <sec id="sec-2-1">
        <title>Test Case Selection Approaches</title>
        <p>
          Rothermel and Harrold [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] presented an algorithm that employs control
dependence graphs of two program revisions, and used these graphs to select test
cases that may exhibit changed behavior on a modi ed revision of the program.
The algorithm uses two control dependency graphs to compare changes made
between two revisions where each node in the graph contains an actual program
statement. Then it uses a list of test execution history that identi es regions in
the original program that are reached by each test. If any two children nodes
are di erent, then the algorithm computes and returns a subset of test cases
that may have traversed the change in the modi ed version. A limitation in this
approach is that it only selects tests that execute the modi ed statements but
not the actual uses of variables, which leads to the inclusion of unnecessary tests
for regression testing.
        </p>
        <p>
          A TCS technique proposed by Volkolos and Frankl [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] uses textual di erence
between two versions of system source code and analyzes code modi cations. The
method uses a C program to remove stylistic di erences such as comments and
blank lines from the di ed output. Then, it analyzes modi cations by checking
which statement was modi ed and select all test cases that traversed through
the modi ed statement. A limitation in this tool is that it considers code changes
between versions without any semantic analysis, i.e. changes that do not a ect
the behavior of the system under test will trigger tests that traverse the modi ed
statements to be selected.
        </p>
        <p>
          Dynamic slicing based approaches for TCS use slice executions to determine
which subset of test cases should be exercised. Agrawal et al [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] proposed an
array of techniques that use execution slice (i.e., statements in the program that
were executed by a test case) to decide on selective regression tests. The general
idea can be summarized as follow: given a set of test cases t that were exercised
against some execution slices in the original program execution, statements that
were not reached in the control of set t will not a ect the program's output for
the same set t in future revisions. Based on this, they proposed a technique that
required nding execution slices of the program under test given all test cases
in a test suite. Then selecting test cases whose execution slices contain modi ed
statements in the new revision.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Word Embeddings and BoW in Text Classi cation</title>
        <p>
          In a study conducted by Chao et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], the authors examined the di erence
between WE and traditional BoW in clinical text classi cation using an SVM
model. The ndings suggest that WE vectors outperformed BoW when using
1-gram features, whereas no similar conclusion could be drawn from the same
model when using 2-grams features with BoW.
        </p>
        <p>
          Similarly, Enriquez et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] conducted an experiment to compare the e ect
of WE in document classi cation as compared with a BoW based approach.
The experiment's data-set was a collection of texts from Amazon, covering 11
di erent domains. The classi cation results showed that the use of WE is not
su cient on its own to gain a performance improvement, but rather suggested an
integrative approach of both WE and BoW. The evaluation of the performance
was based on the accuracy of three versions of classi ers: a version based on
BoW, a version based on WE, and a version based on both approaches. The
results showed that the combined approach outperformed the classic BoW and
WE in 9 out of 11 experimented domains.
3
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Background</title>
      <sec id="sec-3-1">
        <title>Method Using Bag of Words for Test Case Selection (MeBoTS)</title>
        <p>MeBoTS is a machine learning based model that aims at predicting test case
verdict using historical test execution results and code churns. The term code
churns is used here to refer to source code changes made between two
checkins. The method is comprised of 3 steps, as shown in Fig 1. This section brie y
describes these steps.</p>
        <p>
          Code Churns Extractor (Step 1) The method uses a code churn extractor
program that collects and compiles churns of source code from one or more
repositories. The program expects one input parameter: a time ordered list of historical
test case execution results queried from a database, where each element in the
list is a metadata state representation of a previously run test case. Each state
contains a hash reference that points to a speci c location in Git's history for
the tested check in. The program performs a le comparison utility (di ) across
pairs of consecutive commit hashes in the list using the GitPython library [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
The output is then arranged in a table-like format and written in a csv le,
named as 'Lines of Code'.
        </p>
        <p>Textual Analysis and Features Extraction (Step 2) The second step in the method
is to extract features from the collected code churns (output of step 1) and
transform the source code into a numerical form. For thWe used an open source tool
(cc ex) that utilizes BoW for modelling textual data. The input to cc ex is the
output of the churn extractor in step 1. cc ex uses each line from the code churn
and:
{ creates a vocabulary for all lines (using the bag of words technique, with a
speci c cut-o parameter)
{ creates a token for the words that are seldom used(i.e. fall outside of the
frequency de ned by the cut-o parameter of the bag of words)
{ nds a set of prede ned keywords in each line
{ checks each word in the line to decide if it should be tokenized or if it is a
prede ned feature
This way of extracting information about the source code is new in our approach,
compared to the most common approaches of analyzing code churns. In contrast
with other approaches, MeBoTS recognizes what is written in the code, without
understanding the syntax or semantics of the code. This means that we can
analyze each line of code separately, without the need to compile the code and
without the need to parse it.</p>
        <p>Training and Applying the Classi er Algorithm (Step 3) We exploit the set of
extracted features provided by the textual analyzer in step 2 as the independent
variables and the verdict of the executed test cases as the dependant variable,
which is a binary representation of the execution result (passed or failed). The
MeBoTS method uses a second Python program that utilizes and trains an ML
model to classify test case verdicts. The program reads the BoW vector space le
in a sequence of chunks, merging the extracted feature vectors and the verdicts
vector into a single data frame that gets split into a training and testing set
before it is fed into the models for training.
Word embeddings is one of the approaches used to represent words in a
machinefriendly way. It has been proposed as an alternative to the bag-of-words model
and quickly become the state-of-the-art method used while training neural
networks. In word embeddings, each word is represented as a dense, low-dimensional
oating-point vector. Such vectors are learned from data. Another important
bene t of using word embeddings is that the geometric relationships between
word vectors should re ect the semantic relationships between these words.
Using the word embeddings algorithm for source code classi cation could be
benecial for two main reasons. Firstly, it allows us to represent each of the tokens in
a line of code and feed them to a neural network capable of processing sequences
(e.g., a convolutional neural network). Secondly, it captures similarities between
the roles of tokens in the code without the need for parsing the code.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Design of Experiment</title>
      <p>4.1</p>
      <sec id="sec-4-1">
        <title>Context of the experiment: Collaborating company</title>
        <p>The study has been conducted at an organization, belonging to a large
infrastructure provider company. The organization develops a mature software-intensive
telecommunication network product. The organization consists of several
hundred software developers, organized in several agile teams. The organization is
mature with regard to measuring. For instance every agile team, as well as
leading functions/roles, uses one or more monitors to display status and progress
in various development and devops areas. A well-established and e cient
measurement infrastructure, automatically collects and processes data, and then
distributes the information needed by the organization.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Code Churns and Test Executions Data Collection</title>
        <p>Our data-set comprised of historical test execution results and code churns for
software that has lived and evolved for over a decade at the collaborating
company. The analyzed software was written in the C language and contained a few
million lines of code and a test pool size of over 10k test cases. In this
experiment, our sample data-set comprised of 150k LOC belonging to 12 test cases,
with 46% of the lines belonging to the 'passed' class and 54% to the 'failed' one.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Experiment Subjects</title>
        <p>The subjects of our study are samples of the original data-set. The strati ed
cross-validation technique was used to partition the data-set into 15 di erent
subsets (k=15), such that the representation of the binary strata have
approximately an equal representation across the 15 samples. Each subset consisted of
9200k LOC for validation and approximately 140k LOC for training. The
representation of the binary classes in each fold followed the same distribution of
classes in the original base set with 46% of lines belonging to the 'failed' class
and 54% to the 'passed' class.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Features Extraction with BoW</title>
        <p>
          In this study, we used an open source measurement tool [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] for transforming
the experimental subjects into feature vectors using BoW. The tool starts by
tokenizing each line in a le using white and special characters:
()[]!@#$%&amp;*=;:' "~,&lt;&gt;j/?. Then it counts the frequency of occurrence of the tokens found
in each line. Depending on whether the frequency count of a token exceeds a
lower threshold value, the token gets either selected as a feature or discarded. In
our experiment, we kept the frequency threshold value to its default - 25% and
set the BoW n-gram to 2 to generate features of two adjacent tokens that are
originally separated by white spaces. The resulting space of feature vectors for
each subset comprised of a total of 2248 features.
        </p>
      </sec>
      <sec id="sec-4-5">
        <title>Features Extraction with WE</title>
        <p>
          In our study, we used the Continuous Bag-Of-Words (CBOW) variant of the
Word2Vec word embedding algorithm proposed by Mikolov et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We used
the implementation available in the Gensim library [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. IN CBOW, the word
embeddings are obtained as a side-e ect of training a single-layered neural
network to predict a given word based on other words in its neighborhood, called
window. In our study, we use the window size equal to 5 and generate embedding
vectors of 70 numbers. We trained 15 Word2Vec models on the 15 generated
subsets and used these models to preprocess lines of code in both the validation and
training sets, for each subset respectively. The resulting vectors for each subset
were saved locally so they can be fed as inputs to a neural network classi er.
After training the Word2Vec models, tokens that share similar semantic
orientation are closely placed in the vector space. Fig 2 illustrates an example of how
the tokens in the original data-set (before partitioning) are placed. 4. The gure
was generated using the t-distributed stochastic neighbor embedding (t-SNE)
visualization algorithm in Python. The representation of vectors are plotted in
two dimension for a set of positive and negative words. As the words are spread
across the entire diagram, we can expect that it is possible to nd vectors that
are unique and therefore are good predictors. The processing pipeline used in
this study is presented in Figure 3. In the rst step, a line of code is tokenized.
Then, each token is replaced by its identi er in the vocabulary. In the following
step, we pad each sequence with zeros so all of them contain 55 numbers. The
generated sequences can be provided to a neural network as an input. In its rst
layer, the nn replaces token identi ers with their embedding vectors stored in
the so-called embedding matrix. Therefore, an input sequence of 55 numbers is
transformed into a 70x55 matrix.
4.6
        </p>
      </sec>
      <sec id="sec-4-6">
        <title>Evaluation with Random Forest and Neural Network</title>
        <p>
          All the vectors representing the 15 subsets of code churns were categorized into
two groups, one group containing the BoW vector les and another for the
Word2Vec generated outputs. To evaluate the e ect of WE, we ran 15 trials of
training a random forest model on the BoW vector representations in the BoW
group and another 15 trials for training a CNN model on the les in the WE
group. Our choice of training a random forest model on the BoW vectors is based
on the fact that RF is known for performing well with high dimensional data,
such as those generated by BoW transformations [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. We used the
implementation of Random Forest available in the scikit-learn library [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and the
implementation of CNN in the Keras library [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] to implement the CNN model. Since the
WE model results in multidimensional array, we could not use the combination
of WE and random forest classi er. Our experiments with the CNN architecture
4 Each point in the gure represents a word used in the source code. As the gure
represents the actual code, and due to a non-disclosure agreement with our industrial
partner, words that are not language speci c such as variable and class names are
not visualized in the gure
and the BoW feature extraction, on the other hand, provided results that were
too poor to consider in the paper (BoW does not provide the feature set that is
rich enough for CNN). Therefore, we selected two pairs which are best suited for
each other, rather than forcing the algorithms to work with the feature
extraction technique that is not suitable for them. For each training trial, we recorded
two performance metrics: precision and recall. The architecture of the CNN is
presented in Figure 4. It accepts input as a sequence of vectors, each containing
55 numbers representing identi ers of tokens in the Word2Vec vocabulary. In
the rst layer, these vectors are transformed into matrices (70x55) using word
embeddings (see Figure 3 for details). We use two convolutional layers consisting
of 20 and 16 lters, respectively. The output of each is subjected to maximum
pooling (pooling size = 3) to reduce the dimensionality of the features maps. The
output of the last maximum pooling layer is attened to a vector of 96 numbers
and processed in the dense layer to produce a 1x1 output with the use of the
sigmoid activation function. The precision and recall scores of the two models
across the 15 folds as shown in Table 2.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>To decide whether to use a parametric or non-parametric statistical test, we
checked if the data sample was normally distributed. We plotted frequency
histograms for the precision and recall scores of both models (RF and CNN) for
int \space a
\space
=
\space 10</p>
      <p>;</p>
      <p>Encoding tokens with their identifiers from the vocabulary
1 2 3 2 4 … 0</p>
      <p>…
the 15 folds and examined the distribution of the points. We decided to run a
Shapiro-Wilk test to check if the distribution of points follow a Gaussian curve.
The test results were statistically signi cant for the RF models when using BoW
(Precision with BoW: Test statistic = 0.484, p-value = 0.000, Recall with BoW:
Test statistic = 0.538, p-value = 0.000), which means that the assumption of
normality in both samples can be rejected. Conversely, the results of precision
and recall with WE suggest that the distribution of both samples were normal
(Test statistic = 0.929, p-value = 0.262, Recall with WE: Test statistic = 0.893,
p-value = 0.075). Since the statistical results of the Shapiro-Wilk test suggest
that we have issues with normality in the precision and recall results, we decided
to run a non-parametric test for comparing the di erence between the precision
and recall scores obtained by both models. The Mann whitney rank-based test
was selected as an appropriate method since it can handle skewed data where
the data does not follow a normal distribution.</p>
      <p>We found a signi cant di erence between the precision and recall scores for
both the RF and CNN models. The results of the comparison for the precision
scores showed a test statistics of 12.5 and a p-value below 0.001. Similarly, the
comparison between the recall scores for the same models reported a test
statistics of 32.5 and a p-value of less than 0.001, suggesting a signi cant di erence in
the recall metrics. Table 1 summarizes the mean scores of the 4 performance
met</p>
      <p>Input Embedding layer Convolution 1D
Indices of tokens Embedding vectors layer
padded with zeros of tokens</p>
      <p>Max
Pooling 1D</p>
      <p>Convolution 1D
layer</p>
      <p>Max
Pooling 1D</p>
      <p>Flatten Dropout Dlaeynesre
rics (precision with BoW, precision with WE, recall with BoW, and recall with
WE). The results show a higher mean precision and recall for the RF model than
those obtained by the CNN model. This brings us to believe that using BoW for
features extraction in the MeBoTS is more e ective than WE in this context.
In this paper we have only used a single industrial data-set that belongs to
software, while other industrial software written in di erent languages and
domains might reveal considerably di erent results. This was a design choice as we
wanted to understand the dynamics of test execution and be able to use
statistical methods alongside the machine learning algorithms. However, we are aware
that the generalization of the results for di erent types of systems require further
investigations using tests and churns from di erent systems. Another limitation
comes from the randomness in selecting code churns and test cases without being
ascertained about the nature of their failures. For example, there is a chance of
encountering one or more tests that had failed due to non-functional related
issues, for instance, a machinery failure at execution time. Likewise, the possibility
of having tests that failed due to defects in the test script code and not the base
source code exists. To minimize this threat, we collected data for multiple tests,
thus minimizing the probability of identifying tests which are not representative.
7</p>
    </sec>
    <sec id="sec-6">
      <title>Recommendations</title>
      <p>This section provides our recommendations to practitioners who would like to
use the MeBoTS method:
{ We recommend practitioners to try a variation of classi cation models with
di erent hyper-parameter tuning to assess models' e ectiveness when using
the MeBoTS method.
{ Using a BoW based classi er in the MeBoTS method surpasses that used
with WE, and therefore, we recommend the use of the simple BoW modelling
technique for extracting features from code churns for training a classi er.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion and Future Work</title>
      <p>In this study, we experimented with a set of industrial code churns and test
execution results the e ectiveness of using WE as an alternative feature extraction
approach to BoW in the MeBoTS method. By conducting a total of 30 trials of
training and validating two prediction models on the BoW and WE vector
representations, we empirically compared the di erence between the models' precision
and recall scores. The results con rm with statistical signi cance (p-value less
than 0.001) that modelling code churns with BoW results in higher prediction
performance as compared with a WE-based model. In terms of future work,
more empirical studies with larger industrial data are needed to validate the
e ectiveness of both techniques in the context of MeBoTS. Moreover, training a
series of word embeddings using a di erent variation of parameters such as the
vector and window sizes is needed to draw more conclusive results about the
e ectiveness of WE in this context.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horgan</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krauser</surname>
            ,
            <given-names>E.W.</given-names>
          </string-name>
          , London,
          <string-name>
            <surname>S.A.</surname>
          </string-name>
          :
          <article-title>Incremental regression testing</article-title>
          .
          <source>In: 1993 Conference on Software Maintenance</source>
          . pp.
          <volume>348</volume>
          {
          <fpage>357</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Al-Sabbagh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Staron</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hebig</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meding</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Predicting test case verdicts using textual analysis of commited code churns</article-title>
          .
          <source>EasyChair Preprint no. 1177</source>
          (
          <issue>EasyChair</issue>
          ,
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>Combining svms with various feature selection strategies</article-title>
          .
          <source>In: Feature extraction</source>
          , pp.
          <volume>315</volume>
          {
          <fpage>324</fpage>
          . Springer (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chollet</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , et al.: Keras. https://github.com/fchollet/keras (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Enr quez, F.,
          <string-name>
            <surname>Troyano</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez-Solaz</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>An approach to the use of word embeddings in an opinion classi cation task</article-title>
          .
          <source>Expert Systems with Applications 66</source>
          ,
          <issue>1</issue>
          {
          <issue>6</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kho</surname>
          </string-name>
          , J.:
          <article-title>Why random forest is my favorite machine learning model</article-title>
          . https://towardsdatascience.com
          <article-title>/why-random-forest-is-my-favorite-machinelearning-model-</article-title>
          <string-name>
            <surname>b97651fa3706</surname>
          </string-name>
          ,
          <source>accessed: 2019-06-28</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient estimation of word representations in vector space</article-title>
          .
          <source>arXiv preprint arXiv:1301.3781</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Ochodek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Staron</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bargowski</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meding</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hebig</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>Using machine learning to design a exible loc counter</article-title>
          .
          <source>In: Machine Learning Techniques for Software Quality Evaluation (MaLTeSQuE)</source>
          , IEEE Workshop on. pp.
          <volume>14</volume>
          {
          <fpage>20</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cournapeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duchesnay</surname>
          </string-name>
          , E.:
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          ,
          <volume>2825</volume>
          {
          <fpage>2830</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Rehurek</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sojka</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Software Framework for Topic Modelling with Large Corpora</article-title>
          .
          <source>In: Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks</source>
          . pp.
          <volume>45</volume>
          {
          <fpage>50</fpage>
          . ELRA, Valletta, Malta (May
          <year>2010</year>
          ), http://is.muni.cz/publication/884893/en
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Rothermel</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harrold</surname>
            ,
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>A safe, e cient algorithm for regression test selection</article-title>
          .
          <source>In: 1993 Conference on Software Maintenance</source>
          . pp.
          <volume>358</volume>
          {
          <fpage>367</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Thie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Gitpython documentation</article-title>
          . https://gitpython.readthedocs.io/en/stable/, accessed:
          <fpage>2019</fpage>
          -07-30
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Vokolos</surname>
            ,
            <given-names>F.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frankl</surname>
            ,
            <given-names>P.G.</given-names>
          </string-name>
          :
          <article-title>Pythia: a regression test selection tool based on textual di erencing</article-title>
          . In: Reliability,
          <article-title>quality and safety of software-intensive systems</article-title>
          , pp.
          <volume>3</volume>
          {
          <fpage>21</fpage>
          . Springer (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>