<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identifying Vulnerable Functions from Source Code using Vulnerability Reports</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rabaya Sultana Mim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Toukir Ahammed</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kazi Sakib</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Information Technology, University of Dhaka</institution>
          ,
          <addr-line>Dhaka</addr-line>
          ,
          <country country="BD">Bangladesh</country>
        </aff>
      </contrib-group>
      <fpage>66</fpage>
      <lpage>73</lpage>
      <abstract>
        <p>Software vulnerability represents a flaw within a software product that can be exploited to cause the system to violate its security. In the context of large and evolving software systems, developers find it challenging to identify vulnerable functions efectively when a new vulnerability is reported. Existing studies have underutilized vulnerability reports which can be a good source of contextual information in identifying vulnerable functions in source code. This study proposes an information retrieval based method called Vulnerable Functions Detector (VFDetector) for identifying vulnerable functions from source code and vulnerability reports. VFDetector ranks vulnerable functions based on the textual similarity between the vulnerability report corpora and the source code corpora. This ranking is achieved modifying conventional Vector Space Model to incorporate the size of a function which is known as the tweaked Vector Space Model (tVSM). As an initial study, the approach has been evaluated by analysing 10 vulnerability reports from six popular open-source projects. The result shows that VFDetector ranks the actual vulnerable function at first position in 40% cases. Moreover, it ranks the actual vulnerable function within rank 5 in 90% cases and within rank 7 for all analysed reports. Therefore, developers can use these results to implement successful patches on vulnerable functions more quickly .</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;vulnerability identification</kwd>
        <kwd>vulnerable function</kwd>
        <kwd>vulnerability report</kwd>
        <kwd>source code</kwd>
        <kwd>vector space model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        report corpora. In addition, programming language spe- out manual expert intervention. Recently, DL based
techcific keywords is removed for generating code corpora. niques [
        <xref ref-type="bibr" rid="ref1 ref2">11, 12, 13</xref>
        ] has gained extensive use in detecting
Finally, to rank the vulnerable functions, similarity scores source code vulnerabilities due to its ability to
automatiare measured between the code corpora of the functions cally extract features from source code. DL based
methand report corpora by a modified version of Vector Space ods can be categorized into text-based and graph-based
Model (tVSM) where larger methods get more weight methods.
while ranking. Text based methods: The text-based approach in
      </p>
      <p>In experiments, as an initial study ten Common Vul- vulnerability detection treats a program’s source code as
nerabilities and Exposures (CVE) reports are chosen ran- text and employs natural language processing techniques
domly from six open source GitHub repositories. Based to identify vulnerabilities. Russell et al. [3] introduced
on the commit link available in reports we crawled the TokenCNN model, which utilizes lexical analysis to
the corresponding projects before the vulnerability was acquire source code tokens and employs a Convolutional
patched. The result analysis shows that VFDetector ranks Neural Network (CNN) to detect vulnerabilities.
the vulnerable functions at the first position in 40% cases, Li et al. [4] proposed Vuldeepecker, a method that
whereas it ranks the actual vulnerable function within collects code gadgets by slicing programs and transforms
top 5 in 90% cases and within top 7 in 100% cases. them into vector representations, training a Bidirectional</p>
      <p>It is evident from the results that VFDetector performs Long Short Term Memory (BLSTM) model for
vulnerapromisingly in detecting vulnerable functions against a bility recognition.
vulnerability report in a large scale software systems. It Zhou et al. [5] introduced µVulDeePecker, which
enis also observed that in Top 5 and Top 7 ranking, the hances Vuldeepecker by incorporating code attention
functions which ranks above the actual vulnerable func- with control dependence to detect multi-class
vulneration are the related functions of that vulnerability which bilities. However, the performance of these text-based
acquires higher similarity. It guides a developer to patch approaches is limited because they rely solely on static
those related functions too in order to mitigate that vul- source code analysis and do not account for the program
nerability from the system. semantics.</p>
      <p>The remainder of this paper is structured as follows: Graph based methods: To address the limitations
Section 2 gives an overview of previous studies on vul- of text-based methods, researchers have turned to
dynerability detection at file level or function level. Section namic program analysis to convert a program’s source
3 describes our methodology for detecting vulnerable code semantics into a graph representation facilitating
functions in a project. Section 4 reports our experimental vulnerability detection through graph analysis. Zhou et
ifndings and the analysis thereof. Section 5 demonstrates al. [6] introduced Devign which employs a graph neural
the threats to validity of our work. Section 6 motivates network for vulnerability identification. This approach
future research directions and concludes this paper. includes a convolutional module that eficiently extracts
critical features for graph-level classification from the
learned node representations. By pooling the nodes, a
2. Related Work comprehensive representation for graph-level
classification is achieved.</p>
      <p>In recent years, the research community has directed Cheng et al. [7] introduced a diferent approach named
significant attention toward the issue of vulnerability Deepwukong which divides the program dependency
detection, primarily due to the complex challenges it graph into various subgraphs after distilling the program
presents. The existing body of literature has introduced semantics based on program points of interest. These
numerous methodologies in response to these challenges. subgraphs are then utilized to train a vulnerability
deThese methods can be classified into three distinct cate- tector through a graph neural network. While these
gories based on the degree of automation: manual, semi- graph-based techniques prove more efective in
identifyautomatic, and fully automatic techniques. ing vulnerabilities but it is important to note that their</p>
      <p>Manual techniques rely on human experts to create scalability is worse than text based methods due to large
vulnerability patterns. However, all patterns can not be number of graph nodes in complex program.
generated manually, which leads to reduced detection efi- Exploring the existing literature, it is evident that
textciency in practical scenarios. In contrast, semi-automatic based methods lacks in incorporating program semantics
techniques involve human experts in the extraction of while graph-based methods achieve high accuracy
considspecific features like API symbols [ 9] and function calls ering source code semantics but have scalability issues in
[10], which are then fed into traditional machine learn- complex scenarios. Moreover, due to the underutilization
ing models for vulnerability detection. Full-automatic of contextual information like vulnerability reports with
techniques utilize Deep Learning (DL) to automatically source code existing methods fails to detect complicated
extract features and construct vulnerability patterns with- vulnerabilities in real-world projects. Because whenever
a new vulnerability is reported in a system it is hard to
detect in which function the vulnerability exist as the
system consist of huge volume of functions. Before using
vulnerable reports as a source of contextual information
in existing methods, it is important to verify whether
vulnerable functions can be identified efectively using
these reports. Moreover, identifying vulnerable functions
using vulnerability reports can play an efective role to
minimize the search space in existing methods.</p>
      <sec id="sec-1-1">
        <title>Source code corpora consist of source code terms used</title>
        <p>to assess similarity with vulnerability report corpora.</p>
        <p>Therefore, the precision of code corpora generation
directly impacts the precision of matching, consequently
enhancing the accuracy of vulnerability localization. In
this step all the folders are extracted from a system with
their corresponding C/C++ files. From each of these files
all functions are extracted automatically in individual C
ifles which ensures function level analysis. For Example:
3. Methodology CVE-2014-2038 of Linux version 3.13.5 consist of 15,675
ifles which has total 229,682 functions.</p>
        <p>
          This study proposes an approach which detects vulner- This stage generates a vector of lexical tokens by
doable functions from huge volume of files of a large soft- ing lexical analysis on every source code file. There are
ware system using vulnerability reports. The overall unnecessary tokens in source code which do not contain
process of this approach consist of three distinct steps any vulnerability related information. These tokens are
and those are Source Code Corpora Generation, Vulnera- discarded from source code such as programming
lanbility Report Corpora Generation, Ranking Vulnerable guage specific keywords (e.g., int, if, float, switch, case,
Functions. Each of these steps encompasses a series of struct), stop words (e.g., all, and, an, the). Many words in
tasks as illustrated in Figure 1. At first, all files and their the source code may include multiple words. For
examcorresponding functions are extracted from a particular ple, the term "addRequest" consists of the keywords "add"
version of a software system. Then these source code is and "Request". Mutiwords are separated using multi word
processed to create code corpora. Similarly vulnerability identifier. Furthermore, statements are divided according
report is processed to produce report corpora. Finally, to certain syntax-specific separators like ‘ , ’, ‘=’, ‘(’, ‘)’, ‘{’,
similarity between the report and code corpora is mea- ‘}’, ‘/’, and so on. WordNet2 is used to derive each word’s
sured using tweaked Vector Space Model (tVSM) to rank semantic meanings because a term might have more than
the vulnerable source code functions. one synonym. In specific cases, developers and Quality
Assurance (QA) personnel may employ diferent
termi3.1. Dataset nology, even though they are referring to the same
scenario with equivalent meanings. For example, the term
We used the benchmark dataset Big-Vul1 developed by ‘finalize’ may have multiple synonyms such as ‘conclude’
Fan et al. [
          <xref ref-type="bibr" rid="ref3">14</xref>
          ]. This dataset comprises reliable and com- or ‘complete.’ When describing a situation, if a developer
prehensive code vulnerabilities which are directly linked uses ‘finalize’ but QA opts for ‘conclude’, it’s challenging
to the publicly accessible CVE database. Notably, the cre- for the system to identify these variances without
considation of this dataset involved a significant investment of ering the semantic meanings of these words. Therefore,
manual resources to ensure its high quality. Furthermore, the extraction of semantic meaning is crucial in achieving
this dataset is noteworthy for its substantial scale, being accurate rankings for vulnerable functions.
one of the most extensive vulnerability datasets avail- The final stage of code corpora generation
incorpoable. It is derived from a collection of 348 open-source rates WordNet lemmatization, a technique that
normalGithub projects, encompassing a time span from 2002 izes words to their base or dictionary form. WordNet
to 2019, and covers 91 distinct Common Weakness Enu- lemmatization utilizes the comprehensive WordNet
lexmeration (CWE) categories. This comprehensive dataset ical database, organizing words into synonymous sets
comprises approximately 188,600 C/C++ functions, with called synsets. This method identifies word lemmas based
5.6% of them identified as vulnerable (equivalent to 10,500 on the word’s part of speech and context within
Wordvulnerable functions). This dataset provides granular Net, ofering a more context-aware approach to
lemmaground-truth information at the function level, specify- tization. As a result, it considers a word’s meaning and
ing which functions within a codebase are susceptible to contextual usage, allowing for precise reduction of words.
vulnerabilities. For instance, it transforms "running" to "run" and "better"
to "good" based on their meanings and parts of speech,
unlike standard lemmatization that typically relies on
sufix removal.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>1https://github.com/ZeoVan/MSR_20_Code_vulnerability_</title>
        <p>CSV_Dataset</p>
      </sec>
      <sec id="sec-1-3">
        <title>2https://wordnet.princeton.edu/</title>
        <sec id="sec-1-3-1">
          <title>3.3. Vulnerability Report Corpora</title>
          <p>A software vulnerability report contains information like
description about the vulnerability, severity rating,
vulnerability identifier (CVE-ID), reference to additional
sources which gives valuable insights about a software
vulnerability issue. However, these reports can also
include irrelevant terms such as stop words and words in
various tenses (present, past, or future). To refine
vulnerability reports, pre-processing is necessary. In the initial
stage of vulnerability report corpora creation, stop words
are eliminated. We apply WordNet Lemmatizer, similar
to what’s used for source code corpora generation, to
generate refined report corpora containing only relevant
terms.</p>
        </sec>
        <sec id="sec-1-3-2">
          <title>3.4. Ranking Vulnerable Functions</title>
        </sec>
      </sec>
      <sec id="sec-1-4">
        <title>In this step, relevant vulnerable functions are ranked</title>
        <p>based on the textual similarity between the query
(report corpus) and each of the function in the code corpus.
Vulnerable functions are ranked by applying tVSM. We
employ tVSM, which modifies the Vector Space Model
(VSM) by emphasising large-scale functions. In
traditional VSM, the cosine similarity is used to measure the
ranking score between the associated vector
representations of a report corpus (r) and function (f), according to
Equation 1.</p>
        <p>(,  ) = (,  ) =</p>
        <p>
          Here, ⃗ and ⃗ are the term vectors for the
vulnerability report (r) corpus and function (f) corpus
respectively. Throughout the years, numerous adaptations of
the tf(t,d) function have been introduced with the aim of
enhancing the VSM model’s efectiveness. These
encompass logarithmic, augmented, and Boolean modifications
of the traditional VSM [
          <xref ref-type="bibr" rid="ref4">15</xref>
          ]. It has been noted that the
logarithmic version can yield improved performance, as
indicated by prior studies [
          <xref ref-type="bibr" rid="ref5 ref6 ref7">16, 17, 18</xref>
          ]. From that point of
view, tVSM modified Equation 1 and uses the logarithm
of term frequency (tf) and if(inverse function frequency)
to give more importance on rare terms in the functions.
 · ⃗
⃗
⃗ ⃗
|| · |  |
(1)
Thus tf and if are calculated using Equation 2 and 3
respectively.
        </p>
        <p>(,  ) = 1 + 
  = (
# 

)
Here,  represents the frequency of a term  appearing
in a function  , #  denotes the total count of
functions within the search space,  signifies the overall
number of functions that include the term . Thus in
equation 4 each term weight is calculated as follows:
ℎ∈ = ( ) × (  )
= (1 +  ) × (
# 

)
The VSM score is calculated using equation 5.
(2)
(3)
(4)
calculate the length value for each source code function
based on the number of terms contained within the
function. Here we apply the normalized value of ’#terms’
as the argument for the exponential function − . The
normalization process is defined in Equation 7.</p>
        <p>Let z denote a set of data, with  and 
representing the maximum and minimum values of z term,
respectively. The normalized value for z term is
determined as:
 () =  − 
 − 
Considering the above analysis, tVSM score is calculated
by multiplying the weight of each function, denoted as
x(terms), with the cosine similarity score represented by
cos(r, f), as described in Equation 8:
  (,  ) = () × (,  )
(7)
(8)
(,  ) =
√︀∑︀(1 + log ) ×   2 ×
∑︁ (1 + log ) × (1 + log  ) ×   2×
∈∩
1</p>
      </sec>
      <sec id="sec-1-5">
        <title>Once the tVSM score for each function has been com</title>
        <p>puted, a list of vulnerable functions is arranged in
de1 scending order of scores. The function with the highest
√︀∑︀(1 + log  ) ×   2 similarity score is positioned at the top of the ranked list.</p>
        <p>(5)</p>
        <p>
          Traditional VSM tends to give preference to smaller
functions when ranking them, which can be problem- This section provides information on the practical
impleatic for large functions because they may receive lower mentation, the criteria used for evaluation and
experisimilarity scores. Past research [
          <xref ref-type="bibr" rid="ref10 ref8 ref9">19, 20, 21</xref>
          ] has indicated mental result analysis of this study.
that larger source code files are more likely to contain
vulnerabilities. Therefore, in the context of vulnerability 4.1. Implementation
localization, it’s crucial to prioritize larger functions in
our ranking. To address this issue, we introduce a func- The proposed method is implemented in python (version
tion denoted as ’x’ (as shown in Equation 6) within the 3.11.5). The experiment was conducted on an Windows
tVSM model, aiming to account for the function’s length. server equipped with an Intel(R) Core(TM) i5-10300H
CPU processor @3.0GHz and having 64GB of RAM. The
() = 1 − −(# ) (6) implementation involves various python libraries and
Equation 6 represents a logistic function, specifically an NLTK (Natural Language Toolkit) libraries for text
proinverse logit function, designed to ensure that larger func- cessing and feature extraction. It takes function files
tions receive higher rankings. We employ Equation 6 to as input and provides ranking of suspicious vulnerable
functions as output.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Experiment and Result Analysis</title>
      <sec id="sec-2-1">
        <title>4.2. Evaluation</title>
        <sec id="sec-2-1-1">
          <title>To conduct this research we used the extensive Big-Vul</title>
          <p>dataset which contains large scale vulnerability reports
of C/C++ code from open source GitHub projects. Other
C/C++ datasets can also be used. Based on the highest
number of vulnerabilities reported, we choosed top six
well known projects from this dataset which are Chrome,
Linux, Radare2, ImageMagick, Tcpdump and FFmpeg
as shown in Table 1. As the selected projects are
opensource in nature and are hosted on GitHub, serving as
the primary platform for storing code and managing
version control. It allows us to extract all essential
commits for our analysis. Additional information about the
repositories can be found in Table 2. As an initial study,
VFDetector was evaluated using ten vulnerability reports
from these six open-source projects which are chosen
randomly from the dataset. Table 1 lists the analysed
project name, CVE ID of report, and the source code link.</p>
          <p>To measure the efectiveness of the proposed
vulnerability detection method, we use the Top N Rank metric.
This metric signifies the count of vulnerable functions
ranked in the top N (where N can be 1, 5, or 7) in the
obtained results. When assessing a reported vulnerability,
if the top N query results include at least one function
that corresponds to the location where the vulnerability
needs to be addressed, we determine that the vulnerable
function is detected successfully. Table 2 includes ten
vulnerability reports from six open source projects with
their number of commits, total files, total functions,
actual vulnerable functions name and finally VFDetector
ranking in Top N ranked functions in output. The
results of Table 2 shows that among the ten CVE reports
VFDetector ranks the actual vulnerable function at the
1st position for four (40%) reports which are CVE ID
#13000, #14470, #15033, #16359. For five reports (50%)
with CVE ID #3916, #2094, #2038, #10190 and #11339 it
ranks the vulnerable function in Top 5 rank. It indicates
that nine (90%) reports are ranked in Top 5. For one
report CVE-2013-6763 of Linux Kernel version 3.12.1 it
ranks the vulnerable function in Top 7 rank i.e., in 7th
position out of total 273,898 functions. Upon manual
inspection, we observed that the six functions preceding
the vulnerable function exhibit a higher similarity score
compared to the actual vulnerable function. The reason
behind this can be the inter-connectedness of these six
functions with the vulnerable function through function
calls. It is also noticeable that projects with less number
of functions ranks the vulnerable function in 1st position
and with large number of functions the ranking decreases
slightly. The reason behind this is larger projects might
contain more associated functions which are needed to
be fixed in order to address a particular vulnerability.</p>
          <p>In summary, the experimental results show that
VFDetector can detect vulnerable functions from a huge
volume of functions and can also suggest developers with
the related functions having highest similarity scores
which might need to be patched to address the reported
vulnerability. Moreover, to the best of our knowledge
we are the first to incorporate vulnerability reports in
software vulnerability detection from the concept that
vulnerability report’s description contain conceptual
information about a reported vulnerability. Based on the
promising results in this initial evaluation, the future
work can be analyzing more vulnerable reports from
diverse projects to make the approach comparable and
generalizable.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Threats to Validity</title>
      <sec id="sec-3-1">
        <title>In this section, we discussed the potential threats which may afect the validity of this study.</title>
        <p>Threats to external validity: The generalizability
of the acquired results poses a threat to external validity.
The dataset that we used in our research was gathered
from open-source. Open-source projects may contain
data that difers from those created by software
companies with sound management practices. Seven Apache
projects are examined in this study. More projects from
other systems are needed to be evaluated for the
generalisation. However, to overcome this threat large-scale
diversified projects with long change history is to be
chosen.</p>
        <p>Threats to internal validity: One limitation of our
approach is its reliance on sound programming practices
when naming variables, methods, and classes. If a
developer uses non-meaningful names, it could have an
adverse impact on the efectiveness of vulnerability
detection.ay not fully represent the characteristics of the
whole program. Additionally, our model is evaluated
with C/C++ functions and it may encounter challenges
in detecting vulnerabilities in other programming
languages.</p>
        <p>Threats to construct validity: We used the WordNet
database and lemmatizer of NLTK library as essential
components in text pre-processing to extract word
semantics and reduce words to their base forms. Since
these resources are well known for their usefulness in
NLP, we relied on their accuracy. Moreover,
vulnerability reports ofer essential information that developers
rely on to address and patch vulnerable functions. A bad
vulnerability report delays the fixing process. It’s worth
noting that our approach is dependent on the quality of
these reports. If a vulnerability report lacks suficient
information or contains misleading details, it can have a
detrimental impact on the performance of VFDetector.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Vulcnn:</surname>
          </string-name>
          <article-title>An image-inspired scalable vulnerability detection system</article-title>
          ,
          <source>in: Proceedings of the 44th International Conference on Software Engineering</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>2365</fpage>
          -
          <lpage>2376</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Vuldeelocator: A deep learning-based system for detecting and locating software vulnerabilities</article-title>
          ,
          <source>IEEE Transactions on Dependable and Secure Computing</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. N.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , A c/c++
          <article-title>code vulnerability dataset with code changes and cve summaries</article-title>
          ,
          <source>in: Proceedings of the 17th International Conference on Mining Software Repositories</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>508</fpage>
          -
          <lpage>512</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>H.</given-names>
            <surname>Schütze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Raghavan</surname>
          </string-name>
          , Introduction to information retrieval, volume
          <volume>39</volume>
          , Cambridge University Press Cambridge,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Croft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Metzler</surname>
          </string-name>
          , T. Strohman,
          <article-title>Search engines: Information retrieval in practice</article-title>
          , volume
          <volume>520</volume>
          ,
          <string-name>
            <surname>Addison-Wesley</surname>
            <given-names>Reading</given-names>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sakib</surname>
          </string-name>
          ,
          <article-title>An appropriate method ranking approach for localizing bugs using minimized search space</article-title>
          .,
          <source>in: ENASE</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>303</fpage>
          -
          <lpage>309</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Rahman</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Sakib</surname>
          </string-name>
          ,
          <article-title>A statement level bug localization technique using statement dependency graph</article-title>
          .,
          <source>in: ENASE</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>171</fpage>
          -
          <lpage>178</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>N. E.</given-names>
            <surname>Fenton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ohlsson</surname>
          </string-name>
          ,
          <article-title>Quantitative analysis of faults and failures in a complex software system</article-title>
          ,
          <source>IEEE Transactions on Software engineering 26</source>
          (
          <year>2000</year>
          )
          <fpage>797</fpage>
          -
          <lpage>814</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>T. J.</given-names>
            <surname>Ostrand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Weyuker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Bell</surname>
          </string-name>
          ,
          <article-title>Predicting the location and number of faults in large software systems</article-title>
          ,
          <source>IEEE Transactions on Software Engineering</source>
          <volume>31</volume>
          (
          <year>2005</year>
          )
          <fpage>340</fpage>
          -
          <lpage>355</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [21]
          <string-name>
            <surname>H. Zhang,</surname>
          </string-name>
          <article-title>An investigation of the relationships between lines of code and defects</article-title>
          ,
          <source>in: 2009 IEEE international conference on software maintenance, IEEE</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>274</fpage>
          -
          <lpage>283</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>