<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Benchmarking the vulnerability detection capabilities of software analysis tools</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elena Baninemeh</string-name>
          <email>e.baninemeh@uu.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Slinger Jansen</string-name>
          <email>slinger.jansen@uu.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lappeenranta University</institution>
          ,
          <addr-line>Lappeenranta</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Utrecht University</institution>
          ,
          <addr-line>Utrecht</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Code cloning and copy-pasting code fragments is common practice in software engineering. If security vulnerabilities exist in a cloned code segment, those vulnerabilities may spread in the related software, potentially leading to security incidents. Code similarity is one efective approach to detect vulnerabilities hidden in software projects. However, due to the complexity, size, and diversity of source code, current methods sufer from low accuracy, and poor performance. Moreover, most existing clone detection techniques focus on a limited set of programming languages in the detection process. We propose to solve these problems using SearchSECO, a software analysis tool that detects vulnerabilities in multiple programming languages.</p>
      </abstract>
      <kwd-group>
        <kwd>Software vulnerability</kwd>
        <kwd>code clone detection</kwd>
        <kwd>software security</kwd>
        <kwd>open-source software</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The rapidly growing demands for software lead to the increasing popularity of code reuse,
including existing code templates and components. Open-source software (OSS) has become
one of the best solutions to improve both the eficiency and the quality of development at the
meanwhile reducing cost. However, a considerable number of vulnerabilities in OSS programs
would naturally lead to many software vulnerabilities caused by code cloning, which poses
a severe threat to system security [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In fact, OSS has increased the rate of vulnerabilities
because, as the name implies, the code is open-source and available to everyone. Most software
developers copy the code from other software systems and reuse them without significant
modification. This type of reuse code is called code cloning [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Code cloning is expected to rise,
especially with tools such as GitHub co-pilot, which uses code templates and auto-completion
features to support software engineers.
      </p>
      <p>Information about known vulnerabilities is published through diferent resources such as the
National Vulnerability Database (NVD) in the form of Common Vulnerabilities and Exposures
(CVE). Existing techniques for vulnerable code clone detection fall into two categories: code
similarity and functional similarity. In code similarity approaches, the target source code is
https://www.slingerjansen.nl/ (S. Jansen)</p>
      <p>
        © 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
compared against a set of known vulnerable code samples and determined to be vulnerable if
a threshold of similarity is met. Code similarity approaches are typically classified based on
four types of detection coverage. type-1 (identical), type-2 (syntactically equivalent), type-3
(syntactically similar), and type-4 (semantically similar) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        On the other hand, functional similarity approaches seek to generate abstract functional
patterns of code which model vulnerable behavior. However, due to the complexity of building
such a pattern, these techniques are typically specialized to only a small class of vulnerabilities
or a particular source code project, rendering them inefective as general-purpose vulnerable
code clone detection techniques [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>In this work, we introduce SearchSECO, a code-similarity technique capable of identifying
modified vulnerable code clones while remaining generic to type-1 and type-2 and supporting
multiple languages. Additionally, we built a database by mining vulnerable and patched source
code from GitHub. In this paper, we present the main two processes, including vulnerabilities
collection and vulnerabilities detection.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Research Approach</title>
      <p>SearchSECO is a large database of methods of the top rated projects (with “stars”) on Github.
SearchSECO clones a git project, extracts a number of versions, and extracts the files and authors
from those versions. The method’s abstract syntax tree is extracted and a representation of
this abstract syntax tree is hashed. SearchSECO currently parses Java, Javascript, C/C++, and
Python. SearchSECO is itself a project on Github and can be found via: https://github.com/
SecureSECO/SearchSECOController. Furthermore, the database can be accessed through a
portal: https://secureseco.science.uu.nl/portal/. In this portal visitors can enter their own project
link and email address. After the project has been processed and matched, the visitor receives a
report of the matches in the SearchSECO database and can determine if there are any potentially
vulnerable fragments in their project. Currently (June 14th 2022), the database contains 16
million unique methods from approximately 100 thousand projects from Github.</p>
      <p>The database until recently only had matching capability, but as it is the meta-data that
makes the method database interesting, we have started by matching vulnerability data from
vulnerability database and directly from open source project. In this paper, our research objective
is to benchmark SearchSECO’s performance to other tools. It must be noted that software
engineers using SearchSECO are not time constrained. However, they do care about accurate
feedback about their projects and therefore we benchmark SearchSECO against other available
state-of-the-art tools in terms of precision, rather than performance speed.</p>
      <p>The main research question in this study is as follows:(MRQ) How efective is the vulnerability
detection feature of SearchSECO compared to other vulnerability detection approaches? We
formulated the following research questions to address the MRQ:  1: Is detection reporting in
vulnerability detection approaches suficient for benchmarking?  2: How can SearchSECO be
compared accurately to other vulnerability detection approaches?  3: How is the scalability
of SearchSECO in detecting vulnerabilities compared to state-of-the-art approaches?  4: How
efective is SearchSECO in detecting the latest vulnerabilities published by CVE and GitHub?</p>
      <p>
        We employed a literature study using the snowballing method, combined with document
analysis and replication study, and performing an experiment to compare SearchSECO with
other vulnerability detection tools. Table 1 shows the mapping between the research questions
and the research methods. The preferred literature study method is snowballing. Wohlin [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
presents several guidelines for this method which will use during the literature study. Document
analysis is a systematic procedure for reviewing or evaluating documents, including manuscripts
and illustrations, that have been published without a researcher’s intervention [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Document
analysis is one of the analytical methods in qualitative research that requires data investigation
and interpretation to elicit meaning, gain understanding, and develop empirical knowledge [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
The preferred literature study method is snowballing. Wohlin [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] presents several guidelines
for this method which will use during the literature study. Document analysis is a systematic
procedure for reviewing or evaluating documents, including manuscripts and illustrations, that
have been published without a researcher’s intervention [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Document analysis is one of the
analytical methods in qualitative research that requires data investigation and interpretation
to elicit meaning, gain understanding, and develop empirical knowledge [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Like many other
empirical disciplines, replication has been seen as an essential means of assessing reliability
and confidence in empirical findings. A key component of experimentation is replication. To
consolidate a body of knowledge built upon experimental results, they must be extensively
verified. This verification is carried out by replicating an experiment to check if its results can
be reproducible [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. We aim to use the ACM SIGSOFT Empirical Standards 1 for benchmarking
procedure.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. SearchSECO</title>
      <p>In this section, we describe the design and implementation of our proposed approach
SearchSECO for method-level vulnerability detection. We aim to accurately discover the code clones
between a set of vulnerable codes and a target program using the code clone detection technique.
In SearchSECO, we focus on the detection of vulnerable code fragments accurately, Scaling to a
large code base, and Supporting multiple languages.</p>
      <p>1https://github.com/acmsigsoft/EmpiricalStandards/blob/master/docs/Benchmarking.md</p>
      <sec id="sec-3-1">
        <title>3.1. Vulnerabilities collection process</title>
        <p>We collected vulnerability data of each project from two sources: NVD and public Git repositories
on GitHub. NVD is a vulnerability database built upon and fully synchronized with the CVE list.
In addition to a large amount of vulnerability data, it also provides enhanced information (e.g.,
vulnerability type, references to solutions) for each record. GitHub provides a larger quantity
and wider variety of code, which can help us supplement the vulnerability dataset. We built the
SearchSECO vulnerability database in the following steps:
• We crawled all of the vulnerability entries in the CVE database and NVD, such as the
descriptive information for each vulnerability. Specifically, we parse the Github web
pages to extract CVE Details such as vulnerable lines and hash commits.
• The NVD receives its vulnerability listings directly from the CVE. Therefore,
vulnerabilities that are not reported to the CVE, so they would not publish in the NVD. Hence,
beside extracting vulnerabilities from NVD, we have to extract vulnerabilities from Github
(See figure 1). First, SearchSECO clone the repository by using the ”git clone repository”
command. Then, it will search for the commits regarding CVEs for each repository by
using the ” git log –grep=“CVE-20” command. This process of collecting vulnerable code
and extraction required data is fully automated,</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Vulnerabilities detection process</title>
        <p>In this section, we describe our approach to vulnerability detection, which is a scalable approach
to code clone detection. The types of code clones have to be clarified in order to explain the
process. Four diferent types of code clones are, Type-1: Exact clones, Type-2: Renamed clones,
Type-3: Restructured clones, and Type-4: Semantic clones.</p>
        <p>Type-1: Identical code fragments, but may have some variations in whitespace, layout, and
comments.</p>
        <p>Type-2: Syntactically equivalent fragments with some variations in identifiers, literals, types,
whitespace, layout, and comments.</p>
        <p>Type-3: Syntactically similar code with inserted, deleted, or updated statements.
Type-4: Semantically equivalent but syntactically diferent code.</p>
        <p>We designed SerachSECO to detect Type-1 and Type-2 clones because our goal is to reduce
false positives and negatives and increase scalability.</p>
        <sec id="sec-3-2-1">
          <title>3.2.1. Code clone detection</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2.2. Prepossessing</title>
          <p>The following steps will perform in the preprocessing when SearchSECO receives the code
fragment or project to detect vulnerabilities.
1. Method extraction: The process start with retrieving functions from a given program by
using a robust parser.
2. Abstraction and normalization: we used an abstraction and normalization feature, so
every formal parameters, local variables, data types, and function calls that appear in the body
of a function are replaced with symbols such as FPARAM in level 1, LVAR in level 2, DTYPE in
level 3, and FUNCCALL in level 4.
3. Generating hash value: In this step, the hash value generate based on the MD5 algorithm.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Related work</title>
      <p>
        Many approaches have been proposed to detect the vulnerabilities brought by code clones. Kim
et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] proposed VUDDY, a highly eficient method for detecting vulnerable code cloning,
which is achieved by leveraging function-level granularity and a length-filtering technique
that reduces the number of signature comparisons. However, it does not support common
code modification methods such as word order modification and redundant code insertion,
which causes its limitation in practice. VFDETECT [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] proposed an approach based on an
innovative fingerprint model to detect vulnerable code. VulPecker [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] developed a technique
that identifies a vulnerability-to-similarity-algorithm mapping. This way, each algorithm can
be applied to the vulnerabilities to which they are best suited. However, this approach is still
limited by the underlying accuracy of the similarity algorithms and only achieves a recall score
of 60%, meaning many vulnerable clones were left undetected. VCIPR [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is a scalable system
for vulnerability detection in unpatched source code. That uses a fast, token-based approach to
detect vulnerabilities at function level granularity.
      </p>
      <p>
        Akram and Luo [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] developed a quantitative vulnerability detection technique based on the
code clone detection technique at the source code level. They retrieved vulnerable source code
ifles from the various web source code repositories by tracking the patch file of vulnerabilities.
Then, the vulnerable source code files are retrieved using common vulnerabilities and exposures
(CVE) numbers.
      </p>
      <p>
        ReDeBug [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] is a technique that does use the information in both the vulnerable code and
the patched code. ReDeBug performs sequence-based matching utilizing the dif files associated
with a particular vulnerability. A dif file contains the lines that were explicitly modified during
the transition of the code from vulnerable to patched, as well as some context code within close
textual proximity. This allows ReDeBug to detect some type-3 clones; however, if the code
modification is near the location of the lines modified during the patch process, this technique
will fail to detect the vulnerable clone.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this study, we propose a vulnerability detection tool to benchmark with diferent approaches
and methodology from state-of-the-art research on vulnerability detection. We aim to design
our approach for scalable and accurate detection of vulnerable code clones. Moreover, we aim
to address an automated way to collect vulnerable functions and implement SearchSECO to
demonstrate its eficacy and efectiveness to detect numerous vulnerable clones from a large
code base with unprecedented scalability and accuracy.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>A novel vulnerable code clone detector based on context enhancement and patch validation</article-title>
          ,
          <source>Wireless Communications and Mobile Computing</source>
          <year>2022</year>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mondal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. K.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <article-title>Identifying code clones having high possibilities of containing bugs</article-title>
          ,
          <source>in: 2017 IEEE/ACM 25th International Conference on Program Comprehension (ICPC)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C. K.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Cordy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Koschke</surname>
          </string-name>
          ,
          <article-title>Comparison and evaluation of code clone detection techniques and tools: A qualitative approach</article-title>
          ,
          <source>Science of computer programming 74</source>
          (
          <year>2009</year>
          )
          <fpage>470</fpage>
          -
          <lpage>495</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F. P.</given-names>
            <surname>Viertel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Brunotte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Strüber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <article-title>Detecting security vulnerabilities using clone detection and community knowledge</article-title>
          .,
          <source>in: SEKE</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>245</fpage>
          -
          <lpage>324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wohlin</surname>
          </string-name>
          ,
          <article-title>Guidelines for snowballing in systematic literature studies and a replication in software engineering</article-title>
          ,
          <source>in: Proceedings of the 18th international conference on evaluation and assessment in software engineering</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Bowen</surname>
          </string-name>
          ,
          <article-title>Document analysis as a qualitative research method, Qualitative research journal (</article-title>
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Corbin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Strauss</surname>
          </string-name>
          ,
          <article-title>Basics of qualitative research: Techniques and procedures for developing grounded theory</article-title>
          , Sage publications,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Juristo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. S.</given-names>
            <surname>Gómez</surname>
          </string-name>
          ,
          <article-title>Replication of software engineering experiments, in: Empirical software engineering</article-title>
          and verification, Springer,
          <year>2010</year>
          , pp.
          <fpage>60</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Software systems at risk: An empirical study of cloned vulnerabilities in practice</article-title>
          ,
          <source>Computers &amp; Security</source>
          <volume>77</volume>
          (
          <year>2018</year>
          )
          <fpage>720</fpage>
          -
          <lpage>736</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <article-title>Vfdetect: A vulnerable code clone detection system based on vulnerability fingerprint</article-title>
          ,
          <source>in: 2017 IEEE 3rd Information Technology and Mechatronics Engineering Conference (ITOEC)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>548</fpage>
          -
          <lpage>553</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <article-title>Vulpecker: an automated vulnerability detection system based on code similarity analysis</article-title>
          ,
          <source>in: Proceedings of the 32nd Annual Conference on Computer Security Applications</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>213</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Akram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <article-title>Vcipr: vulnerable code is identifiable when a patch is released (hacker's perspective)</article-title>
          ,
          <source>in: 2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST)</source>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>402</fpage>
          -
          <lpage>413</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Akram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <article-title>Sqvdt: A scalable quantitative vulnerability detection technique for source code security assessment</article-title>
          ,
          <source>Software: Practice and Experience</source>
          <volume>51</volume>
          (
          <year>2021</year>
          )
          <fpage>294</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Jang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brumley</surname>
          </string-name>
          ,
          <article-title>Redebug: finding unpatched code clones in entire os distributions</article-title>
          ,
          <source>in: 2012 IEEE Symposium on Security and Privacy</source>
          , IEEE,
          <year>2012</year>
          , pp.
          <fpage>48</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>