<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Measuring Software Security Using Improved CWE Base Scores</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sabrina Mamtaz Nourin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>George Karabatis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Foteini Cheirdari Argiropoulos</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Systems, University of Maryland</institution>
          ,
          <addr-line>Baltimore County (UMBC), Baltimore, Maryland</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ML4Cyber</institution>
          ,
          <addr-line>Baltimore, Maryland</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Increasing the security of a software system by decreasing the number of its vulnerabilities has been a major objective of any organization. Therefore, it is important to identify a measure that indicates the security level of the software system. This paper presents a scoring method to measure the security posture of a software system. This novel scoring method for Common Weakness Enumeration (CWE)s considers semantic information in order to increase the accuracy of the score and provides a better outlook of the security posture of a software system using full automation.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;CWE base score</kwd>
        <kwd>software security score</kwd>
        <kwd>context information</kwd>
        <kwd>semantic comparison</kwd>
        <kwd>natural language processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Common Vulnerabilities and Exposures (CVE) [4] pro</title>
        <p>vides a list that contains actual instances of the
weakAs software systems are quite prevalent in today’s digital nesses listed in CWE. The CVE list comes with a
pubsociety, the existence of vulnerabilities in the software is licly available score, along with impact and exploitability
often the culprit for malicious attacks, resulting in secu- sub-scores for each vulnerability. Impact represents how
rity breaches, leading to loss of information and computer much damage a weakness can cause if exploited, and
assets, and inflicting disruption of operations, and finan- exploitability measures the ease of causing damage by
cial losses. The more we examine software systems the exploiting the weakness. Practically, the base score is the
more we discover potential vulnerabilities. Therefore, sum of impact and exploitability sub-scores [4]. CVSS
the security of software is an issue of high priority, al- (Common Vulnerability Scoring System) [5] is the
ofithough it is a labor intensive and complex task. The first cial guideline to generate scores for CVE. CVSS scores
step is to gauge the level of security in a software system are publicly available, therefore, most related research
and calculate it. uses them to generate software security scores. However,</p>
        <p>In this paper we propose a novel method to com- the same CVSS score is typically assigned to the same
pute software security score using semantic knowledge type of vulnerabilities (all nearly similar vulnerabilities
and Natural Language Processing (NLP). We introduce have the same score despite certain few yet substantial
a new scoring system for Common Weakness Enumer- diferences, especially when operated in diferent
enviation (CWE), which is a community-developed list of ronments). Sometimes a CVSS score is not available for
software and hardware weakness types. Identified weak- all possible CWEs. Based on our understanding and
exnesses in software code written in any programming perimentation, using a combination of semantic
informalanguage, are commonly found in CWE, which serves as tion of CWE along with CVE, one can generate a more
a measuring stick for security tools, and as a baseline for accurate software security score, because CWE provides
weakness identification, mitigation, and prevention ef- additional information for each vulnerability and it is the
forts [1]. There is no publicly available list of CWE scores basis for CVE.
[2]. MITRE [3] provides guidelines on how to calculate The CWSS (Common Weakness Scoring System) [6]
the CWE scores. But it is a time-consuming process, very is a system that provides a mechanism for prioritizing
user dependent and requires manual input from the user. CWEs based on a weakness scoring concept. CWSS
proTherefore, our objective is to provide an accurate and poses three metrics for weakness scoring – base metric,
automated scoring method for CWEs, by implementing attack surface metric, and environmental metric. The
context similarity using NLP. base metric captures the inherent risk of the weakness,
and it does not change over time and environment. Also,
the base score for each CVE is generated using this base
metric, and it does not change over time and system
environment. The attack surface metric represents the
barriers that an attacker must overcome in order to
exploit the weakness. Finally, the environmental metric
PSTCI2021: 3rd International Workshop on Privacy, Security, and
Trust in Computational Intelligence , November 01–05, 2021,
Queensland, Australia
" snourin1@umbc.edu (S. M. Nourin); georgek@umbc.edu
(G. Karabatis); contact@ml4cyber.com (F. C. Argiropoulos)
© 2021 Copyright for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).</p>
        <p>CEUR Workshop Proceedings (CEUR-WS.org)</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>represents characteristics of the weakness that are
specific to a particular environment or operational context.</p>
      <p>In this paper, we propose a solution to automatically We categorize the related work in two topics:
improvecalculate the CWE base score using NLP, without any ment of vulnerability scores, and calculating the overall
user input. We augment the components of base metrics security of a system. Most of these works use CVE as the
by utilizing semantic similarity of the CWE and CVE basis for scoring. Very few works have considered CWE
descriptions. By doing so, we obtain an accurate CWE for measuring software security.
score for each weakness in the system, and hence, we
generate an accurate score that represents the security 2.1. Vulnerability score improvement
posture of the software.</p>
      <p>We first follow the scoring guidelines from MITRE [ 2] A lot of work has been performed to predict and acquire
to generate a security score, which is used as a baseline a score for vulnerabilities (CVEs). In [7], the authors
for comparison purposes against our method. The base- use logistic regression to detect the probability of
havline yields a base score of a CWE by mapping CWE to ing a vulnerability and binomial distribution to detect
CVE using their id, and taking the average of the base the severity of vulnerabilities, on vulnerability datasets
scores of mapped CVEs. Then, the baseline is used to gen- of Adobe flash player and Firefox. The authors used
erates the security score by taking the sum of base scores mean time to vulnerability (temporal factor), local risk
of all CWEs that exist in the software system, multiplied rate, mean risk rate, and overall risk value as important
by their respective percentage of appearance. A higher factors to predict the severity of a vulnerability. In [8],
security score means that the system is more secure and the authors used attack surface entity point connection
less vulnerable to possible attacks. However, this direct with vulnerability location to predict vulnerability
exmapping between CWE and CVE may not provide an ploitation potential. All of these works used the CVSS
accurate score, because a large number of CWEs cannot (Common Vulnerability Scoring System) [5] score. In
be mapped with any CVE, while a small number of CWEs [9], temporal and environmental metrics were used to
is mapped with multiple CVEs, introducing inaccuracies obtain an improved CVSS (Common Vulnerability
Scorin the base score of CWE. ing System) score. In [10], the authors used temporal</p>
      <p>We propose a novel algorithm that leads to more accu- factor with Pareto and Weibull distribution [11] on
disrate way to score the CWEs by mapping to CVE using cover, disclosure, exploit and patch date to add a temporal
context similarity. We use the available CVE scores [4] to metric. Subsequently, [12] and [13] were inspired from
generate a base score for each CWE. We discover similar [10] for temporal metrics. In [14], the authors also used
CVEs and CWEs based on their respective description the temporal factor to calculate software integrity level.
using NLP techniques, and using this similarity, we can These temporal factors consider the time passed since
generate the CWE base scores, and hence, an overall soft- the discovery of that vulnerability. The temporal factor
ware security score. Software developers and managers is adjusted as newer versions of CVE come in, sometimes
can observe the continuous improvement of the software the base scores change, some other times a resolved CVE
based on the changes of the scores of consecutive phases is even removed from the list. So, for the time being, we
of fixing, testing, and debugging. In addition, the score can ignore the temporal metrics. In [13], the authors used
generated for each CWE can be used to generate a rank- software context information to prioritize CVSS based
ing of weaknesses in the system, and determine which vulnerabilities, where the context information works as
ones to address first, making a positive impact on pro- equivalent to temporal and environmental metrics. Our
grammers and managers. The final result of the entire goal is to make a better use of context information by
process is a scored list of the actual vulnerabilities that incorporating it with CWE to generate a better CWE
exist in the software system, and its overall security score. score. Recently, information extraction techniques are
In summary, the contributions of this paper are: being used to predict vulnerability score from the
vulnerability description. In [15], the authors used word
• Generate a more accurate base score for software embedding (using skip-gram [16]) and one-layer shallow
weaknesses (CWEs) using context similarity of Convolutional Neural Network (CNN) [17] to predict
vulCWE and CVE. nerability severity using the CVE database. They used
• Calculate the security score of a software system probabilistic distribution to predict the severity level of
using the generated CWE base score in a fully the vulnerabilities. In [18], the authors also used word
automated way. embedding (using word2vec [19]) with CNN to create
• Evaluate the accuracy of the scoring mechanism sentence embedding of CVE, CWE and Common Attack
through experimentation. Pattern Enumeration and Classification (CAPEC) [ 20],
and knowledge graph to predict security entity
relationship. The vectors generated by word embedding do not
provide the entire context information. Hence, there
is room for mistakes, therefore, the semantic process
does not produce the correct result. To resolve this, we
are using sentence embedding. It does a better, due to
comparison between sentences, because it considers the
context of sentence. In [21], the authors used PageRank
[22] algorithm on CVSS to calculate the CWE score. They
tried to improve it by using the child-parent relationship
between diferent CWE. This can be useful to identify the
possibility of a weakness being present in the system. But
as we certainly know which CWEs exist in the system
from our prior work [23], we can directly calculate the
CWE score without considering its possible children.
compare the results of both methods.</p>
      <sec id="sec-2-1">
        <title>3.1. Baseline</title>
        <sec id="sec-2-1-1">
          <title>For the baseline, we generate the CWE base score by</title>
          <p>mapping the CWE ID with its corresponding CVE ID,
using the CVE list provided by MITRE [4]. We have
followed the scoring suggestions from MITRE [2], which
are:</p>
          <p>Base Score for a CWE: Every CVE can be mapped
to a CWE (multiple CVEs can be mapped with the same
CWE), but not every CWE can be mapped to a CVE.</p>
          <p>Therefore, some CWE (CWEX) can be mapped with one or
more CVEs, and some CWE cannot be mapped with any
2.2. Security score CVE. The CVE list provided by MITRE [4] includes the
following information for each CVE: a CWE ID, a CVSS
As the vulnerability score gained from CVSS only pro- base score, and impact and exploitability sub-scores. For
vides information about individual CVEs, and not the each CWE, we take the average of the CVSS base scores
total security posture of the software system, we need an of its mapped CVEs. Then we normalize this average
approach to calculate the software security score accu- score by dividing it by the range of CVSS base scores
rately. There are only few existing works that concentrate (diference of the maximum and minimum CVSS base
on calculating the security score of the entire software. scores among all the mapped CVEs) for that particular
In [24], the authors proposed a method to assess the po- CWE. The result is the base score Sc(CWEX) of a CWE.
tential risk of cyber-attacks utilizing network-wide
compliance reports. They consider vulnerability distribution, (  ) = ( ) − ( ) (1)
dependency between the vulnerabilities, and network ( ) − ( )
configuration. But they do not consider the important
impact factor. In [25], the authors propose that security
metrics are proportional to the weighted total of base
score and the weight of the vulnerabilities. In [26], the
authors used aspect-oriented Stochastic Petri nets for
threat modeling, where the authors followed the
security metrics calculation proposed in [25]. They consider
threat categorization of STRIDE for creating a way for
threat mitigation. CVSS is used in [26], [24], [25] as the
basis of vulnerability scoring. Our goal is to improve
the security score compared to the above methods by
using CWE along with CVE for scoring, and by using
environmental metrics.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>In [2], weight is represented as frequency, and it is calcu</title>
          <p>lated as the number of CVEs that can be mapped to a CWE
within the National Vulnerabilities Database (NVD) [27].</p>
          <p>The NVD is the U.S. government repository of
vulnerability management data. This mapping with the entire
database of CVE was performed to rank the entire CWE
list. As every software system has a small number of
repeated CWEs, we do not need to consider the entire
list of CWE to generate that system’s security score. So, static analysis scanning tools, which scan the code and
we modified the weight calculation as follows: instead generate a list of potential CWEs. Based on our prior
of counting the number of times a CWE maps with any research, we observe that many of these potential CWEs
CVE in the entire NVD database, we count the number are in fact false positives. So, we designed and
impleof times a particular CWE appears in the system. mented the VINCI (Vulnerability IdeNtifiCatIon) tool,</p>
          <p>Weakness Score of a CWE: After generating the which identifies and labels the false positive
vulnerabilCWSS base score for each CWE in the system, we take ities using the output report of the scanning tools [23].
their weighted average to determine the security score VINCI uses author information along with other criteria
[25] [26]. To get the base score of each CWE, we multiply to accurately identify false positive CWEs, allowing only
their mapped base score with their weight. The sum of true positive CWEs to be considered for security
scorthe base scores of all existing CWE in a system is the ing. Consequently, using VINCI, we get an actual list of
weakness score of a system. weaknesses (CWE) in a system. Then, we calculate an
accurate base score for the CWEs, and using this base
  = ∑︁((  ) × (  )) (3) score, generate the software security score.</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>Security Score: In this method, the CWE base score is</title>
          <p>generated from CVE base scores. As the CWE base score
is the weighted average of base scores of the mapped
CVEs, the maximum value of CWE base scores must also
be 10. Since the highest value of base scores is 10, and
the weights are less than 1 (weight in this case is defined
as the percentage of times a CWE appears in a system),
this score identifies how weak the overall system is, and
ranges from 1-10. We multiply the base score by 10 to
generate a more conventional score out of 100. Therefore,
the security score of the system, based on the weakness
score as follows:
  = 100 −  
(4)
In the baseline, we map a CVE to a CWE, to get the
CWE base score. Though this method is widely used,
it does not give us an accurate CWE score, due to the
fact that many CWEs map to the same CVE. Therefore, Measuring similarity between a CWE and a CVE
these CWEs will have similar base score. In real life, using NLP: At this point we need to identify how similar
CWEs that map to the same CVE sometimes may not a given CWE is to a CVE. Hence, we need to perform a
operate in a similar way, and hence may not cause similar contextual comparison between the descriptions of CWE
security breach. This diference leads to an error in the and CVE to find out which vulnerability (CVE) is
conbase score calculation. In order to have a more precise textually similar to a weakness (CWE) by using their
base score, we make use of the contextual information descriptions. This is not only important but also critical
to map the CWEs to CVEs; consequently, we minimize since we can get a better idea whether a CVE and a CWE
this error. Also, considering the severity level as the base work in the similar way and cause security breach in
scores for unmapped CWEs generates an error in the the same way, or not. To do this, we create a universal
security score, because the severity level is generic, and vector of all the CVE and CWE descriptions using
sencan vary from one scanning tool to another. However, tence embedding. Sentence embedding is an NLP process
if we use context similarity, we can generate a more that embeds each sentence into an n-dimensional vector
accurate score for any non-mapped CWEs. Therefore, space, by inheriting the features from underlying word
for a more accurate estimation of the security posture embeddings. So, each sentence vector is created based
of the system, we propose to create a general method to on the words in that sentence, and the context of those
generate CWE base score using context similarity, impact, words. After representing two sentences as vectors (for
and exploitability sub-scores for each weakness. example sentences like ‘I drive car’ and ‘I can drive’), we
can calculate their cosine similarity.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>3.2. Proposed Approach</title>
        <p>the two sentences. If similarity(A,B) increases, it means The entire process to calculate the base score is shown
that the angular diference between these two sentences in Algorithm 1.
decreases, meaning the sentences are similar. Using this
process, we perform a semantic search between a CWE Algorithm 1
and a CVE. Semantic search is the task of finding similar Input: the list of the CWE and CVE descriptions
sentences from a corpus to a given sentence, by using Output: CWE base scores
the context of the sentence. Therefore, we conduct a se- 1: Create universal sentence vectors (vectx) for all CWE
mantic search to find the closest CVEs to any given CWE and CVE descriptions using sentence embeddings
based on similarity, which is represented by the sentence 2: for all CWE do
vector. This is represented in lines 1 to 3 in algorithm 1. 3: Find the most similar CVEs by comparing the
senWe use Python NLP libraries to create sentence vectors, tence vectors.
such as the sentence-transformers [28] library. As there
is not enough CVE/CWE description data to train the (, ) (.)
model, and the vulnerability/weakness descriptions are ||||
well formatted, we decided to use a pre-trained model for
sentence embedding. This model is trained on well for- 4: Calculate the average impact and exploitability
matted text, which means it is good at creating a vector sub-scores as in equations 6 and 7
for grammatically and syntactically correct texts, and it 5: Calculate CWE base score using impact and
exworks well for our purpose. We have manually checked ploitability sub-scores as in equation 8
each CWE and its similar CVEs to confirm the accuracy 6: end for
of the model.</p>
        <p>Using the generated sentence vectors for CVE and An advantage of this process is that, unlike the
baseCWE description in the universal vector space, we calcu- line, every CWE gets scored, depending on its similarity
late how much two sentences difer from each other. By with a CVE. Even if a CWE does not properly map to
subtracting their vectors using distance modules from any CVEs, it gets assigned an appropriate score. The
Scipy [29], we get a score that represents how diferent sub-scores of the mapped CVEs are multiplied by the
these two sentences are, i.e. we obtain the distance be- similarity score, which give us the weighted sub-scores
tween them. This distance score is always less than or that help us generate accurate base score for a CWE
equal to 1, (1 signifies that the sentences are completely with respect to the CVE. We can also generate CWE base
diferent, and 0 signifies that both sentences are the same). scores directly from CVE base scores (without generating
By subtracting the distance score from 1, we get the sim- impact and exploitability sub-scores separately, and then
ilarity score between two sentences. Higher similarity adding them) using the same method. But impact and
score means that two sentences are more similar. When exploitability of two CWEs can sometimes vary largely,
we compare the sentence vector of a CWE against the even if they have the same base score. So, it is better to
universal sentence vector of the CVEs, this similarity derive impact and exploitability score first, and then add
score helps us find the most similar CVEs for that CWE. them to generate the base score.</p>
        <p>This is equivalent to the probability of a CWE being re- Finally, when we have all the base weakness scores
lated to a particular CVE. For each weakness in CWE list, for all CWEs, we can calculate the total weakness score
we consider the 5 most similar CVEs. The CVE list [4] of the system by taking the average of their base scores
comes with base score of a CVE, along with its impact using the following equation:
and exploitability sub-score. For each CWE we take the
weighted average of the impact and exploitability sub-   = ∑︀( ) (9)
score of the 5 most similar vulnerabilities, where weight     
is the similarity score between the CWE and the CVE. where, n represents the total number of CWEs in a
sysThe impact score is represented as, tem.</p>
        <p>∑︀5=1( × ) Like the baseline, this weakness score always ranges
 = (6) from 1 to 10. We multiply the resulting score by 10 to get
5 a more conventional score out of 100. It is more
desirAnd the exploitability score is represented as, able for a higher score to reflect a more secure system,</p>
        <p>Therefore, the security score of the system, based on the
 = weakness score is as follows:
∑︀5
=1( × )
5</p>
        <p>(7)</p>
        <sec id="sec-2-2-1">
          <title>Therefore, the base score is given by Equation 8 below.</title>
          <p>_ =  +</p>
          <p>= 100 − 10 ×   (10)</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>This calculation process is represented in lines 4 to 5 in Algorithm 1.</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experimental Evaluation</title>
      <sec id="sec-3-1">
        <title>In order to evaluate our technique, compare it against the most commonly used ones, and to measure the improvement, we performed several experiments and recorded the results using various actual and synthetic datasets.</title>
        <sec id="sec-3-1-1">
          <title>4.1. Datasets</title>
          <p>We applied our method on two real-world open source
software systems (WhatsApp [30] and Atom [31]) and
a synthetic dataset (SARD [32]). After running static
analysis tools on these software systems, we obtain a
list of possible weaknesses. However, this list contains
false positives, which are removed by running the VINCI
tool [23]. The result contains only actual CWEs that
exist in the system (true positives). WhatsApp had 490
actual CWEs, Atom had 160 CWEs and SARD had 30
CWEs. In every system, one CWE might be repeated
multiple times. There are also a lot of CWEs that cannot
be directly mapped with any CVE, especially in the case
of SARD and Atom. The unmapped CWEs add significant
efect on the security score. The base scores of CWEs are
calculated as the methods discussed in section 3.2.
non-mapped CWEs is very large, it can greatly afect the
actual security score. The more non-mapped CWE a
system contains, the more its final security score varies from
the baseline score. So, even though the baseline approach
4.2. Experiments and Results is the most commonly used method, it provides
inaccurate results for most of the cases, especially in the SARD
In our experiments we derive the security posture of a dataset. As it is synthetic, it may not reflect real word
software system using the security score described in our scenarios, even lesser number of distinct CWEs. Among
approach. We used a 64 bit Dell laptop with Windows 10 our datasets, SARD has the least number of unique CWEs,
Pro operating system, Intel(R) Core(TM) i7- 1065G7 CPU and a higher number of CWEs that are non-mapped to a
(2 GHz) with 8GB RAM and 500 GB hard disk. Figure 2 CVE. Therefore, the baseline security score for SARD is
shows a comparison of the final security score baseline very low, resulting in a large diference from the security
method versus two versions of our approach. The first score generated using NLP. We also experimented on
version calculates weakness scores acquired directly by Atom version 1.33.0, which contains a large number of
averaging the base score of similar CVEs, without calcu- unmapped CWEs. We observe that there is a big
diferlating the impact and exploitability sub-scores separately. ence between the baseline score and the final security
The second version calculates the impact and exploitabil- score for Atom too, with a low baseline score and much
ity sub-scores separately, and then adds them to obtain higher security score. The base scores achieved by our
the base score for CWE. Diferent methods of calculating process utilize context and are available for all CWEs. For
security score for diferent systems are represented in the the two methods that we implemented, we observe that,
X-axis, and the security score for each case is represented the one calculating impact and exploitability sub-score
in the Y-axis. The security score ranges from 1 to 100, gives slightly better result than directly calculating the
where higher scores identify better security. base score. This happens because even if two CVEs have</p>
          <p>We observe a large diference of security score be- the same base score, their impact on the system can be
tween the baseline and our proposed methods. The final diferent, and their exploitability can vary too. So, the
security score generated by our proposed method is sig- CWE base score can be diferent sometimes if we
calcunificantly higher than the baseline score. This is due to late impact and exploitability sub-scores of CVE first and
the fact that each CWE in the baseline is mapped to one add them later, as opposed to calculating it directly from
or more CVEs, and sometimes there is no equivalent CVE CVE base scores. The diference of these scores occurs
for a given CWE. For the non-mapped CWEs, the base- for some CWEs, but in some other cases, impact and
exline considers the severity level as the base score, which ploitability sub-scores do not vary much when they have
is often not accurate. This can sometimes increase and the same base score. So, the overall security score, which
sometimes decrease the actual security score, injecting consists of all CWEs, might not always be very diferent
more inaccuracy in the final score. Also, if the number of for these two methods.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion</title>
      <sec id="sec-4-1">
        <title>Bureau of US Department of Commerce.</title>
        <p>It is very important to estimate how secure a software
is, in order to determine how much efort and attention References
should be allocated in terms of having a more secure
software system. The more CWEs a system has, and the [1] MITRE, Common weakness enumeration (cwe),
more severe those CWEs are, the more vulnerable the https://cwe.mitre.org/, 2006. Accessed: 2021-06-28.
system is to attacks. Using the CWEs that exist in a sys- [2] MITRE, 2019 cwe top 25 most dangerous software
tem, we can generate a score that identifies the security errors, 1.https://cwe.mitre.org/top25/archive/2019/
posture of a software system, therefore providing idea 2019_cwe_top25.html, 2019. Accessed: 2021-05-17.
about the vulnerability of the software system against [3] T. M. Corporation, The mitre corporation, https:
cyber attacks. //www.mitre.org/, 1997. Accessed: 2021-08-12.</p>
        <p>In this paper, we have designed and implemented a [4] NVD, Common vulnerabilities and exposures (cve),
method to provide a more accurate software security https://cve.mitre.org/, 1999. Accessed: 2021-06-28.
score based on the existing CWE information. This is [5] NVD, Common vulnerability scoring system, https:
accomplished by scoring each CWE in the software as //www.first.org/cvss/, 1999. Accessed: 2021-06-29.
accurately as possible, utilizing contextual information. [6] MITRE, Common weakness scoring system (cwss),
We have used semantics and NLP to map CWEs to CVEs, https://cwe.mitre.org/cwss/cwss_v1.0.1.html, 2006.
in order to generate a base score for CWE. In order to Accessed: 2020-07-07.
guarantee the correctness of the semantic mapping pro- [7] X. Zhu, C. Cao, J. Zhang, Vulnerability severity
cess, we manually checked the correctness of sentence prediction and risk metric modeling for software,
similarity. In the future, we plan to reduce the need of Applied Intelligence 47 (2017) 828–836.
manual checking by generating and collecting suficient [8] A. A. Younis, Y. K. Malaiya, Using software
strucdata for supervised learning, which will make the process ture to predict vulnerability exploitation potential,
even more automated. in: 2014 IEEE Eighth International Conference on</p>
        <p>In addition, we have evaluated our methods through Software Security and Reliability-Companion, IEEE,
experimentations, where we have compared the results 2014, pp. 13–18.
with the most generally used baseline, which is currently [9] J. A. Wang, F. Zhang, M. Xia, Temporal metrics
the state-of-art [2]. We have tested our proposed method for software vulnerabilities, in: Proceedings of the
and the baseline method on diferent real world and syn- 4th annual workshop on Cyber security and
inforthetic software systems. In every case, our proposed mation intelligence research: developing strategies
method has produced significantly more accurate and to meet the cyber security and information
intellilogical scores compared to the popularly used baseline gence challenges ahead, 2008, pp. 1–3.
method. [10] S. Frei, M. May, U. Fiedler, B. Plattner, Large-scale</p>
        <p>The resulting procedure can help system managers vulnerability analysis, in: Proceedings of the 2006
and developers get an accurate estimation about the se- SIGCOMM workshop on Large-scale attack defense,
curity posture of the system, leading towards more se- 2006, pp. 131–138.
cure software (which are less prone to cyber-attacks), [11] A. Alzaatreh, F. Famoye, C. Lee, Weibull-pareto
and therefore increasing the quality of software. This distribution and its applications, Communications
method for CWE scoring can also be used to prioritize in Statistics-Theory and Methods 42 (2013) 1673–
weaknesses that exist in a system, and determine which 1691.
ones to address first. In a whole, our method can reduce [12] R. Wang, L. Gao, Q. Sun, D. Sun, An improved
the workload relevant to ensuring software security, and cvss-based vulnerability scoring mechanism, in:
help the organizations create more secure software sys- 2011 Third International Conference on Multimedia
tems using less time and resources. Information Networking and Security, IEEE, 2011,
pp. 352–355.
[13] C. Fruhwirth, T. Mannisto, Improving cvss-based
Acknowledgments vulnerability prioritization and response with
context information, in: 2009 3rd International
Symposium on Empirical Software Engineering and
Measurement, IEEE, 2009, pp. 535–544.
[14] T. Fujiwara, J. M. Estevez, Y. Satoh, S. Yamada, A
calculation method for software safety integrity level,
in: Proceedings of the 1st Workshop on Critical
AuThis research has been partially supported by the State
of Maryland through TEDCO Maryland Innovation
Initiative (MII) grant # 0719-003.</p>
        <p>This research has been partially funded by a grant from
the the Ofice of Innovation and Entrepreneurship of
the US Economic Development Administration by the
tomotive applications: Robustness &amp; Safety, 2010, ciation for Computational Linguistics, 2019. URL:
pp. 31–34. http://arxiv.org/abs/1908.10084.
[15] Z. Han, X. Li, Z. Xing, H. Liu, Z. Feng, Learning [29] E. Jones, T. Oliphant, P. Peterson, et al., Scipy: Open
to predict severity of software vulnerability using source scientific tools for python (2001).
only vulnerability description, in: 2017 IEEE Inter- [30] WhatsApp Inc. (Facebook, Inc.), Whatsapp, 2009.
national Conference on Software Maintenance and URL: https://whatsapp.com.</p>
        <p>Evolution (ICSME), IEEE, 2017, pp. 125–136. [31] MIT, Atom, https://atom.io/, 2014. Accessed:
2021[16] T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, 07-16.</p>
        <p>J. Dean, Distributed representations of words and [32] P. Black, Sard: Thousands of reference programs
phrases and their compositionality, in: Advances for software assurance (2017). URL: https://tsapps.
in neural information processing systems, 2013, pp. nist.gov/publication/get_pdf.cfm?pub_id=923127.
3111–3119.
[17] K. O’Shea, R. Nash, An introduction to
convolutional neural networks, arXiv preprint
arXiv:1511.08458 (2015).
[18] H. Xiao, Z. Xing, X. Li, H. Guo, Embedding and
predicting software security entity relationships:
A knowledge graph based approach, in:
International Conference on Neural Information
Processing, Springer, 2019, pp. 50–63.
[19] T. Mikolov, K. Chen, G. Corrado, J. Dean, Eficient
estimation of word representations in vector space,
arXiv preprint arXiv:1301.3781 (2013).
[20] MITRE, Capec, https://capec.mitre.org/, 2007.
Ac</p>
        <p>cessed: 2021-08-10.
[21] Y. Du, Y. Lu, A weakness relevance evaluation
method based on pagerank, in: 2019 IEEE Fourth
International Conference on Data Science in
Cyberspace (DSC), IEEE, 2019, pp. 422–427.
[22] A. N. Langville, C. D. Meyer, Google’s PageRank</p>
        <p>and beyond, Princeton university press, 2011.
[23] F. Cheirdari, G. Karabatis, Analyzing false
positive source code vulnerabilities using static analysis
tools, in: 2018 IEEE International Conference on</p>
        <p>Big Data (Big Data), IEEE, 2018, pp. 4782–4788.
[24] M. N. Alsaleh, E. Al-Shaer, Enterprise risk
assessment based on compliance reports and vulnerability
scoring systems, in: Proceedings of the 2014
Workshop on Cyber Security Analytics, Intelligence and</p>
        <p>Automation, 2014, pp. 25–28.
[25] J. A. Wang, H. Wang, M. Guo, M. Xia, Security
metrics for software systems, in: Proceedings of
the 47th Annual Southeast Regional Conference,
2009, pp. 1–6.
[26] N. H. Sherief, A. A. Abdel-Hamid, K. M. Mahar,</p>
        <p>Threat-driven modeling framework for secure
software using aspect-oriented stochastic petri nets, in:
2010 The 7th International Conference on
Informatics and Systems (INFOS), IEEE, 2010, pp. 1–8.
[27] NIST, National vulnerability database, https://nvd.</p>
        <p>nist.gov/vuln, 1999. Accessed: 2021-06-21.
[28] N. Reimers, I. Gurevych, Sentence-bert:
Sentence embeddings using siamese bert-networks, in:
Proceedings of the 2019 Conference on Empirical
Methods in Natural Language Processing,
Asso</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>