<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>PEVuln: A Benchmark Dataset for Using Machine Learning to Detect Vulnerabilities in PE Malware</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nathan Ross</string-name>
          <email>nross12@qub.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oluwafemi Olukoya</string-name>
          <email>o.olukoya@qub.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesús Martínez del Rincón</string-name>
          <email>j.martinez-del-rincon@qub.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Domhnall Carlin</string-name>
          <email>d.carlin@qub.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Secure Information Technologies (CSIT), Queen's University Belfast</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2000</year>
      </pub-date>
      <fpage>6</fpage>
      <lpage>35</lpage>
      <abstract>
        <p>In this paper, we present a benchmark dataset for training and evaluating static PE malware machine learning models, specifically for detecting known vulnerabilities in malware. Our goal is to enable further research in defense against malware by exploiting their bugs or weaknesses. After recognising limitations in current malware datasets regarding exploitable malware, our dataset addresses these gaps by utilizing the malware vulnerability database Malvuln, and software vulnerability database ExploitDB to create a new malware dataset with 684 vulnerable malware samples, 35,241 non-vulnerable malware samples, 1,425 vulnerable benign samples, and 7,905 non-vulnerable benign samples, detailed with timestamps, families, threat mapping, vulnerability mapping, and obfuscation analysis. This 4-class dataset lays the foundation for advancing future research in analysis and vulnerability exploitation in malware using machine learning. We also provide baseline results using state-of-the-art models for malware classification to benchmark the performance of the dataset, where the binary tasks achieve F1 scores above 0.90, while the multi-class task attains an F1-Score of 0.958.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Dataset</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Malware</kwd>
        <kwd>Vulnerabilities</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>but will support the identification of vulnerabilities within them. As an initial step, this approach will
not only detect but also stop the spread of malware once it is in the system.</p>
      <p>This paper addresses an existing shortfall in malware vulnerability datasets by creating the first
comprehensive dataset for static analysis of Windows portable executable (PE) malware vulnerabilities
for detection and analysis. Our dataset contains 45,255 samples comprised of malware and benign
samples that are labeled vulnerable and non-vulnerable to allow for the development of ML models for
detecting malware, vulnerabilities, or both. The creation of this dataset aims to underpin innovative
approaches to enhance the security of cyber systems.</p>
      <p>The main contributions of this paper are:
• A comprehensive and curated malware vulnerability dataset, which is publicly available1.
• Extensive annotations on the data including timestamps, families, threat details,
vulnerability details, CWEs, packer types and complexities, vulnerability payloads, and Mitre ATT&amp;CK
framework mapping.
• A detailed feature set ready to be used that utilizes the well-spread Ember feature extraction2,
which provides greater flexibility to researchers and developers.
• Data visualization of our dataset was performed using dimensionality reduction techniques to
depict the data distribution in our dataset.
• Timestamp collection, obfuscation and vulnerability analysis, and CWE and ATT&amp;CK mappings
between benign and malware samples, were performed for accurate annotation and richer analysis
of the dataset and potential defenses.
• A comprehensive ML benchmark using our dataset for binary (malware vs. benign, vulnerable vs.</p>
      <p>non-vulnerable) and multi-class classification tasks.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>In this section, we describe approaches for identifying vulnerabilities within malware in Section 2.1,
the application of machine learning techniques for malware classification and vulnerability detection in
Section 2.2, and finally, a comparative analysis of relevant datasets used in malware classification and
detection, which emphasizes the contribution of our dataset in Section 2.3.</p>
      <sec id="sec-2-1">
        <title>2.1. Malware Vulnerabilities</title>
        <p>
          Research into vulnerabilities within malware is a promising yet under-explored domain. The idea that
malware can harbour exploitable flaws makes for an imperative defensive mechanism [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. In 2010,
Caballero et al.[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] introduced a novel approach, termed stitched dynamic symbolic execution, aimed at
uncovering vulnerabilities within the malware. They aimed to facilitate the discovery of exploitable
bugs as a significant step in automated malware analysis. Caballero et al. identified six distinct bugs
in four common malware families, highlighting their persistence across multiple evolutions. These
weaknesses were exploitable to disrupt or at least hinder their operations defensively. These findings
were corroborated in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], where the authors highlight that malware is prone to exploitable bugs, just
like benign software. The authors suggest that most malware does not go through quality control
processes, meaning not only are they vulnerable to common issues during the software development
lifecycle, but exposed to them in a greater manner. Defensive exploitation of such vulnerabilities can
allow defenders to delay, mitigate or frustrate malware-based attacks.
1https://github.com/nross12/PEVuln/
2https://github.com/elastic/ember
        </p>
        <p>
          Given that these vulnerabilities exist, they can be identified similarly to how software is labelled in
lists such as MITRE Common Weakness Enumeration (CWE)3. The studies in [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ] also demonstrate
that no proper quality assurance is carried out when developing malware since it contains multiple
bugs persistent in families across a timeline. This exemplifies how inherited flaws within malware can
be exploited defensively. The research in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] also highlights a bug that crashes Command &amp; Control
(C2) servers in the Mirai code, and how the vulnerability persists in many variants. Similar case studies
have shown that a DLL-hijack vulnerability was found and exploited in the LockBit ransomware [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
which worked against nearly every other ransomware family. This vulnerability has been likened to
a Pandora’s box of vulnerabilities [
          <xref ref-type="bibr" rid="ref13 ref14 ref15">13, 14, 15</xref>
          ]. Apart from the lack of quality control, code reuse is
likely another reason for the persistence of malware vulnerabilities. A study on malware evolution
and code reuse over four decades[
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] found many instances of code reuse. Findings from research
conducted by Intezer and McAfee, which involved analyzing malware samples and cyber campaigns,
revealed substantial evidence of code reuse spanning 10 years (2007-2017) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. This investigation
helped uncover previously unknown connections among North Korea’s malware families, indicating
that the reuse of malware code is widespread in cybercrime. The lack of quality control, the rise of
malware variants, code reuse, and malware-as-a-service ensure that malware vulnerabilities remain.
        </p>
        <p>
          One of the well-known cases of using exploitable vulnerabilities in malware for ofensive security
in detection technology is WannaCry, the biggest ransomware attack in history, which spread within
days to more than 250,000 systems in 150 countries and was stopped by registering a web domain
found in the malware’s code [18]. Once the ransomware checked the URL and found it was active, it
was shut down – buying precious time and giving organizations room to update their systems. Such
vulnerabilities can often persist in malware and its variants for a long time across diferent target
platforms [
          <xref ref-type="bibr" rid="ref11">11, 19</xref>
          ]. Similar studies have identified and exploited flaws in ransomware encryption
techniques, assisting victims of ransomware and defending against the threats [20, 21, 22]. Other studies
have investigated the identification and exploitation of malware vulnerabilities for covert monitoring
of C&amp;C servers through protocol infiltration for botnet disruption and breakdown [ 23, 19, 24, 25]. To
develop an automated system that rapidly uncovers exploitable flaws and defects in malicious software
as a kill-switch approach, PE malware machine learning models for vulnerability identification and
detection could be designed and utilized, under the assumption that benchmark annotated datasets are
available for training.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Machine Learning</title>
        <p>ML has become a cornerstone for malware detection and classification with various techniques being
applied by researchers to identify patterns and behaviors that can discern malware from typical benign
software [26, 27, 28]. ML models can be trained on features derived from three primary analytical
methods: static, dynamic and hybrid analysis [29]. Static analysis involves examining an executable
without explicitly executing it, making it one of the safest and most eficient methods to obtain
information [30]. Dynamic Analysis, on the other hand, allows you to gather features based on
the output observed by the behavior of the executable during runtime on the system [31]. This method
is more intensive on the system and can take much longer to generate data, especially when working
with a large dataset. Hybrid analysis is a time-consuming and resource-intensive method that combines
the strength of static and dynamic techniques to provide a comprehensive understanding of malware
samples[32]. Additionally, there are memory analysis techniques that create a memory image of the
malware during dynamic execution for analysis [19, 33, 34]. This paper proposes specifically the
extraction of features by static analysis of PE files.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Datasets</title>
        <p>Machine learning models for malware detection rely heavily on datasets that are curated, annotated and
comprehensive. Many benchmark datasets are instrumental for malware detection, which can be seen
Features
in Table 1, but they lack vulnerabilities and metadata within these datasets on vulnerable executables.
Additionally, a comparison of the features, vulnerable samples and metadata proposed in our dataset
and other open datasets for PE malware is presented in Table 1.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset Description</title>
      <p>We created this dataset by systematically collecting data from several malware repositories, data sources,
and techniques which are categorized into four main classes as outlined below:
• Vulnerable Malware (VM): The samples and metadata for vulnerable malware were collected
from four main sources - Malvuln4, VirusTotal5, VulDB6 and VirusShare7. Malvuln is a resource
dedicated to malware security vulnerability research and provides information on identifying
and exploiting malware, including details such as the threat, vulnerability, description, family,
hash, exploit POC, etc. The metadata of each vulnerable malware sample is collected, and the
binary executables are downloaded from VirusTotal and VirusShare using the MD5 of the sample
and VulDB using the MVID allocated by Malvuln.
• Non-vulnerable Malware(NVM): An abundance of malware was collected, including samples
caught by honeypots and from large data dumps on VirusShare. Similarly to the vulnerable
malware data, the non-vulnerable malware data contains additional information extracted about
each sample from VirusTotal and VirusShare.
• Vulnerable Benign Software(VB): We leveraged Exploit-DB8, which is an exploit database similar
to Malvuln, except that it holds labelled conventional software vulnerabilities with options for
downloading the vulnerable application, metadata, and exploitation payload.
4https://malvuln.com/
5https://www.virustotal.com/
6https://vuldb.com/
7https://virusshare.com/
8https://www.exploit-db.com/
• Non-vulnerable Benign Software (NVB): To create a dataset of benign applications, we extracted
executables from a fresh installation of Windows 10 (located in C:\Windows\System32). This
is a popular approach in the malware community for creating a benign dataset [30]. Further
vulnerability scans were carried out to ensure that the executables were not vulnerable.
The data sources and types of information collected are summarized in Table 2. Creating an augmentation
of an existing dataset or using a combination of multiple sources to create a dataset is consistent with
the literature. For instance, an enhanced version of the EMBER dataset, EMBERSim, was developed
in [35] to include similarity information, addressing the problem of binary code similarity search in
Windows PE files. The expanded tags were created using a combination of automatic tagging tools
such as AVClass[36, 37] to include class, family, behavior, and file properties. In this research, we
utilized AVClass for tagging our data which can provide additional context for addressing the issue of
vulnerabilities in PE malware by highlighting the prevalence of specific vulnerabilities across diferent
threat types or families. Table 1 expands on these classes with details on the number of samples in the
classes and the number of unique families, vulnerabilities, and CWEs. The vulnerable malware subset
is naturally the minority class given how little analysis exists of vulnerabilities in malware samples.
This results in a highly imbalanced dataset, and may preclude specific training strategies as we will
describe in Section 4.</p>
      <p>This dataset will be regularly updated as new vulnerable PE malware and benign samples are collected
from the repository.</p>
      <sec id="sec-3-1">
        <title>3.1. Framework Overview</title>
        <p>
          We have developed a systematic approach, depicted in Figure 1, for the creation of this malware
vulnerability dataset, from meticulously seeking out raw data that has gone through various stages
of processing and cleaning, followed by ML model training catering to both binary and multi-class
classification tasks to obtain preliminary results and data visualization. Firstly, both vulnerable and
nonvulnerable malware and benign software executables, with their corresponding metadata, were sourced
and subsequently pre-processed, ensuring that clean and well-structured outputs capture the attributes
and features necessary to perform a deeper analysis of the performance of the dataset. In the executable
#
#
#
#
pre-processing phase, feature extraction is performed using Ember[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] to derive meaningful features
from the executables for each class. Obfuscation scanning and static application security testing (SAST)
were performed to obtain additional information about the executables. Simultaneously, metadata
pre-processing focuses on the supplementary data associated with the executables. This involves
mapping hashes such as md5 and sha256 where appropriate so that the samples can be identified and
extracting timestamps, threat types, families, and various vulnerability details (such as the vulnerability
type and the payload to exploit it). Additionally, the Mitre ATT&amp;CK framework and common weakness
enumeration (CWE) were mapped, and obfuscation details (such as the types of packer used) were
analyzed and categorized based on complexity. To generate an ML benchmark for the community, as
well as to prove useful models can be generated from our dataset, a comprehensive set of common
machine learning algorithms were trained, including LightGBM (LGBM), Random Forest (RF), k-Nearest
Neighbors (KNN), Support Vector Machines (SVM), and Artificial Neural Networks (ANN). These models
are trained to perform both binary and multi-class classification tasks to distinguish between benign
and malware data and further categorize if it is vulnerable or not. Full details on the ML baseline
models are given in Section 4. Finally, we also visualize our dataset and its analysis in Section 3.7. This
presents detailed visualizations of the full dataset by applying dimensionality reduction techniques,
namely Principal Component Analysis (PCA) and t-SNE. These techniques allow us to produce 2D
representations of our large dimensional data. Data samples are depicted using diferent labels and
colors according to the previously mentioned metadata including vulnerabilities, obfuscation details,
CWEs, etc. Furthermore, we have visualized the Mitre ATT&amp;CK mappings to better understand the
attack surfaces in which the vulnerable malware targets develop relationships between these attack
surfaces and the vulnerabilities present.
        </p>
        <sec id="sec-3-1-1">
          <title>Dataset Collection and Pre-processing</title>
          <p>Data Sources
Malvuln
ExploitDB
VirusTotal
VirusShare
VulDB
Honeypots
PE pre-processing
Feature Extraction
Obfuscation Scanning
Static Application Security Testing
Metadata pre-processing
Timestamps
Hashes (md5 and sha256)
Threat Type
Family
Vulnerability Type
Vulnerability Details
Vulnerability Payload
CWEs
Mitre ATT&amp;CK
Obfuscation Details (Type/Complexity)
Train/Test
models</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Dataset Visualisation</title>
          <p>VM &amp; VB
Vulnerability Types
CWEs
Obfuscation Details
VM:
ATT&amp;CK
All: VM, VB, NVM, NVB
Obfuscation Details</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Preliminary Model Benchmarks</title>
          <p>LGBM
Random Forest
KNN
SVM
ANN
Binary Classification
VM vs VB
VM vs NVM
VM vs NVB
VB vs NVB
Multi-class Classification
(All Classes)</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. SAST Tools Scanning</title>
        <p>When working with ML, the data needs to be curated and precisely annotated to avoid outliers, training
discrepancies and even potential poisoning and/or backdoor attacks. In our case, this is especially
relevant regarding the vulnerabilities that could theoretically exist in the non-vulnerable data as it
had not gone through rigorous testing in the compilation of this dataset. So for that, Source Code
Analysis Tools (or Static Application Security Testing, SAST) were used on the non-vulnerable class
executables to find potential security flaws through the means of signature-based pattern matching,
semantic analysis, and taint analysis. This allows for the removal of data that can be ruled out as
vulnerable from the non-vulnerable classes which, in theory, should make the vulnerable classes more
prominent when training models. To perform this SAST on the data, two tools were used; cve-bin-tool9
and Vulnscan10.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Packing in Vulnerable &amp; Non-vulnerable Malware</title>
        <p>As we are working with executables, particularly user software and malware, they can be afected by
evasion techniques, more specifically in our case obfuscation. To check for this, we used two tools:
Bintropy[38], a Python tool that detects obfuscation based on entropy, and Detect It Easy (DIE)11 which
works similarly to Bintropy12 but also provides information on the packer, linker, and compiler. It
is possible to categorize the diferent types of packers based on their complexities which can be seen
in [39]. The types of packers range from Type-I to Type-VI, with higher values indicative of a greater
complexity. This method of categorization provides additional information which allows us to conduct
a more nuanced analysis of the diferent packers used in malware compared to benign software. This is
critical as it allows an investigation into whether malicious actors lean towards using more complex
packers with specialized techniques over others for evasion. These patterns could also inform us of a
pattern or correlation between the use of specific packers in malware with vulnerabilities compared to
those with no vulnerabilities.</p>
        <p>As shown in Table 4 and Table 5, the most often used packers in the vulnerable and control classes of
the benign and malicious samples are rather simple, spanning from Type-I to Type-III based on the
packer taxonomy proposed by Ugarte-Pedrero et al.[39]. This packer distribution result is consistent
with multiple longitudinal studies [39, 19, 40, 30] investigating the complexity of custom and
of-theshelf run-time packers in the wild. The implication of this study is two-fold. Firstly, packer complexity
in control (non-vulnerable) and vulnerable malware follows the same pattern, making it possible to
use analysis and detection techniques for packed malware to packed vulnerable malware. Secondly,
there is an overlap between packers used in vulnerable and controlled malware samples, so we can
9https://github.com/intel/cve-bin-tool
10https://github.com/zhutoulala/vulnscan
11https://github.com/horsicq/Detect-It-Easy
12https://github.com/packing-box/bintropy
# V-Benign
# NV-Benign
1
0
0
1
3400
develop classifiers that do not associate specific packers with vulnerability. This finding is significant,
as research on the impact of machine-learning-based malware detection on packed samples utilizing
static analysis features has observed that classifiers frequently link particular packers to malicious
activity because there is insuficient overlap between packers utilized in malicious and benign samples
[30]. Additionally, observations of the count of unique packers in the control benign are restricted to
one, which can be attributed to how we collected the data. That being, Windows 10 likely does not
need to implement various packers as their executables and DLLs are distributed and managed with
the Operating System (OS), and their consistent employment of UPX for all packed executables and
DLLs is likely driven by several factors13; mainly concerning maintaining consistency and widespread
applicability across various executables and DLL files. We can therefore conclude that packers are not a
feature for determining the nature of an executable as controlled or vulnerable malware.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. CWE Mapping</title>
        <p>For richer vulnerability analysis, we have mapped the CWEs for both the vulnerable malware and
vulnerable benign classes in this dataset and conducted strategic similarity analysis between them so
that we can understand the crossover between vulnerabilities present in benign software or malware.
By doing this, we can further analyze the correlation between the vulnerabilities with other attributes
such as system architecture, programming languages used, attack vectors, or exploitation complexities.
Mapping CWEs in vulnerable malware is also crucial for linking vulnerabilities across certain malware
families to demonstrate both the persistence of vulnerabilities in newer malware of the same family and
the possible introduction of newer vulnerabilities. Another advantage of CWE mapping for malware
vulnerabilities is that it enables the creation of automated exploits based on the identified weakness
category[41]. This is similar to the automated exploit generation for vulnerabilities seen in commercial
and open-source programs, such as Stack-based bufer overflow( CWE-121)[42, 43], PHP object injection
vulnerabilities(CWE-502, CWE-915)[44] and XML injection vulnerabilities(CWE-91)[45].</p>
        <p>In Table 6, an analysis of comprehensive weakness categories is presented for vulnerable malicious
and benign samples. The analysis reveals that the top 10 most frequent vulnerability categories in the
vulnerable malware dataset have a 20% similarity with the top 10 most frequent vulnerabilities in the
vulnerable benign dataset, with CWE-200 and CWE-120 being the common vulnerability categories.
This suggests that both classes contain vulnerabilities that can lead to information exposure and bufer
overflow. The findings provide insight into the categories of vulnerabilities present in vulnerable
malware and their prevalence. Additionally, the analysis highlights the overlap and diferences
between the vulnerabilities in legitimate software applications and malicious binaries, indicating that
vulnerable malicious binaries often exhibit diferent categories of weaknesses compared to their benign
counterparts.</p>
        <p>
          Approximately 71.69% of vulnerabilities in the malware samples are attributed to five main classes of
weaknesses. These classes include Permission Issues, Hidden Functionality, Bufer Overflows , Improper
Authentication, and Use of Hard-coded Credentials. They have consistently been the primary sources of
vulnerabilities, making them favoured targets for defenders seeking to exploit these security issues.
The Permission Issues category refers to weaknesses related to improper assignment or handling of
permissions. A comprehensive study on covert monitoring of C&amp;C servers [19] found that
overpermissioned protocols are prevalent in the malware operational landscape, with nearly 1 in 3 malware
bots exhibiting this vulnerability, confirming the significance of this weakness category. Another
study on IoT botnets [23] highlighted weak and default passwords in C2 servers, aligning with the
identification of Improper Authentication and Use of Hard-coded Credentials as primary sources of
malware vulnerabilities. It is also expected that CWE-912, Hidden Functionality, is prevalent in our
dataset, as it can take the form of embedded malicious code and is useful for attacks that modify
the control flow of the application, aligning with malicious behavior. Another noteworthy weakness
category in the Top 10 for malware is CWE-426, a weakness category associated with DLL Hijacking,
which is prevalent and exploitable in most ransomware families for preventing file encryption [
          <xref ref-type="bibr" rid="ref13 ref14 ref15">13, 14, 15</xref>
          ].
        </p>
        <p>In benign executables, 76.64% of vulnerabilities are attributed to two main classes of weaknesses. The
primary sources of vulnerabilities are Improper Restriction of Operations within the Bounds of a Memory
Bufer and Improper Input Validation. These weaknesses are commonly targeted by threat actors when
trying to exploit security issues. Our ranking, based on our dataset, is further supported by the 2023 CWE
Top 10 KEV weaknesses14, which is a catalogue of Known Exploited Vulnerabilities maintained by the
Cybersecurity and Infrastructure Security Agency (CISA)15 for vulnerability management prioritization.
Notably, four of the Top 10 CWEs identified in the vulnerable benign dataset, CWE-787 (#3), CWE-20
(#4) and CWE-22 (#9), are among the top 10 class of weaknesses exploited by threat actors as observed
in the wild. This suggests that the vulnerabilities and prevalence identified in the dataset accurately
reflect real-world observations.</p>
        <p>
          We have further conducted a detailed analysis of the weakness categories present in the vulnerable
samples by aligning the weaknesses with the OWASP Top 1016 and the 2023 CWE Top 2517, as shown
in Table 7. The identified malware CWEs correspond to 8 of the OWASP Top 10 application security
risks, whereas benign CWEs correspond to 6 categories. Similarly, the malware CWEs align with 9 of
the top 25 most dangerous software weaknesses, while benign CWEs align with 8 categories. The CWE
classifies vulnerabilities in benign and malicious binaries into the same four categories in the OWASP
Top 10. This indicates that vulnerability detection and identification tools developed for these categories
in benign software may apply to malicious binaries due to the overlap. For example, the technique
of symbolic execution, used in TaintScope[46] to find bugs in benign software, was also applied to
identify bugs in prevalent families of bots and other malware in a study by Caballero et al[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Moreover,
the CWE from both classes sometimes maps to unique categories of critical risk in the OWASP Top
10 without overlapping. This supports the initial assertion about the diferences in vulnerabilities in
both samples. For instance, only the malware samples have vulnerabilities associated with A02:2021
Cryptographic Failures. One characteristic of this category is a failure in the encryption mechanism,
which is significant as one common vulnerability exploited in ransomware is encryption failures[ 21, 22].
When the first ransomware bug bounty operation was launched, Locker Bugs, which covers encryption
errors, was prioritized[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. This analysis demonstrates that although there is awareness of software
vulnerabilities, it is essential to recognize that malware possesses inherent vulnerabilities that are critical
and impactful. These vulnerabilities can be exploited for ofensive security in defense technology.
14https://cwe.mitre.org/top25/archive/2023/2023_kev_list.html
15https://www.cisa.gov/known-exploited-vulnerabilities-catalog
16https://owasp.org/www-project-top-ten/
17https://cwe.mitre.org/top25/index.html
18https://nvd.nist.gov/vuln/categories
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. MITRE AT T&amp;CK Mapping</title>
        <p>The Mitre ATT&amp;CK framework is essential in understanding and leveraging tactics and techniques
for both vulnerability exploitation and defense mechanisms. Tactics19 within the ATT&amp;CK framework
represent the strategic objectives or the “why" behind an adversary’s actions, whereas techniques20
detail the “how" by describing the specific methods used to accomplish these tactical goals. With this
19https://attack.mitre.org/tactics/enterprise/
20https://attack.mitre.org/techniques/enterprise/
knowledge, potential attack paths in malware can be established and exploited, or conversely can
be used for counteracting adversarial actions in benign software. In our dataset, these tactics and
techniques are mapped according to the vulnerabilities in the vulnerable malware class, which can be
seen in Table 8.</p>
        <p>Table 8 shows the top ATT&amp;CK Tactics and Techniques adopted in the vulnerable malware, which
are overwhelmingly File and Directory Permissions Modification, Obtain Capabilities (malware), Brute
Force (Password Guessing), Hijack Execution Flow and Valid Accounts (Default Accounts) and makes
up over 82% of the techniques. These techniques show that the adversaries want to infiltrate the
victim’s network, evade detection throughout the compromise process, establish resources they can use
to support operations, steal account credentials, obtain higher-level permissions and maintain their
foothold. It is important to note that the tactics and techniques in Table 8 are only representative of our
dataset, which is influenced by the number of malware families and samples. Under diferent conditions,
the order will be diferent. In any case, the present tactics and techniques of vulnerable malware are the
ones you would expect from malware generally [47, 48], which shows that the vulnerability in malware
is not some sophistication of tactics and techniques. Similarly to every piece of software code, malware
is also prone to vulnerabilities and weaknesses that can be exploited. To find the TTPs that were
present in over 600,000 malware samples that were gathered between January 2023 and December 2023,
Picus Labs [48] conducted an analysis. The study identified the top 10 MITRE ATT&amp;CK techniques,
which difer from the top 10 ATT&amp;CK techniques used in this study’s vulnerable malware samples.
A related study that looked into ATT&amp;CK Trends and Techniques in 951 Windows malware families
between 2017 and 2018 revealed that the examined dataset had a varied number of observations for
each ATT&amp;CK [47]. For example, only three of the dataset’s top ten ATT&amp;CK approaches were adopted
by threat groups and malware, according to the report by Picus Labs. Top ATT&amp;CK techniques can
vary depending on the specific malware family being considered, as diferent families have diferent
why and how due to their tactical goals and actions. For instance, the techniques used by ransomware
will difer from those used by spyware. For example, the Centre for Threat Informed Defense, MITRE
Engenuity [49], analyzed 22 ransomware groups over 3 years and compiled a list of the Top 10 ATT&amp;CK
techniques. As anticipated, the Top 10 ATT&amp;CK Techniques for ransomware only had three techniques
(T1486 - Data Encrypted for Impact, T1027 - Obfuscated Files or Information, T1055 - Process Injection) in
common with the top ATT&amp;CK techniques identified by Picus Labs [ 48] and the trends in Windows
Malware [47].</p>
        <p>
          Vulnerable malware is still malware and adheres to the same characteristics and propagation strategies.
Therefore, its vulnerabilities can persist in malware families and variants for a long time because the
authors either do not recognize the vulnerabilities, have poor operational security, or just lack a quality
control process and are more likely to have a bug that can persist in malware families and variants
for a long time, which persists as a notable observation that we have highlighted by Singh in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
Vulnerable or not, malware tactics and techniques are still adversarial. To prioritize ATT&amp;CK techniques
for defending against malware attacks, creating a top ATT&amp;CK techniques list in Table 8 can be a
helpful starting point. This approach is similar to MITRE’s methodology21, which takes into account
the prevalence of techniques, common attack choke points, and actionability to help defenders focus
on the most relevant techniques for their organization. Including the number of observations per
technique allows us to measure how frequently an attacker uses a specific MITRE ATT&amp;CK technique,
which can be useful in identifying important techniques when dealing with vulnerable malware for
exploitation. Additionally, defenders have the opportunity to mitigate or defend against each ATT&amp;CK
based on publicly available threat intelligence derived from real-world observations. MITRE ATT&amp;CK
ofers additional information for mitigation, detection, procedure examples, and references for each
identified significant technique, which can serve as a knowledge base for ofensive security and detection
technology.
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6. EMBER Feature Set</title>
        <p>
          For the data to be used with ease as a benchmark for classification models, Ember, an open-source dataset
and feature extraction research project that uses the LIEF22 project to extract features from PE files for
static malware detection, was utilized to systematically generate a comprehensive set of vectorized
features. The resulting feature vector per sample contains 2381 elements, using the version 2 extractor,
from the PE files in our data. This feature set is widely applied in the ML malware detection literature
[35]. Thus, the extraction of static EMBER features from our dataset is consistent with other benchmark
datasets for malicious PE detection in the literature[
          <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
          ]. For the features extracted using Ember, data
normalization is applied to allow for the values between the diferent classes to be distributed across a
common scale. This allows us to work with data within the desired range without losing much of the
original distribution. The MinMaxScaler is the method used to normalize input features to the [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ]
range based on the training data. This normalization technique has been applied for feature embedding
when working with other benchmark PE malware datasets such as the EMBER dataset to enhance ML
training [50, 51].
        </p>
      </sec>
      <sec id="sec-3-7">
        <title>3.7. Visualizations of Data Distributions</title>
        <p>To visualize the data distribution and have a better understanding of our dataset, Principal Component
Analysis (PCA) was applied. PCA is a statistical method that reduces the dimensionality of large datasets,
in our case our complexity was derived from the 4 classes and a vector length of 2381 from the Ember
features, by transforming them into smaller linear and uncorrelated variables, known as principal
components. By reducing the dimensionality to only 2 variables (the two first principal components),
the dataset can be visualized as an image. Similarly, t-distributed Stochastic Neighbor Embedding
(t-SNE)[52] was also used, by employing an unsupervised, non-linear reduction of dimensionality.
However, unlike PCA, t-SNE preserves the local structure of the data which allows for more visible
clusters.</p>
        <p>From Figure 2, it is clear that there is a discernible separation between the malware and benign
software classes. Notably, there is an overlap between the vulnerable and control malware. This is to be
expected as the patterns and behavior in malware should be consistent between samples. The same is
expected and seen in the vulnerable and control benign samples. There is, however, an existing overlap
between the vulnerable malware and vulnerable benign samples. This raises the assumption that PCA
is categorizing these two classes based on what we hope to be their vulnerabilities.</p>
        <p>The analysis derived from Figure 2 translates well to Figure 3. However, there is even more discernible
separation, especially between the diferent samples within each class, seen from the prominent clusters
formed. Similar to before, the malware and benign software classes are separated. However, the
grouping for vulnerable malware and vulnerable benign software is more noticeable, which again we
can assume to be features that capture their vulnerabilities.
21https://top-attack-techniques.mitre-engenuity.org/methodology
22https://github.com/lief-project/LIEF
3
2
0
1
2 1
A
C
P</p>
        <p>Visualization of dataset using PCA</p>
        <p>From those visualizations, it can be concluded that our dataset represented as Ember feature vectors
will allow for the efective separation of malware and benign samples and the training of a malware
detector. On the contrary, the identification of vulnerabilities within the program may result in a more
challenging problem.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Model Benchmarks</title>
      <sec id="sec-4-1">
        <title>4.1. Model Setup</title>
        <p>As previously described, a comprehensive set of common ML-based models were used, including
LightGBM (LGBM)23, Random Forest (RF)24, K-Nearest Neighbors (KNN)25, Support Vector Machine
(SVM)26, and an Artificial Neural Network (ANN), specifically a Multi-layer Perceptron (MLP) 27. By
using these models, which operate on diferent principles and assumptions, we can observe distinct
perspectives on the dataset and provide a comprehensive set of baseline models that the scientific
community can use to benchmark their novel approaches.</p>
        <p>In our experimental setup, these classifiers were trained initially with an 80:20 train and test split.
Additionally, we implemented stratified K-fold 28 as a way to mitigate any overfitting that may exist
in our models and provide a more comprehensive and realistic evaluation. The reasoning for using a
23https://lightgbm.readthedocs.io/en/stable/
24https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html
25https://scikit-learn.org/stable/modules/generated/sklearn.neighbors.KNeighborsClassifier.html
26https://scikit-learn.org/stable/modules/svm.html
27https://scikit-learn.org/stable/modules/generated/sklearn.neural_network.MLPClassifier.html
28https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html</p>
        <p>Visualization of dataset using t-SNE
stratified K-fold ensures that each fold maintains approximately the same percentage of samples for
each target class in the entire dataset, especially considering the class imbalance that exists where the
vulnerable malware class is underrepresented. We used a k-fold of “K=5" and the cross-validation model
provides train/test indices to split data into train/test sets as a default. After the results were calculated,
we took a mean of each estimator to obtain the final values. Performance metrics such as accuracy,
weighted F1-score, weighted precision, and weighted recall were calculated to evaluate the eficacy of
the models on both default parameters and hyperparameter optimization. As for stratified K-fold, the
accuracy metric was replaced with balanced_accuracy, defined as the average of recall obtained on each
class, specifically designed for dealing with unbalanced classes.</p>
        <p>A set of ten diferent tasks were addressed and validated using our dataset, which includes:
• Vulnerable Malware vs Non-vulnerable Malware in Table 12
• Vulnerable Malware vs Vulnerable Benign in Table 13
• Vulnerable Malware vs Non-vulnerable Benign in Table 14
• Vulnerable Benign vs Non-vulnerable Benign in Table 15
• Vulnerable Benign vs Non-vulnerable Malware in Table 16
• Non-vulnerable Benign vs Non-vulnerable Malware in Table 17
• Vulnerable Malware vs All Benign (VB+NVB) in Table 18
• Vulnerable Benign vs All Malware (VM+NVM) in Table 19
• Malware (VM+NVM) vs Benign (VB+NVB) in Table 10
• Multi-class (All four classes: VM, NVM, VB and NVB) in Table 11</p>
        <p>
          Hyperparameter optimization was applied to provide the best possible performance for each machine
learning technique. While some studies use default settings and parameters, such as those in EMBER[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ],
a better approach involves extensive hyperparameter tuning, as seen in the UCSB Packed Malware
dataset[30]. Our model parameter settings align with these approaches. Subsequently, we implemented
hyperparameter optimization using gridsearch29 for all models. These parameters can be seen in Table 9.
Additionally, we used the optimized parameters when implementing the estimators for cross validation.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Binary Classification</title>
        <p>
          Firstly, combining all classes into a binary task (malware vs. benign) allows us to benchmark our
dataset against others that exist in [
          <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7">4, 5, 6, 7</xref>
          ] for malware detection. This not only aligns with the
conventional practices seen in the literature but also simplifies this classification task so that we can
analyze a more straightforward model performance.
29https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Multi-class Classification</title>
        <p>Subsequently, we can compare the results from Table 10 with the results we have obtained in Table 11
to discern between how the ML classifiers distinguish between an added complexity of vulnerabilities.</p>
        <p>
          Across all tasks and techniques, the LGBM model consistently outperformed others, achieving an
average F1-Score of between 0.929 and 0.999. Albeit, all models demonstrated a strong performance,
especially considering the notable variance in class sizes and similar data structure between the related
classes, that being either the vulnerable and non-vulnerable malware or benign classes. This tells us that
these models can find patterns in the data that can recognise if a sample is either 1) malware or benign
and 2) vulnerable or non-vulnerable. One alternative method for establishing model benchmarks could
involve utilizing representation learning[
          <xref ref-type="bibr" rid="ref18 ref19">53, 50, 54, 55</xref>
          ]. This approach aims to diferentiate samples
that share similarities in the feature space. This could be a viable research direction as the t-SNE plots
in Figure 3 show some overlap among the VM, VB, and NVM classes.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Conclusion</title>
      <sec id="sec-5-1">
        <title>5.1. Finding Vulnerabilities in Malware: Applications and Ethical Considerations</title>
        <p>
          The public disclosure of vulnerabilities in commercial and open-source programs has sparked
conversations and led to the creation of responsible disclosure programs. These programs provide guidelines
for reporting and cataloguing vulnerabilities. For instance, CISA KEV[
          <xref ref-type="bibr" rid="ref20">56</xref>
          ] is a catalogue of Known
Exploited Vulnerabilities that updates its entries only when a vulnerability has been assigned a common
identifier for publicly known cybersecurity vulnerability, actively exploited, and there is clear
remediation guidance. On the other hand, VulnCheck KEV[
          <xref ref-type="bibr" rid="ref21">57</xref>
          ] adds vulnerabilities to their catalogue as long
as the vulnerability is publicly reported as exploited in the wild. Unlike CISA KEV, VulnCheck KEV
does not have additional criteria such as an identifier and clear remediation guidance. It is still unclear
whether there might be a similar requirement for vulnerability disclosure in malware. Currently, the
Malvuln project is leading that charge by cataloguing vulnerabilities found in malware and providing
exploitation details.
        </p>
        <p>The dataset and benchmarks in this research focus on using exploitable vulnerabilities in malware for
ofensive security in defense technology, rather than public disclosure of malware vulnerabilities. While
it is known that attackers exploit vulnerabilities in benign software, we are looking at the opposite
scenario, where defenders can identify vulnerabilities in malware to enhance threat intelligence for
cyber defense. The dataset aims to advance research on using exploitable malware vulnerabilities as a
defense layer against malware.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Future Work</title>
        <p>The contributions to the POC vulnerabilities in malware within this dataset are supplied only by Malvuln.
This single-author database, although regularly updated, requires subsequent future work to aid in
obtaining samples to investigate the impact of the sample size on the performance metrics which is
necessary for the research of vulnerability detection in malware.</p>
        <p>
          One potential enhancement from Section 3.4 is to the CWE mapping of malicious binaries and the
incorporation of CWE chains and composite[
          <xref ref-type="bibr" rid="ref22">58</xref>
          ]. This involves identifying the relationships between
diferent CWEs in a vulnerable sample, which can be implicit, named or composite as defined by
MITRE’s knowledge base. By identifying these relationships, we can understand how weaknesses
can be combined to create vulnerabilities which can then inform the process of generating automated
exploits. For example, focusing on only one weakness in the chain or one composite component might
limit the comprehensive understanding of vulnerabilities. RedHat has adopted the practice of chaining
multiple CWEs together for the root cause analysis of security flaws in their products, aiming to go
beyond tracking a single root cause, which is the current approach to CWEs in the industry[
          <xref ref-type="bibr" rid="ref23 ref24">59, 60</xref>
          ].
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Conclusion</title>
        <p>This paper presents a new PE malware dataset that leverages the use of Malvuln to open new doors
in malware vulnerability research. Our dataset introduces information unavailable in the other
complementary PE malware datasets, such as the presence and annotation of vulnerabilities, obfuscation
techniques or ATT&amp;CK. Our contribution brings together the vulnerabilities found in malware from
the Malvuln dataset and vulnerabilities in benign software from the ExploitDB database, CWE mapping
in both vulnerable classes, Mitre ATT&amp;CK Tactics and Techniques from VulDB, and analysis of the use
of obfuscation across all classes. Additionally, our dataset utilizes Ember to extract the static PE features
from all compatible samples to form feature vectors across the four classes to be used for Machine
Learning. We obtained benchmark and baseline results on binary malware and benign classifiers,
vulnerability detection and multi-class (vulnerable/non-vulnerable, malware/benign) classifiers using
these features. From our results, we can derive various assumptions. For all tasks, it is clear that the
classifiers could efectively and accurately distinguish between the two cases; 1. malware and benign
software, and also 2. vulnerable and non-vulnerable samples but to a lesser degree. We also provide an
in-depth analysis of the threat and vulnerability mapping and the correlation of vulnerabilities between
malware and benign software. By providing these insights and a deep analysis of the importance of
exploitable malware vulnerabilities and the potential of their identification as a defensive mechanism,
we hope that our contributions encourage further studies into exploitability for defense.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We wish to acknowledge funding from the UK Government through the New Deal for Northern Ireland.
The funding is delivered on behalf of the Northern Ireland Ofice and the Department for Science,
Innovation and Technology by Innovate UK. Domhnall Carlin is funded by the UKRI EPSRC Research
Software Engineering Fellowship (EP/V052284/1).
[18] MalwareTech, Finding the kill switch to stop the spread of ransomware, 2017. URL: https://www.</p>
      <p>ncsc.gov.uk/blog-post/finding-kill-switch-stop-spread-ransomware-0, date Accessed: 07-06-2024.
[19] J. Fuller, R. P. Kasturi, A. Sikder, H. Xu, B. Arik, V. Verma, E. Asdar, B. Saltaformaggio, C3po:
large-scale study of covert monitoring of c&amp;c servers via over-permissioned protocol infiltration,
in: Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security,
2021, pp. 3352–3365. doi:10.1145/3460120.3484537.
[20] SecurityScorecard, When hackers get hacked: A cybersecurity triumph, 2023. URL: https://app.</p>
      <p>daily.dev/posts/HpFtuXxKC, date Accessed: 07-06-2024.
[21] A. Gdanski, L. Kessem, From thanos to prometheus: When ransomware encryption goes wrong,
2021. URL: https://securityintelligence.com/posts/ransomware-encryption-goes-wrong/, date
accessed: 07-06-2024.
[22] P. Arntz, Oops! black basta ransomware flubs encryption, 2024. URL: https://www.threatdown.</p>
      <p>com/blog/oops-black-basta-ransomware-flubs-encryption/, date accessed: 07-06-2024.
[23] A. Anubhav, Several iot botnet c2s compromised by a threat actor due to weak credentials., 2019.</p>
      <p>URL: https://www.ankitanubhav.info/post/c2bruting, date Accessed: 07-06-2024.
[24] S. Walla, C. Rossow, Malpity: Automatic identification and exploitation of tarpit vulnerabilities in
malware, in: 2019 IEEE European Symposium on Security and Privacy (EuroS&amp;P), IEEE, 2019, pp.
590–605. doi:10.1109/EuroSP.2019.00049.
[25] H. Grifioen, C. Doerr, Could you clean up the internet with a pit of tar? investigating tarpit
feasibility on internet worms, in: 2023 IEEE Symposium on Security and Privacy (SP), IEEE, 2023,
pp. 2551–2565. doi:10.1109/SP46215.2023.10179467.
[26] G. Apruzzese, P. Laskov, E. Montes de Oca, W. Mallouli, L. Brdalo Rapa, A. V. Grammatopoulos,
F. Di Franco, The role of machine learning in cybersecurity, Digital Threats: Research and Practice
4 (2023) 1–38. doi:10.1145/3545574.
[27] J. Singh, J. Singh, A survey on machine learning-based malware detection in executable files,
Journal of Systems Architecture 112 (2021) 101861. URL: https://www.sciencedirect.com/science/
article/pii/S1383762120301442. doi:10.1016/j.sysarc.2020.101861.
[28] D. Ucci, L. Aniello, R. Baldoni, Survey of machine learning techniques for malware analysis,</p>
      <p>Computers &amp; Security 81 (2019) 123–147. doi:10.1016/j.cose.2018.11.001.
[29] R. Sihwail, K. Omar, K. Z. Arifin, A survey on malware analysis techniques: Static, dynamic,
hybrid and memory analysis, Int. J. Adv. Sci. Eng. Inf. Technol 8 (2018) 1662–1671. doi:10.18517/
ijaseit.8.4-2.6827.
[30] H. Aghakhani, F. Gritti, F. Mecca, M. Lindorfer, S. Ortolani, D. Balzarotti, G. Vigna, C. Kruegel,
When malware is packin’heat; limits of machine learning classifiers based on static analysis
features, in: Network and Distributed Systems Security (NDSS) Symposium 2020, 2020.
[31] R. Sihwail, K. Omar, K. A. Zainol Arifin, S. Al Afghani, Malware Detection Approach Based
on Artifacts in Memory Image and Dynamic Analysis, Applied Sciences 9 (2019) 3680. URL:
https://www.mdpi.com/2076-3417/9/18/3680. doi:10.3390/app9183680, number: 18 Publisher:
Multidisciplinary Digital Publishing Institute.
[32] K. A. Roundy, B. P. Miller, Hybrid analysis and control of malware, in: Recent Advances in Intrusion
Detection: 13th International Symposium, RAID 2010, Ottawa, Ontario, Canada, September 15-17,
2010. Proceedings 13, Springer, 2010, pp. 317–338. doi:10.1007/978-3-642-15512-3_17.
[33] B. Cheng, J. Ming, J. Fu, G. Peng, T. Chen, X. Zhang, J.-Y. Marion, Towards paving the way
for large-scale windows malware analysis: Generic binary unpacking with orders-of-magnitude
performance boost, in: Proceedings of the 2018 ACM SIGSAC Conference on Computer and
Communications Security, 2018, pp. 395–411. doi:10.1145/3243734.3243771.
[34] O. Alrawi, M. Ike, M. Pruett, R. P. Kasturi, S. Barua, T. Hirani, B. Hill, B. Saltaformaggio, Forecasting
malware capabilities from cyber attack memory images, in: 30th USENIX security symposium
(USENIX security 21), 2021, pp. 3523–3540.
[35] D. G. Corlatescu, A. Dinu, M. P. Gaman, P. Sumedrea, Embersim: A large-scale databank for
boosting similarity search in malware analysis, Advances in Neural Information Processing
Systems 36 (2024).
[36] M. Sebastián, R. Rivera, P. Kotzias, J. Caballero, Avclass: A tool for massive malware labeling,
in: Research in Attacks, Intrusions, and Defenses: 19th International Symposium, RAID 2016,
Paris, France, September 19-21, 2016, Proceedings 19, Springer, 2016, pp. 230–253. doi:10.1007/
978-3-319-45719-2_11.
[37] S. Sebastián, J. Caballero, Avclass2: Massive malware tag extraction from av labels, in: Proceedings
of the 36th Annual Computer Security Applications Conference, 2020, pp. 42–53. doi:10.1145/
3427228.3427261.
[38] R. Lyda, J. Hamrock, Using entropy analysis to find encrypted and packed malware, IEEE Security
&amp; Privacy 5 (2007) 40–45. doi:10.1109/MSP.2007.48.
[39] X. Ugarte-Pedrero, D. Balzarotti, I. Santos, P. G. Bringas, Sok: Deep packer inspection: A
longitudinal study of the complexity of run-time packers, in: 2015 IEEE Symposium on Security and
Privacy, IEEE, 2015, pp. 659–673. doi:10.1109/SP.2015.46.
[40] T. Muralidharan, A. Cohen, N. Gerson, N. Nissim, File packing from the malware perspective:
techniques, analysis approaches, and directions for enhancements, ACM Computing Surveys 55
(2022) 1–45. doi:10.1145/3530810.
[41] T. Avgerinos, S. K. Cha, A. Rebert, E. J. Schwartz, M. Woo, D. Brumley, Automatic exploit generation,</p>
      <p>Communications of the ACM 57 (2014) 74–84. doi:10.1145/2560217.2560219.
[42] S. Xu, Y. Wang, Bofaeg: Automated stack bufer overflow vulnerability detection and exploit
generation based on symbolic execution and dynamic analysis, Security and Communication
Networks 2022 (2022) 1251987. doi:10.1155/2022/1251987.
[43] V. A. Padaryan, V. Kaushan, A. Fedotov, Automated exploit generation for stack bufer
overlfow vulnerabilities, Programming and Computer Software 41 (2015) 373–380. doi: 10.1134/
S0361768815060055.
[44] S. Park, D. Kim, S. Jana, S. Son, {FUGIO}: Automatic exploit generation for {PHP} object injection
vulnerabilities, in: 31st USENIX Security Symposium (USENIX Security 22), 2022, pp. 197–214.
[45] S. Jan, A. Panichella, A. Arcuri, L. Briand, Automatic generation of tests to exploit xml injection
vulnerabilities in web applications, IEEE Transactions on Software Engineering 45 (2017) 335–362.
doi:10.1109/TSE.2017.2778711.
[46] T. Wang, T. Wei, G. Gu, W. Zou, Taintscope: A checksum-aware directed fuzzing tool for automatic
software vulnerability detection, in: 2010 IEEE Symposium on Security and Privacy, IEEE, 2010,
pp. 497–512. doi:10.1109/SP.2010.37.
[47] K. Oosthoek, C. Doerr, Sok: Att&amp;ck techniques and trends in windows malware, in: Security
and Privacy in Communication Networks: 15th EAI International Conference, SecureComm
2019, Orlando, FL, USA, October 23-25, 2019, Proceedings, Part I 15, Springer, 2019, pp. 406–425.
doi:10.1007/978-3-030-37228-6_20.
[48] Picus Security, Picus red report 2024: The top 10 most prevalent mitre att&amp;ck techniques
the rise of hunter-killer malware, 2024. URL: https://www.picussecurity.com/resource/report/
picus-red-report-2024.
[49] Centre for Threat Informed Defense, MITRE ENGENUITY., Top att&amp;ck techniques, 2023. URL:
https://top-attack-techniques.mitre-engenuity.org/, date accessed: 11-06-2024.
[50] P. Xu, Y. Zhang, C. Eckert, A. Zarras, Hawkeye: cross-platform malware detection with
representation learning on graphs, in: Artificial Neural Networks and Machine Learning–ICANN 2021: 30th
International Conference on Artificial Neural Networks, Bratislava, Slovakia, September 14–17,
2021, Proceedings, Part III 30, Springer, 2021, pp. 127–138. doi:10.1007/978-3-030-86365-4_
11.
[51] A. T. Nguyen, F. Lu, G. L. Munoz, E. Raf, C. Nicholas, J. Holt, Out of distribution data detection
using dropout bayesian neural networks, in: Proceedings of the AAAI Conference on Artificial
Intelligence, volume 36, 2022, pp. 7877–7885. doi:10.1609/aaai.v36i7.20757.
[52] L. Van der Maaten, G. Hinton, Visualizing data using t-sne., Journal of machine learning research
9 (2008).
[53] S. Chakraborty, R. Krishna, Y. Ding, B. Ray, Deep learning based vulnerability detection: Are we
there yet?, IEEE Transactions on Software Engineering 48 (2021) 3280–3296. doi:10.1109/TSE.</p>
    </sec>
    <sec id="sec-7">
      <title>A. Appendix</title>
      <sec id="sec-7-1">
        <title>A.1. Timestamps</title>
        <p>
          It’s important to consider the timeline of a dataset in machine-learning-based malware detection and
classification systems. This helps in capturing temporal dependencies within the data, thus avoiding
experimental bias or incorrect conclusions [
          <xref ref-type="bibr" rid="ref25 ref26">61, 62</xref>
          ]. Benchmark datasets such as BODMAS, EMBER, and
SOREL-20M release their datasets with timestamped samples. For instance, BODMAS uses the first-seen
time of a sample based on VirusTotal reports, EMBER uses the timestamp in the header from the COFF
header to split its dataset, and SOREL-20M curates the first and last-seen times of the samples, eventually
using the first-seen time for temporal splitting of the data. Similar to BODMAS and SOREL-20M, our
dataset is annotated using the first seen time from VirusTotal reports, as it has been demonstrated to be
robust to alteration, reliable, and accurate compared to other timestamp metrics[
          <xref ref-type="bibr" rid="ref27">63</xref>
          ]. The timeseries
graphs in Figures 4, 5, 6, and 7 are scaled between the years 2000 and 2024.
        </p>
        <p>Time Series of Values for Vulnerable Malware</p>
        <p>Classes
Vulnerable Malware</p>
        <p>Classes
Vulnerable Benign
120
100
80
cyeun
q
rFe 60
40
20
70
60
50
cy40
n
e
u
rFeq30
20
10
20000
cyenu
q
rFe15000</p>
      </sec>
      <sec id="sec-7-2">
        <title>A.2. Binary Classification</title>
      </sec>
      <sec id="sec-7-3">
        <title>A.3. Feature Importance</title>
        <p>
          Feature importance is a concept in machine learning that quantifies the contribution of each feature
to the prediction of the model so we can derive explanations for what features influence the model’s
determinate output. Various studies of use methods such as correlation analysis [
          <xref ref-type="bibr" rid="ref28">64</xref>
          ] and Shapley values
[
          <xref ref-type="bibr" rid="ref29 ref30 ref31">65, 66, 67</xref>
          ] to assess feature importance. Correlation measures the strength and direction of a linear
relationship on a scale from -1 to 1 between features and the target variable, ofering a straightforward
but overly simplistic view. Whereas Shapley values, which are built from cooperative game theory
as a framework explaining the output of a model in correlation to its input features provide a more
nuanced and comprehensive evaluation by considering all possible combinations of features and their
interactions, which are only bounded by the output magnitude range of the model. This ensures that
there is a fair distribution of importance among features which can not only identify key predictive
features but also aid in the interpretation of the model and its results. We focused on using Shapley
when working with feature importance for our classes, as it also has many methods for plotting graphs
in its library to better visualize the data. Typically in Shapley graphs, it introduces a gradient scale from
red to blue which corresponds to the raw values of the variables for each instance.
• Red points indicate high-impact SHAP values, suggesting that the feature value led the model to
increase its prediction.
• Blue points indicate low-impact SHAP values, suggesting that the feature value led the model to
decrease its prediction.
It is worth noting that SHAP gives insights into the model’s behavior in the context of the data used,
but causal relationships or generalizable patterns with unseen data could be overlooked. SHAP could
also generate unexpected relationships between features and predictions which indicate possible data
issues or model artefacts.
        </p>
        <p>From the SHAP results in Figure 8, which are derived from the LGBM model performing multi-class
classification on the data, many observations can be made. The first of which we can see is that feature
637 had a high positive impact on both vulnerable and non-vulnerable malware predictions. Whereas
feature 637 had a significantly low positive impact in its prediction for non-vulnerable benign with
more weight being in the negative direction. After investigating this feature, we found that it exists
in the COFF File Header, specifically the machine type 30. There were 2 machine types identified; I386
and AMD64. The I386 machines consistently dominate across all classes compared to the AMD64, with
557/3 in VM, 1,408/4 in VB, 32,615/213 in NVM, and 2,910/2222 in NVB. From this distribution, we can
see that, while I386 machines are predominantly more common in VM, VB, and NVM classes, the NVB
30https://learn.microsoft.com/en-us/windows/win32/debug/pe-format
class has a relatively higher proportion of AMD64 machines, which explains the data distribution in
Figure 2. A broader comparison was made in Table 20 by demonstrating the distribution of features
across the top 15 SHAP values for each class.
SHAP value (impact on model output)</p>
        <p>SHAP value (impact on model output)
0
10</p>
        <p>Low</p>
        <p>Feature 626
Feature 757
Feature 520</p>
        <p>Feature 679
e Feature 505
lu Feature 2355
va Feature 509
re Feature 1627
u Feature 503
ta Feature 2359
F Feature 613
e</p>
        <p>Feature 132</p>
        <p>Feature 196
Feature 1561</p>
        <p>Feature 95
Feature 637
Feature 626
Feature 691</p>
        <p>Feature 677
e Feature 2371
lu Feature 613
va Feature 89
e Feature 683
ru Feature 204
ta Feature 2373
F Feature 784
e</p>
        <p>Feature 502
Feature 986
Feature 132
Feature 658
e
u
l
a
v
e
r
u
t
a
e
F
e
u
l
a
v
e
r
u
t
a
e
F
0
2</p>
        <p>Low
SHAP value (impact on model output)</p>
        <p>SHAP value (impact on model output)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Elhadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maarof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hamza</surname>
          </string-name>
          <string-name>
            <surname>Osman</surname>
          </string-name>
          ,
          <article-title>Malware detection based on hybrid signature behaviour application programming interface call graph</article-title>
          ,
          <source>American Journal of Applied Sciences</source>
          <volume>9</volume>
          (
          <year>2012</year>
          )
          <fpage>283</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          , D.-G. Feng,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.-R.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <article-title>Semantics-based malware behavior signature extraction and detection method</article-title>
          ,
          <source>Journal of Software</source>
          <volume>23</volume>
          (
          <year>2012</year>
          )
          <fpage>378</fpage>
          -
          <lpage>393</lpage>
          . doi:
          <volume>10</volume>
          .3724/SP.J.
          <volume>1001</volume>
          .
          <year>2012</year>
          .
          <volume>03953</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jalilian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Narimani</surname>
          </string-name>
          , E. Ansari,
          <article-title>Static signature-based malware detection using opcode and binary information</article-title>
          , in: Data Science: From Research to Application, Springer,
          <year>2020</year>
          , pp.
          <fpage>24</fpage>
          -
          <lpage>35</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -37309-
          <issue>2</issue>
          _
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ciptadi</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Laziuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ahmadzadeh</surname>
          </string-name>
          , G. Wang,
          <string-name>
            <surname>Bodmas:</surname>
          </string-name>
          <article-title>An open dataset for learning based temporal analysis of pe malware</article-title>
          ,
          <source>in: 2021 IEEE Security and Privacy Workshops (SPW)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>78</fpage>
          -
          <lpage>84</lpage>
          . doi:
          <volume>10</volume>
          .1109/SPW53761.
          <year>2021</year>
          .
          <volume>00020</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H. S.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <article-title>EMBER: an open dataset for training static PE malware machine learning models</article-title>
          , CoRR abs/
          <year>1804</year>
          .04637 (
          <year>2018</year>
          ). URL: http://arxiv.org/abs/
          <year>1804</year>
          .04637. arXiv:
          <year>1804</year>
          .04637.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Harang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Rudd</surname>
          </string-name>
          , SOREL-20M:
          <article-title>A large scale benchmark dataset for malicious PE detection</article-title>
          , CoRR abs/
          <year>2012</year>
          .07634 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2012</year>
          .07634. arXiv:
          <year>2012</year>
          .07634.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ronen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Radu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Feuerstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yom-Tov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ahmadi</surname>
          </string-name>
          , Microsoft Malware Classification Challenge,
          <year>2018</year>
          . URL: http://arxiv.org/abs/
          <year>1802</year>
          .10135, arXiv:
          <year>1802</year>
          .10135 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Henriquez</surname>
          </string-name>
          ,
          <article-title>Bugs in malware creating backdoors for security researchers</article-title>
          ,
          <year>2021</year>
          . URL: https://www.securitymagazine.com/articles/ 96348-bugs
          <article-title>-in-malware-creating-backdoors-for-security-researchers</article-title>
          , date accessed:
          <fpage>07</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Caballero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Poosankam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>McCamant</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Babi ć, D. Song, Input generation via decomposition and re-stitching: Finding bugs in malware</article-title>
          ,
          <source>in: Proceedings of the 17th ACM conference on Computer and communications security</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>413</fpage>
          -
          <lpage>425</lpage>
          . doi:
          <volume>10</volume>
          .1145/1866307.1866354.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Pratap Singh</surname>
          </string-name>
          ,
          <article-title>VB2021 paper: Bugs in malware - uncovering vulnerabilities found in malware payloads</article-title>
          , Virus
          <string-name>
            <surname>Bulletin</surname>
          </string-name>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          . URL: https://vblocalhost.com/uploads/ VB2021-Singh-Singh.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Anubhav</surname>
          </string-name>
          , Crash and Burn ::
          <article-title>How to crash a Mirai C2 server &amp; why it works</article-title>
          .,
          <year>2019</year>
          . URL: https://www.ankitanubhav.info/post/crash.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Abrams</surname>
          </string-name>
          ,
          <article-title>Vulnerabilities allow hijacking of most ransomware to prevent ifle encryption</article-title>
          ,
          <year>2022</year>
          . URL: https://www.bleepingcomputer.com/news/security/ lockbit-30
          <article-title>-introduces-the-first-ransomware-bug-bounty-program/</article-title>
          , date accessed:
          <fpage>15</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Page</surname>
          </string-name>
          ,
          <article-title>Ransomlord anti-ransomware exploit tool</article-title>
          .,
          <year>2024</year>
          . URL: https://github.com/malvuln/ RansomLord, date accessed:
          <fpage>15</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>I. Ilascu</surname>
          </string-name>
          , Conti, revil, lockbit ransomware bugs exploited to block encryption,
          <year>2022</year>
          . URL: https://web.archive.org/web/20220601204439/https://www.bleepingcomputer.com/news/ security/conti-revil
          <article-title>-lockbit-ransomware-bugs-exploited-to-block-encryption/</article-title>
          , date accessed:
          <fpage>15</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kovacs</surname>
          </string-name>
          ,
          <article-title>Vulnerabilities allow hijacking of most ransomware to prevent file encryption</article-title>
          ,
          <year>2022</year>
          . URL: https://web.archive.org/web/20220504180432/https://www.securityweek.
          <article-title>com/ vulnerabilities-allow-hijacking-most-ransomware-prevent-file-encryption/</article-title>
          , date accessed:
          <fpage>15</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Calleja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tapiador</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Caballero,</surname>
          </string-name>
          <article-title>The malsource dataset: Quantifying complexity and code reuse in malware development</article-title>
          ,
          <source>IEEE Transactions on Information Forensics and Security</source>
          <volume>14</volume>
          (
          <year>2018</year>
          )
          <fpage>3175</fpage>
          -
          <lpage>3190</lpage>
          . doi:
          <volume>10</volume>
          .1109/TIFS.
          <year>2018</year>
          .
          <volume>2885512</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Rosenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Beek</surname>
          </string-name>
          ,
          <article-title>Examining code reuse reveals undiscovered links among north korea's malware families</article-title>
          ,
          <year>2018</year>
          . URL: https://www.mcafee.com/blogs/other-blogs/
          <article-title>mcafee-labs/ examining-code-reuse-reveals-undiscovered-links-among-north-koreas-malware-families/</article-title>
          , date Accessed:
          <fpage>09</fpage>
          -
          <lpage>07</lpage>
          -
          <year>2024</year>
          .
          <year>2021</year>
          .
          <volume>3087402</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bilot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. El</given-names>
            <surname>Madhoun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Al Agha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zouaoui</surname>
          </string-name>
          ,
          <article-title>A survey on malware detection with graph representation learning</article-title>
          ,
          <source>ACM Computing Surveys</source>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .1145/3664649.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hasegawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yamaguchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shimada</surname>
          </string-name>
          ,
          <article-title>Malware detection by control-flow graph level representation learning with graph isomorphism network</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>111830</fpage>
          -
          <lpage>111841</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2022</year>
          .
          <volume>3215267</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [56]
          <article-title>CISA.gov, Reducing the significant risk of known exploited vulnerabilities</article-title>
          ,
          <year>2022</year>
          . URL: https: //www.cisa.gov/known-exploited-vulnerabilities, date accessed:
          <fpage>15</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [57]
          <string-name>
            <surname>Vulncheck</surname>
            <given-names>Kev</given-names>
          </string-name>
          , Coverage criteria,
          <year>2024</year>
          . URL: https://docs.vulncheck.com/community/ vulncheck-kev/coverage-criteria, date accessed:
          <fpage>15</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [58]
          <string-name>
            <surname>MITRE</surname>
          </string-name>
          , Chains and composites,
          <year>2023</year>
          . URL: https://cwe.mitre.org/data/reports/chains_and_ composites.html, date Accessed:
          <fpage>16</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [59]
          <article-title>Red Hat Product Security, cwe-toolkit - cwe chaining concept</article-title>
          and tools,
          <year>2020</year>
          . URL: https://github. com/RedHatProductSecurity/cwe-toolkit, date Accessed:
          <fpage>16</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [60]
          <string-name>
            <given-names>Red</given-names>
            <surname>Hat Customer Portal</surname>
          </string-name>
          ,
          <article-title>Cwe compatibility for red hat customer portal</article-title>
          ,
          <year>2024</year>
          . URL: https://access. redhat.com/articles/cwe_compatibility, date Accessed:
          <fpage>16</fpage>
          -
          <lpage>06</lpage>
          -
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [61]
          <string-name>
            <given-names>K.</given-names>
            <surname>Allix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Bissyandé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Klein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Le Traon</surname>
          </string-name>
          ,
          <article-title>Are your training datasets yet relevant? an investigation into the importance of timeline in machine learning-based malware detection</article-title>
          ,
          <source>in: International Symposium on Engineering Secure Software and Systems</source>
          , Springer,
          <year>2015</year>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>67</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -15618-
          <issue>7</issue>
          _
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [62]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pendlebury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pierazzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jordaney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kinder</surname>
          </string-name>
          , L. Cavallaro, {TESSERACT}:
          <article-title>Eliminating experimental bias in malware classification across space and time</article-title>
          ,
          <source>in: 28th USENIX security symposium (USENIX Security 19)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>729</fpage>
          -
          <lpage>746</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [63]
          <string-name>
            <given-names>A.</given-names>
            <surname>Guerra-Manzanares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Luckner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bahsi</surname>
          </string-name>
          ,
          <article-title>Concept drift and cross-device behavior: Challenges and implications for efective android malware detection</article-title>
          ,
          <source>Computers &amp; Security</source>
          <volume>120</volume>
          (
          <year>2022</year>
          )
          <article-title>102757</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.cose.
          <year>2022</year>
          .
          <volume>102757</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [64]
          <string-name>
            <given-names>G.</given-names>
            <surname>Chin</surname>
          </string-name>
          ,
          <article-title>Correlation vs SHAP: Understanding Feature Importance in ML Models</article-title>
          ,
          <year>2024</year>
          . URL: https://medium.com/@
          <article-title>gawainchin/ correlation-vs-shap-understanding-feature-importance-in-ml-models-d6b52b1fba28.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [65]
          <string-name>
            <given-names>K.</given-names>
            <surname>Främling</surname>
          </string-name>
          ,
          <article-title>Feature Importance versus Feature Influence and What It Signifies for Explainable AI</article-title>
          ,
          <year>2023</year>
          . URL: http://arxiv.org/abs/2308.03589, arXiv:
          <fpage>2308</fpage>
          .03589 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [66]
          <string-name>
            <given-names>R.</given-names>
            <surname>Boemer</surname>
          </string-name>
          ,
          <source>Selecting Features With Shapley Values</source>
          ,
          <year>2023</year>
          . URL: https://medium.com/
          <article-title>the-ml-practitioner/selecting-features-with-shapley-values-b2da08b5b14c.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [67]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kübler</surname>
          </string-name>
          , Shapley Values Clearly Explained,
          <year>2024</year>
          . URL: https://towardsdatascience.com
          <article-title>/ shapley-values-clearly-explained-a7f7ef22b104.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>