<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Binary Malware Attribution using LLM Embeddings and Topological Data Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kincaid MacDonald</string-name>
          <email>Kincaid.MacDonald@gtri.gatech.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ajai Ruparelia</string-name>
          <email>Ajai.Ruparelia@gtri.gatech.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Boomer Rogers</string-name>
          <email>Herman.Rogers@gatech.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adonis Bovell</string-name>
          <email>Adonis.Bovell@gtri.gatech.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Branden Stone</string-name>
          <email>Branden.Stone@gtri.gatech.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Georgia Institute of Technology</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Georgia Tech Research Institute</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In malware authorship attribution, easily obtained, expressive, stylistically salient features are scarce. Often the only data source is the binary malware executable; this has traditionally constrained static analysis to producing simple statistics on the decompiled binary. Dynamic analysis can obtain more descriptive features by running the malware executable in a sandbox, but this is both resource-intensive and has obvious risks. In this work, we introduce two new sources of expressive static features which, when combined, achieve classification accuracies competitive with leading dynamic analysis models, and set a new state of the art for static malware authorship attribution. Our features stem from a hypothesis that distinctive stylistic coding features are present at or around the level of the function. To capture features at the function level, we use Large Language Model (LLM) embeddings of decompiled source code. To capture features at the function level, we use a Graph Neural Network, trained end-to-end on call graphs to identify stylistically-salient geometric features of the program architecture. We demonstrate that this new paradigm of graph learning combined with LLM code-embeddings can easily be extended to encompass and enhance existing featurizations. We further employ Topological Data Analysis to illustrate that the topological features of the call graph reveal stylistic indicators, which significantly aid in the task of code attribution.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Code Attribution</kwd>
        <kwd>Large Language Model</kwd>
        <kwd>Topological Data Analysis</kwd>
        <kwd>Graph Neural Networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        What is the shape of code – and can it reveal a telltale fingerprint of the coder? Stylometry has a
rich history within source code, where the programmer’s identity can be revealed by quasi-linguistic
features like variable names, spacing conventions, and syntax [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. But in many areas, as with
malware, source code is unavailable, and stylometry must be performed directly on binary executables.
Here, the usual arsenal of linguistic stylometers falls flat. Traditional techniques are entirely dependent
on the quality of the binary decompilation and are efectively left hunting for scraps - like file headers,
import tables, and careless comments not removed during compilation [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>
        These challenges for traditional static analysis of binary malware samples motivate today’s leading
commercial solution: dynamic malware analysis. Here, the malware is run within a sandbox, which
converts its behavior - API calls, network requests - into features used in downstream classification
[
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ]. Unfortunately, malware files are sometimes able to detect when they’re being run in a sandbox
and mask their behavior accordingly. Additionally, running hostile code on one’s computers, even
within a sandbox, is inherently risky.
      </p>
      <p>In this work, we propose a new type of static malware analysis based on the topology and geometry
of the underlying program – i.e. the shape of the code. We show that information about the shape
alone, extracted from the program’s call graph, can achieve reasonable performance on binary malware
classification. When we augment the call
graph with code embeddings that indicate the
relative proximity in feature space of nodes on
the graph, our model achieves a new
state-ofthe-art for static binary malware analysis. In
short, our novel contributions are as follows:
• demonstrate how Topological Data</p>
      <p>Analysis (TDA) on call graphs can
extract distinguishing structures and
signatures of code;
• present a novel method based on Large</p>
      <p>Language Model (LLM) embeddings that
extract stylistic features from
decompiled code;
• introduce a pipeline integrating LLM
and TDA features into Graph Neural
Network (GNN) architecture to produce
a state-of-the-art static malware
attribution model.</p>
      <p>This paper is structured as follows: First,
we describe our ‘function-level’ hypothesis
that motivates our architecture design (see
Figure 1). We then profile our dataset and its
processing pipeline, before detailing our novel
stylistic features: LLM code embeddings, and
persistent homology across the call graph. We
ifnally describe our models, their training, and
their performance.
2. The</p>
    </sec>
    <sec id="sec-2">
      <title>Function-level Hypothesis</title>
      <p>In any program, there are two primary stylistic
influences: the preferences of the programmer
and the demands of the program. The
programmer may favor specific constructs, such as
using for loops instead of while loops, switches
instead of if statements, or dictionaries instead
of objects. Moreover, programmers often have
unique ways of defining functions; one may use several small functions, while another might achieve
the same goal with a single, larger function.</p>
      <p>
        Additionally, the program imposes global design constraints related to its structure and functionality.
For example, ransomware coerces victims into involuntary actions, such as demanding money through
extortion, using techniques like encryption and blockers. Unlike spyware, ransomware doesn’t need to
intercept/modify OS functions or APIs to be efective [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In the context of malware, diferent Advanced
Persistent Threat (APT) groups often specialize in diferent types of malicious attacks (e.g. ransomware
or espionage of the electric utility sector [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]), each imposing distinct strategies and constraints on
the malware architecture. Conversely, the lower-level features, such as individual lines of code, are
influenced by the programming language’s syntax and style guidelines.
      </p>
      <p>The question then arises: where can we observe a programmer’s style most efectively? We
hypothesize that a programmer’s style is most evident at the function level. Functions strike the balance
between global design constraints and local syntactic constraints, making them ideal candidates for
this analysis. Moreover, good programming practice emphasizes making functions simple and modular,
which allows complex problems to be decomposed into smaller, manageable parts. This decomposition
process enables an in-depth and precise analysis of the programmer’s unique approach and style.</p>
      <p>In this paper, we aim to identify a programmer’s style by focusing on the function level, rather
than the global structure dictated by program goals or the local syntax influenced by language and
style guides. We hypothesize that the function level, encompassing both the contents of functions and
their interconnectivity, provides the ideal scale for this analysis. By examining functions and their
interactions, we demonstrate the feasibility of identifying a unique “fingerprint” or “signature” of a
team or developer in the context of malware attribution.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset Overview &amp; Feature Extraction</title>
      <p>
        To ensure the reproducibility and comparability of our approach with common benchmarks, we utilize
a publicly accessible dataset [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] consisting of over 3,500 malware samples attributed to 12 Advanced
Persistent Threat (APT) Groups.. We exclude 721 samples consisting of container files and 20 samples
which our automation could not open. Incidentally, this increases consistency, while reducing imbalance
across our dataset resulting in a total of 2,853 samples across 12 distinct APT Groups (see Table 1).
      </p>
      <sec id="sec-3-1">
        <title>3.1. Feature Extraction</title>
        <p>
          To process our dataset, we leverage Ghidra [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], a software reverse engineering (SRE) framework
created and maintained by the National Security Agency, and develop a comprehensive analysis pipeline
that automates the extraction and processing of key features from each malware sample. Our pipeline
is structured to provide insights at both the structural program level and the granular function level,
facilitating the generation of novel features that capture the stylistic nuances of diferent APT groups.
The main steps in our analysis pipeline (Figure 2) are as follows:
        </p>
        <sec id="sec-3-1-1">
          <title>3.1.1. Function-Level Feature Extraction</title>
          <p>For each malware sample, we begin by extracting function-specific features, such as:
• Number of Input Parameters: The count of input parameters for each function.
• Function Size: The size of the function measured by the number of instructions or lines of code.
• Return Type: The data type of the value returned by the function.</p>
          <p>• Number of Variables: The count of variables for each function</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.1.2. Call Graph Construction</title>
          <p>Next, we construct a call graph for each sample:
• Node Representation: Each node in the graph represents a function. The node attributes include
the function-specific features extracted in the previous step.
• Edge Representation: Directed edges between nodes represent function calls, illustrating the
relationships and dependencies between functions within the sample.</p>
          <p>The call graph not only serves as the basis for our GNN structure but also allows us to build the
TDA features described in Section 4.</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>3.1.3. Function Decompilation and Embedding Generation</title>
          <p>We decompile each function to obtain its raw C code:
• Decompilation: Using Ghidra’s decompiler, we convert binary instructions back into its C
representation.
• String Representation: The decompiled C code is stored as a string, allowing for further
text-based analysis.
• Embedding Creation: These string representations are used to create embeddings via LLMs,
enabling the analyses and comparisons of code stylistic features. See Section 4.1.</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>3.1.4. Dataset Splitting</title>
          <p>To ensure robust evaluation, we divide the processed dataset into training, validation, and test sets:
• Split Ratio: We use an 80/10/10% split for training, validation, and testing datasets respectively.</p>
          <p>This approach ensures that our models are trained on a diverse set of samples while being
evaluated on unseen data to assess generalization performance.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Output</title>
        <p>The output of our pipeline consists of detailed representations of each malware sample, including
function-level features, call graphs, and decompiled function strings. These outputs serve as the
foundation for generating novel features that capture the unique characteristics of diferent APT groups.
In the following section, we describe how we leverage these outputs to create LLM embeddings and
TDA features, which are central to our methodology for distinguishing between the stylistic signatures
of malware authors.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Novel Stylistic Features</title>
      <sec id="sec-4-1">
        <title>4.1. LLM Embeddings of Decompiled Code</title>
        <p>To use our decompiled source code in downstream tasks, we first need to vectorize it. While code
embeddings do not contain suficient information to recover the content of the code, they are useful in
defining a measure of similarity between diferent embedded samples. In particular, we assume two
functions are similar if their code embeddings are near each other.</p>
        <p>
          Code embeddings are traditionally produced with an of-the-shelf embedding model trained
specifically on code, like CodeBERT [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. But the large-language model revolution has also introduced a new
state-of-the-art of text and code embedding. OpenAI’s text-embedding-ada-002 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] is such a model.
Built on the pre-trained core of GPT 3.5 and fine-tuned with an embedding objective, it can embed both
text and code into 1536-dimensional vectors in which Euclidean distance corresponds to conceptual
similarity.
        </p>
        <p>
          There are presently few open-source embedding models based on LLMs. Starencoder [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], based on
Starcoder [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], a GPT-2 based model trained exclusively on code, is the only one of which we’re aware,
and neither Starcoder nor its embedding model claim to match OpenAI’s oferings in performance. In
internal testing on an authorship attribution task using the Google CodeJam dataset [15], OpenAI’s
text-embedding-ada-002 model achieved the best performance. Thus, that model was selected for use
here. However, given the number of increasingly capable open-source LLMs (e.g. Meta’s recently
announced LLAMA 2 [16]) it’s only a matter of time before some are fine-tuned for embedding which
ofer performance competitive with today’s closed-source models.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Topological Data Analysis Features</title>
        <p>Our initial research question was initially conceived around the question: Can the ‘shape’ of code be used
for stylometry? TDA gives a powerful set of tools for measuring this shape, by describing the number
and size of connected components, any loops in the data, and - if in a suficiently complex program
- perhaps even the existence of higher-dimensional features (see [17] for a introduction to the topic).
Such topological landmarks are an example of features found around the function level. The degree
of branching between functions and sub-functions or the length of loops formed between diferent
pathways through the code all seem likely to reveal the style of its author.</p>
        <p>To TDA, there are two choices to be made. First, we must turn the code into a graph. Next, we
must select the appropriate vectorization of persistent homology on the graph. Motivated by our
function-level hypothesis, we chose the call graph as the representation most likely to reveal
highlevel stylistic features. For our vectorization, we utilize a combination of two leading techniques: we
used Carrière’s Heat Kernel Signatures to create several persistence diagrams with diferent difusion
parameters [18], then vectorized them into Bubenik’s Persistence Landscapes [19]. We concatenate and
lfatten the resulting representations into a single 4000-dimensional vector per malware sample. We
refer to these as Global TDA features (see Figure 3).</p>
        <p>Traditional TDA operates globally, describing the shape of the entire call graph. If more localized
topological descriptors are required, one can instead perform TDA on collections of subgraphs, creating
node features which encode the local neighborhood topology. This is the approach of Local TDA. Our
function-level hypothesis suggests that such local descriptors may better encode stylistic information,
so in addition to global TDA features, we compute local TDA features on each call graph. We form
-hop subgraphs around each node, and apply the same combination of Heat Kernel Signatures and
Persistence Landscapes to obtain 4000-dimensional local topology features for each node. For example,
in Figure 3, the local TDA of node  uses a 1-hop subgraph containing nodes , , and .</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Methods &amp; Results</title>
      <p>
        As a baseline comparison we choose a random forest model using binary features processed by fuzzy
hashing and NATO1 encoding to achieve 89% accuracy on the same dataset [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. As far as the authors
know, this is the current state of the art on the APT dataset. The hashing algorithm used for the best
accuracy is called impfuzzy [20]. This fuzzy hashing tool only considers the import API and not the
body of the binary. Our approach considers significantly diferent aspects of the binary and is outlined
in Figure 1.
      </p>
      <sec id="sec-5-1">
        <title>5.1. Feedforward Classifier</title>
        <p>We first tested the usefulness of both LLM and TDA features with a simple feedforward classifier. This
provides a baseline for how useful each of our features is in predicted authorship, and whether the LLM
embeddings and TDA features complement each other – whether they capture diferent descriptions of
an author’s style, and whether they can be combined to boost classification accuracy.</p>
        <p>To this end, we designed a fully-connected deep neural network consisting of three submodules.
One module takes a summary of the LLM embeddings as input (starting in 1536-dimensional space
as mentioned above), passes it through ten linear layers with Leaky ReLU activation’s, and outputs
a dimensionally-reduced embedding of the LLM features. The second model does the same for the
lfattened TDA features. The final module concatenates the resulting embeddings as input to a four-layer
classifier. We train the network with Cross Entropy loss and the Adam optimizer.</p>
        <p>One subtlety here is that this is a file-level classification problem, while our LLM embeddings are
function level features. Diferent malware samples have diferent numbers of functions, hence we need
some manner of summarizing the embeddings of  functions into a single vector of unvarying size.
Here, we copy a technique from the graph neural network literature [21]. Rather than performing
a single pooling operation (e.g. mean pooling), we pool the features according to four statistical
moments: obtaining the mean, median, skew, and kurtosis across each of the 1536 dimensions of the
LLM embeddings. We concatenate the results and use this as input to our LLM-embeddings module.
1https://www.nato.int/cps/en/natohq/news_150391.htm</p>
        <p>
          When trained with both LLM and global TDA features, this simple classifier achieves an accuracy of
88% and is comparable to the current state of the art random forest model using fuzzy hashing (see
Table 2). This diference could be attributed to the number of samples dropped from the dataset. In the
case of [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], as impfuzzy only considers import API, it failed to hash 729 samples due to the lack of
portable executable (PE) headers.
        </p>
        <p>We further analyzed the relative importance of our features by retraining the classifier while omitting
either the LLM features and its accompanying module, or the TDA features and its accompanying
module (see Table 2). The global TDA features, trained alone, achieved an accuracy of 73%. The LLM
embeddings, trained alone, achieved a higher accuracy of 85%. This could be due to the fact that LLM
embeddings are coming from the decompiled source code and have more information about the program.
However, the topological descriptors of the call graph performing at 73% seems to indicate that the
shape of code is a useful indicator of the author’s style.</p>
        <p>Analyzing the predictions made by each independently trained classifier, found that of the malware
samples incorrectly classified by the LLM-Embedding-trained model, the global-TDA-trained model
classified 22% correctly. That is, most of the time, when the models disagree, the global-TDA-trained
model is wrong. But a fourth of the time, it knows something the LLM-Embedding-trained doesn’t.
Reasoning from this, one would expect that in the best case, combining the global TDA and Embedding
features would augment the LLM-Embedding-trained model’s accuracy by about 5%. Indeed, this
(Table 2) is what we find as the model trained on both features simultaneously achieves near
state-ofthe-art accuracy (around 88%), versus 84% from the LLM embeddings alone.</p>
        <p>Admittedly, the four-moment summarization of our LLM embedding features is crude. It can highlight
global variance in each of the embedding dimensions – perhaps detecting, e.g., when an author has many
functions that perform a similar function – but lacks the knowledge of whether those functions call
each other. A Graph Neural Network combines both sources of knowledge while enabling a learnable
summarization that can highlight features relevant for classification.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Graphical Neural Networks</title>
        <p>We based our GNN design on the molecular classification network described in Veličković’s
GraphAttention Networks [22]. Structurally, our GNN is a representation of a malware sample’s call graph
where a node is a function and edges signify a call from one function to another. Additionally, nodes
contain attributes consisting of the LLM Embedding of a function and various aspects of the function’s
metadata such as the number of inputs, function size and return type. It uses four Graph Attention
Layers (employing Brody’s “GAT v2” enhancement [23]) interspersed with normalization layers and
Leaky ReLU activations, and followed by global Mean Pooling layer and a multi-layer feedforward
classifier. We again train with cross-entropy loss using the Adam optimizer and a learning rate of 0.0001.
We use 4 concatenated attention heads on the first three GAT layers, followed by 6 pooled heads on the
ifnal.</p>
        <p>
          Within 100 training epochs, our GNN achieves 92% training accuracy using LLM embeddings and
93.22% on the holdout test data (Table 3). This exceeds the state-of-the-art 89% for static malware
analysis [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], and is competitive with open-source dynamic analysis techniques with 95% accuracy on
the same dataset [24].
        </p>
        <p>We further added global and local TDA features to determine if they contain any stylistic indicators
not captured by the GNN. Table 3 details the results of adding these features. While local TDA features
(with LLM embeddings) out performed everything on the validation set. Training with the LLM features
alone achieved the best accuracy (93.22%). Given these results, it seems that these TDA features do not
contain any extra information that can be learned by a GNN.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>Our results indicate that while function-level analysis is efective, global features also play a crucial role.
The use of TDA features alone achieved a 73% accuracy, highlighting that global structural features
carry significant attributional information. In the malware domain, the specific goals and functions of
malware can be closely tied to their authors, indicating that function and attribution may be highly
correlated. This highlights the importance of a multifaceted approach.</p>
      <p>Our GNN, which integrates both structural features (i.e. call graph) and function-level embeddings
(LLM embeddings), outperforms other configurations and achieves state-of-the-art for static binary
malware analysis. This multifaceted approach underscores the importance of integrating various levels
of abstraction to capture the nuances of programming style and function.</p>
      <p>The main findings of our research show that function-level analysis, enhanced by LLM embeddings,
and global structural analysis, together lead to higher accuracy in malware attribution. Our GNN, which
integrates these features, significantly outperforms configurations that use either feature alone. This
result suggests that both the detailed stylistic elements and the broader structural aspects are crucial
for accurate attribution, filling the gap in existing static analysis techniques.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Future Directions</title>
      <p>To further enhance the accuracy of our approach, several improvements can be considered. While call
graphs are inherently directed, our Graph Attention layers do not account for directional information.
A GNN designed for directed graphs, like Perlmutter’s MagNet [25], may be able to mine this
additional source of information. In addition, while call graphs provide valuable structural insights, other
representations such as Control Flow Graphs, Abstract Syntax Trees, or Program Dependence Graphs
could ofer additional layers of information. Integrating these representations may provide a more
holistic view of the code, albeit with increased complexity in processing and analysis. Finally, we hope
that future work will replicate our results using open-source LLM-based embedders like Starencoder,
enabling this pipeline to be used in situations in which one can’t send the decompiled source code to
commercial servers.</p>
      <p>A critical next step is to evaluate the resilience of our method to real-world challenges, such as
malware obfuscation. Techniques that artificially flatten call graphs, encrypt functions, or otherwise
complicate decompilation need to be tested against our model to ensure its robustness. While we suspect
our GNN might be less susceptible to such manipulations compared to traditional static techniques, the
LLM embeddings’ dependency on decompilation quality remains a concern.</p>
      <p>As it stands, our model is deliberately minimal, ignoring features traditionally employed for malware
attribution such as API calls, linguistic comment analyses and runtime profiling. This approach allows
us to better understand the eficacy of our novel features. Given the efectiveness demonstrated by
these features, future work could integrate our GNN’s output into a broader set of malware features and
apply downstream classification, potentially surpassing even the current state-of-the-art in dynamic
analysis.</p>
      <p>In conclusion, our research advances the field of static malware analysis by demonstrating the eficacy
of combining global and function-level features within a GNN framework. This approach not only
enhances attribution accuracy but also opens new avenues for incorporating additional features and
representations. By addressing the above-identified limitations and exploring further enhancements,
future studies can build upon our findings to develop even more robust and accurate methods for
malware attribution, thereby strengthening cybersecurity defenses.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>We are extremely grateful to the Georgia Tech Research Institute (GTRI) Graduate Research Internship
Program (GRIP)2 for supporting this efort during the Summer of 2023 and beyond. Further, thanks to
Bob Wright for reviewing the work and providing valuable feedback.</p>
      <p>M. Dey, Z. Zhang, N. Fahmy, U. Bhattacharyya, W. Yu, S. Singh, S. Luccioni, P. Villegas, M.
Kunakov, F. Zhdanov, M. Romero, T. Lee, N. Timor, J. Ding, C. Schlesinger, H. Schoelkopf, J. Ebert,
T. Dao, M. Mishra, A. Gu, J. Robinson, C. J. Anderson, B. Dolan-Gavitt, D. Contractor, S. Reddy,
D. Fried, D. Bahdanau, Y. Jernite, C. M. Ferrandis, S. Hughes, T. Wolf, A. Guha, L. von Werra,
H. de Vries, StarCoder: may the source be with you! (2023). URL: http://arxiv.org/abs/2305.06161.
arXiv:2305.06161 [cs].
[15] J. Petrík, Google code jam dataset, https://github.com/Jur1cek/gcj-dataset, 2008-2021. Accessed:</p>
      <p>December 12, 2023.
[16] ’Meta’, https://ai.meta.com/llama/, 2023. Accessed: December 12, 2023.
[17] F. Chazal, B. Michel, An introduction to topological data analysis: fundamental and practical
aspects for data scientists, Frontiers in artificial intelligence 4 (2021) 667963.
[18] M. Carriere, F. Chazal, Y. Ike, T. Lacombe, M. Royer, Y. Umeda, Perslay: A neural network layer
for persistence diagrams and new graph topological signatures, in: S. Chiappa, R. Calandra
(Eds.), Proceedings of the Twenty Third International Conference on Artificial Intelligence and
Statistics, volume 108 of Proceedings of Machine Learning Research, PMLR, 2020, pp. 2786–2796.</p>
      <p>URL: https://proceedings.mlr.press/v108/carriere20a.html.
[19] P. Bubenik, Statistical topological data analysis using persistence landscapes, Journal of Machine</p>
      <p>Learning Research 16 (2015) 77–102. URL: http://jmlr.org/papers/v16/bubenik15a.html.
[20] S. Tomonaga, Classifying malware using import api and fuzzy hashing – impfuzzy –, https:
//blogs.jpcert.or.jp/en/2016/05/classifying-mal-a988.html, 2016. Accessed: 2024-07-09.
[21] A. Tong, F. Wenkel, K. MacDonald, S. Krishnaswamy, G. Wolf, Data-driven learning of geometric
scattering networks, 2020. arXiv:2010.02415.
[22] P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, Y. Bengio, Graph Attention Networks,
International Conference on Learning Representations (2018). URL: https://openreview.net/forum?
id=rJXMpikCZ.
[23] S. Brody, U. Alon, E. Yahav, How attentive are graph attention networks?, arXiv preprint
arXiv:2105.14491 (2021).
[24] C. Boot, Supervised learning for state-sponsored malware (2019). URL: http://www.cs.ru.nl/E.Poll/
papers/MalwareStateAttribution2019.pdf.
[25] X. Zhang, Y. He, N. Brugnone, M. Perlmutter, M. Hirn, Magnet: A neural network for directed
graphs, Advances in neural information processing systems 34 (2021) 27003–27015.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Samadzadeh</surname>
          </string-name>
          ,
          <article-title>Extraction of java program fingerprints for software authorship identification</article-title>
          ,
          <source>Journal of Systems and Software</source>
          <volume>72</volume>
          (
          <year>2004</year>
          )
          <fpage>49</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Caliskan-Islam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Harang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Voss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Greenstadt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yamaguchi</surname>
          </string-name>
          ,
          <article-title>Deanonymizing programmers via code stylometry</article-title>
          ,
          <source>in: Proceedings of the 24th USENIX Security Symposium</source>
          ,
          <year>2015</year>
          . URL: https://www.usenix.org/system/files/conference/usenixsecurity15/ sec15-paper
          <string-name>
            <surname>-</surname>
          </string-name>
          caliskan-islam.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Alsulami</surname>
          </string-name>
          , E. Dauber,
          <string-name>
            <given-names>R.</given-names>
            <surname>Harang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mancoridis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Greenstadt</surname>
          </string-name>
          ,
          <article-title>Source code authorship attribution using long short-term memory based networks</article-title>
          ,
          <source>in: Computer Security-ESORICS 2017: 22nd European Symposium on Research in Computer Security</source>
          , Oslo, Norway,
          <source>September 11-15</source>
          ,
          <year>2017</year>
          , Proceedings,
          <source>Part I 22</source>
          , Springer,
          <year>2017</year>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Alrabaee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Debbabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>A survey of binary code fingerprinting approaches: Taxonomy, methodologies, and features</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>55</volume>
          (
          <year>2022</year>
          )
          <volume>19</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          :
          <fpage>41</fpage>
          . URL: https://dl.acm. org/doi/10.1145/3486860. doi:
          <volume>10</volume>
          .1145/3486860.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H. S.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <string-name>
            <surname>EMBER:</surname>
          </string-name>
          <article-title>An Open Dataset for Training Static PE Malware Machine Learning Models</article-title>
          , ArXiv e-prints (
          <year>2018</year>
          ). arXiv:
          <year>1804</year>
          .04637.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Olukoya</surname>
          </string-name>
          ,
          <article-title>Nation-state threat actor attribution using fuzzy hashing</article-title>
          ,
          <source>IEEE Access 11</source>
          (
          <year>2023</year>
          )
          <fpage>1148</fpage>
          -
          <lpage>1165</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2022</year>
          .
          <volume>3233403</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Javaheri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hosseinzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Rahmani</surname>
          </string-name>
          ,
          <article-title>Detection and elimination of spyware and ransomware by intercepting kernel-level system routines</article-title>
          ,
          <source>IEEE Access 6</source>
          (
          <year>2018</year>
          )
          <fpage>78321</fpage>
          -
          <lpage>78332</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>MITRE</surname>
          </string-name>
          ,
          <article-title>Mitre att&amp;ck groups</article-title>
          , https://attack.mitre.org/groups/,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -07-09.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Boot</surname>
          </string-name>
          , Apt malware dataset, https://github.com/cyber-research,
          <year>2019</year>
          . Accessed: December 12,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Agency</surname>
          </string-name>
          , Ghidra, https://ghidra-sre.org/,
          <year>2023</year>
          . Accessed: December 12,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11] 'Microsoft', https://github.com/microsoft/CodeBERT,
          <year>2023</year>
          . Accessed: December 12,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12] OpenAI, https://platform.openai.com/docs/guides/embeddings,
          <year>2023</year>
          . Accessed: December 12,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13] 'BigCode', https://huggingface.co/bigcode/starencoder,
          <year>2023</year>
          . Accessed: December 12,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. B.</given-names>
            <surname>Allal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Muennighof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kocetkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Akiki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          , E. Zheltonozhskii,
          <string-name>
            <given-names>T. Y.</given-names>
            <surname>Zhuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Dehaene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Davaadorj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lamy-Poirier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Monteiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Shliazhko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gontier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Meade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zebaze</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-H. Yee</surname>
            ,
            <given-names>L. K.</given-names>
          </string-name>
          <string-name>
            <surname>Umapathi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Lipkin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Oblokulov</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Murthy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Stillerman</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Abulkhanov</surname>
          </string-name>
          , M. Zocca,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>