<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Development of a smart personnel security system using machine learning⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maksym Bilychenko</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nataliia Kasianova</string-name>
          <email>nataliia.kasianova@npp.kai.edu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Serhii Smerichevskyi</string-name>
          <email>serhii.smerichevskyi@npp.kai.edu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksandr Lavrynenko</string-name>
          <email>oleksandrlavrynenko@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igor Kryvovyazyuk</string-name>
          <email>krivovyazukigor@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lutsk National Technical University</institution>
          ,
          <addr-line>75 Lvivska str., 43018 Lutsk</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>State University “Kyiv Aviation Institute”</institution>
          ,
          <addr-line>1 Liubomyra Huzara ave., 03058 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>203</fpage>
      <lpage>215</lpage>
      <abstract>
        <p>Insider threats remain one of the most challenging aspects of organizational security, particularly in the era of digital transformation and widespread remote access to sensitive data. This study proposes a machine learning-based approach to personnel security that combines Isolation Forest and Local Outlier Factor algorithms with behavioral features enhanced through the use of large language models (LLMs). To improve detection accuracy, user web activity was classified using LLM-generated labels derived from website content analysis. Experimental results demonstrate strong model performance in identifying insider activity at the user level, with high detection accuracy and minimal false classifications. In addition, time-to-detection analysis revealed that most insider threats were identified before or shortly after the onset of malicious behavior. The findings suggest that the proposed system is not only effective in capturing behavioral anomalies but also feasible for real-time deployment in enterprise environments.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;insider threat detection</kwd>
        <kwd>personnel security</kwd>
        <kwd>anomaly detection</kwd>
        <kwd>large language models</kwd>
        <kwd>isolation forest</kwd>
        <kwd>local outlier factor</kwd>
        <kwd>behavioral profiling1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Digital technologies have rapidly permeated nearly all domains of social life, including education,
mass media, the workplace, daily routines, commerce, sports, healthcare and entertainment. The
development of a digital society entails the creation of a complex yet thoroughly transformed
ecosystem in which humans and technologies coexist in new forms of interaction and cohabitation.
The integration of information and communication technologies (ICT), artificial intelligence, the
Internet of Things (IoT), and cloud computing has facilitated the emergence of an interconnected
digital environment. Within this environment, business processes can be executed with minimal
time and resource expenditure. Digital transformation initiatives are progressively focused on
automating repetitive tasks, leveraging big data analytics to inform strategic management
decisions, expanding electronic commerce, and strengthening customer interaction channels.
Nevertheless, digital transformation remains a complex, resource-intensive, and time-consuming
undertaking for most enterprises—particularly for small and medium-sized businesses. In this
context, it becomes imperative for companies to articulate clear digital strategies and define
priority areas for transformation. Based on these priorities, enterprises must adapt to ongoing
changes and implement innovative solutions across both technological and methodological
dimensions.</p>
      <p>Simultaneously, the advancement of digital transformation introduces not only new
opportunities but also a range of emerging risks and threats. These challenges are inherent both to
the transformation process itself and to the subsequent digital operations of an organization. As</p>
      <p>
        0000-0003-4657-1039 (M. Bilychenko); 0000-0001-7729-2011 (N. Kasianova); 0000-0003-2102-1524 (S. Smerichevskyi);
0000-0002-7738-161X (O. Lavrynenko); 0000-0002-8801-4700 (I. Kryvovyazyuk)
noted in a report by Deloitte [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], one of the key categories of risk involves the exposure of
confidential information and data, that is typically associated with improper handling of personal
or sensitive information related to customers, employees or business partners. Data breaches,
violations of data processing protocols and inadequate storage practices throughout the data
lifecycle constitute primary sources of these threats. Moreover, low levels of digital literacy among
staff and a lack of organizational culture around information security can generate systemic
vulnerabilities.
      </p>
      <p>
        Insider threats, social engineering attacks, and careless handling of sensitive information
frequently lead to serious security incidents. According to IBM’s Cost of a Data Breach report [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
the average global cost of mitigating a data breach reached $4.88 million in 2024—the highest
recorded figure in the past eight years. As illustrated in Figure 1, this represents a 10%
year-overyear increase, marking one of the most significant annual spikes during the observed period.
Additionally, the 2025 Thales Data Threat Report revealed that 45% of surveyed companies
worldwide experiencing a data breach, with 14% of these occurring within the past year [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Human-related risks have become especially important for enterprises undergoing digital
transformation. Ensuring personnel security is essential for stable business operations and requires
a structured approach that includes developing employees’ digital skills, improving human
resource policies, and applying up-to-date cybersecurity standards in day-to-day operations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. To
achieve this, companies need to put in place clear digital safety policies for employees, regularly
monitor how staff handle data, limit access to sensitive information based on user roles and follow
modern procedures for data handling and protection. A key part of personnel security is the ability
to detect possible threats early—especially identifying employees who might, whether intentionally
or by mistake, share confidential business information outside the organization. In other words,
timely detection of insider threats and appropriate follow-up actions are critical to protecting
enterprise security. In other words, timely detection of insider threats and appropriate follow-up
actions are critical to protecting enterprise security. Given the growing complexity and scale of
digital environments, traditional control mechanisms are no longer sufficient. This highlights the
need for advanced, data-driven approaches capable of identifying subtle behavioral patterns that
may indicate insider risk. The following section provides a review of current methodologies and
recent advances in insider threat detection, with a particular focus on algorithmic and ML-driven
solutions.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>Traditionally, personnel security has been understood as a set of measures aimed at ensuring the
stability of human capital, maintaining adequate employee qualifications, fostering effective
motivation and preventing internal threats. However, the rapid advancement of digital
technologies necessitates a reassessment of existing approaches to evaluating and forecasting the
state of this subsystem. In particular, the use of modern modeling techniques is gaining relevance—
ranging from score-based assessments of employee motivation to hierarchical models built on
fuzzy logic. These approaches enable the analysis of how changes in motivation systems, training
programs, and personnel management influence the likelihood of labor-related risks, while also
identifying the most vulnerable areas.</p>
      <p>
        Beyond conventional strategies for ensuring personnel security, insider threats have become
especially critical in the digital economy. Despite the availability of extensive software solutions
designed to guard against external intrusions, internal threats—posed by employees or individuals
with privileged access—remain the most challenging to detect and mitigate. Latest research
increasingly shifts from general personnel system management to the development of preventive
analytics tools capable of identifying risky behavior before it leads to serious consequences. In this
context, particular attention has been directed toward the use of data mining and machine learning
methods for building adaptive personnel risk management systems. These systems analyze
extensive digital traces left by employees such as activity within internal networks, access to
information resources, and behavioral pattern shifts, to detect early signs of potential insider
threats with high accuracy [
        <xref ref-type="bibr" rid="ref5 ref6 ref7 ref8">5–8</xref>
        ].
      </p>
      <p>
        Insider threat detection systems based on anomaly detection techniques rely on statistical
modeling of typical employee interactions with IT resources. Deviations from established
behavioral patterns may signal potential personnel-related threats. For instance, M.
RaissiDehkordi and D. Carr [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] proposed a multi-perspective insider threat detection framework that
integrates data from various components of the corporate network. This architecture employs a
one-class Support Vector Machine algorithm, which reduces the number of false positives and
enhances the accuracy of identifying risky employee behavior. While this method performs well in
detecting group-based insider activities, its effectiveness in forecasting individual insider threats
remains limited.
      </p>
      <p>
        An alternative approach was introduced by T. Rashid and colleagues [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], who utilized Hidden
Markov Models to model habitual employee behavior. These models are well-suited for analyzing
temporal sequences but exhibit high computational complexity as model size increases. Another
technique, developed by Y. Song et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], applies Gaussian Mixture Models to construct user
profiles and detect anomalies. Although this method demonstrates high accuracy, it is primarily
geared toward biometric identification rather than targeted detection of insider threats within
corporate systems.
      </p>
      <p>
        In contrast, G. Gavai and colleagues [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] proposed a behavior-based insider threat detection
model using digital traces of employee activity. Their framework defined 42 features across five
categories: email content and usage, login patterns, and software and web activity. The model
applies both Isolation Forest and Random Forest algorithms to detect abnormal behavior and
identifies high-risk individuals through an interactive visualization dashboard. This approach
supports proactive decision-making in personnel security by offering simplicity, adaptability, and
objectivity. However, its performance may degrade in large-scale organizational settings, requiring
additional tuning for stability.
      </p>
      <p>
        In the study by D. C. Le and N. Zincir-Heywood [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], a framework for insider threat detection
was introduced based on an ensemble of unsupervised learning algorithms, including Isolation
Forest, Autoencoder, Local Outlier Factor (LOF), and Lightweight Online Detector of Anomalies
(LODA). Among these, the Autoencoder and LOF models yielded the highest detection
performance. Additionally, the use of ensemble voting across models was shown to enhance both
the accuracy and robustness of the detection system.
T. Al-Shehari et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] proposed a detection approach using the Density-Based Local Outlier
Factor (DBLOF), specifically adapted for imbalanced datasets such as CERT R4.2. Unlike traditional
methods, DBLOF focuses on analyzing local density, enabling the precise identification of rare yet
potentially harmful behavioral anomalies. The method demonstrated high accuracy and proved
effective in detecting insider threats under real-world corporate conditions.
      </p>
      <p>
        C. Song and colleagues [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] introduced Audit-LLM—a multi-agent insider threat detection
system that analyzes event logs using large language models (LLMs). To overcome the limitations
of LLMs in processing complex activity logs, the system incorporates three specialized agents
alongside a multiagent debate mechanism designed to improve inference accuracy. The approach
demonstrated superior detection performance compared to existing solutions, particularly in terms
of explanatory power and handling large-scale log files.
      </p>
      <p>The comparison of models is presented in Table 1.</p>
      <p>
        Author &amp; Year
M. Raissi-Dehkordi
&amp; D. Carr, 2011 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
      </p>
      <sec id="sec-2-1">
        <title>T. Rashid et al., 2016 [10] Y. Song et al., 2013 [11]</title>
      </sec>
      <sec id="sec-2-2">
        <title>Models One-Class SVM</title>
      </sec>
      <sec id="sec-2-3">
        <title>Hidden Markov Models Gaussian Mixture Models</title>
        <p>
          G. Gavai et al., 2015
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Isolation Forest,</title>
        <p>Random Forests
D. C. Le &amp; N.</p>
        <p>
          Zincir-Heywood,
2021 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]
T. Al-Shehari et al.,
2024 [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]
C. Song et al., 2024
[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Isolation Forest,</title>
        <p>Autoencoder, LOF,</p>
        <p>LODA
Density-Based
Local Outlier</p>
        <p>Factor (LOF)
LLM-based
MultiAgent System
(Audit-LLM)</p>
      </sec>
      <sec id="sec-2-6">
        <title>Data</title>
        <p>Simulated
data via
OPNET
CERT</p>
        <p>RUU
Research
Dataset</p>
        <p>Vegas
CERT R4.2,</p>
        <p>R6.2</p>
      </sec>
      <sec id="sec-2-7">
        <title>CERT R4.2 CERT R4.2, R5.2, PicoDomain</title>
      </sec>
      <sec id="sec-2-8">
        <title>Summary</title>
        <p>High accuracy for group attacks;
moderate performance for individual
threats
Model is easy to train but becomes
computationally intensive with scale
High accuracy, though primarily
focused on biometric profiling</p>
      </sec>
      <sec id="sec-2-9">
        <title>Isolation Forest outperformed other</title>
        <p>models in terms of detection
effectiveness
Autoencoder and LOF achieved the
highest detection accuracy</p>
      </sec>
      <sec id="sec-2-10">
        <title>High accuracy for imbalanced</title>
        <p>datasets; interpretability remains
limited
High detection accuracy and
improved explainability; scalability
constraints noted</p>
        <p>In summary, the most prevalent and effective approaches to insider threat detection within
personnel security systems are based on the use of Isolation Forest and Local Outlier Factor
algorithms. Their effectiveness lies in analyzing behavioral data collected from corporate networks,
such as event logs, access records, and system activity. These algorithms have demonstrated high
performance in identifying anomalous employee behavior that may indicate internal security
threats.</p>
        <p>
          However, much of the existing research primarily focuses on optimizing model accuracy while
paying limited attention to the practical challenges of deploying such solutions in real
organizational settings. Key challenges, including the adaptability of detection models to evolving
operational environments, the scalability of solutions across enterprise contexts, and their seamless
integration into existing IT infrastructures, remain underexplored in the current body of research.
Furthermore, the potential of emerging technologies such as large language models (LLMs) to
enhance system scalability, computational efficiency, and the interpretability of detection outcomes
has received limited scholarly attention. Recent works have highlighted the importance of
integrating real-time monitoring and secure data handling into insider threat detection systems.
For instance, Zhuravchak et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] demonstrated the application of Berkeley Packet Filter for
monitoring ransomware activities, emphasizing proactive threat tracking. Additionally,
comparative studies of deep learning models [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] and secure data classification frameworks [
          <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
          ]
underline the potential of combining advanced analytics with robust storage policies to enhance
both the accuracy and security of personnel risk management systems. Given these limitations, the
present research will focus on combining Isolation Forest and LOF algorithms with LLM-based
techniques to develop a modern, adaptive, accurate and scalable insider threat detection system.
This approach aims to move beyond academic metrics and prioritize practical applicability in
realworld enterprise environments.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Data and methodology</title>
      <p>Building upon the insights from the literature and addressing the identified research gaps, this
study proposes a hybrid insider threat detection approach that integrates both established anomaly
detection techniques and recent advancements in data-driven modeling. Specifically, two
algorithms were selected for model development: Isolation Forest and Local Outlier Factor (LOF),
both of which have demonstrated high accuracy in previous studies while offering complementary
strengths in detecting anomalous behavior.</p>
      <p>The LOF algorithm is based on the concept of local density. Unlike global models that assess
each data point in relation to the entire dataset, LOF evaluates an observation in the context of its
immediate neighborhood. The underlying assumption is that an anomalous instance will exhibit
significantly lower local density compared to its neighbors—indicating its presence in a relatively
sparse region of the feature space. Mathematically, the LOF algorithm is implemented through
several sequential steps. The first involves calculating the distance to the  -nearest neighbors 
distance().</p>
      <p>This parameter defines the radius of the local neighborhood around an observation p within
which the analysis is conducted. Subsequently, the algorithm calculates the reachability distance, a
measure designed to prevent unrealistically low density estimates that may result from the
presence of extremely close neighboring points</p>
      <p>reach dist k ( p , o )= max ( k− distance ( o ) , dist ( p , o )).</p>
      <p>The next step involves computing the local reachability density () of the observation:
where N k ( p) is the collection of the k closest points to p, based on a distance metric which is
typically Euclidean distance. This metric quantifies how densely the pointp is surrounded by its
knearest neighbors. In the final step, the LOF score is computed:
lrdk ( p)=</p>
      <p>∑
o ∈ Nk( p)</p>
      <p>| N k ( p)|
reach dist k ( p , o )</p>
      <p>.</p>
      <p>∑
L O F k ( p)= o ∈ Nk( p) lrdk ( p) .</p>
      <p>| N k ( p)|</p>
      <p>lrdk ( o )</p>
      <p>The LOF score represents the ratio of the average local reachability density of a point’s
neighbors to the local reachability density of the point itself. Formally, if L O F k ( p)≈ 1 the
behavior of point p is considered normal. However, if L O F k ( p)≫ 1, the point is identified as
anomalous, indicating that its local density is significantly lower than that of its surrounding
neighbors.
The second algorithm selected for this study is Isolation Forest. Unlike most traditional methods
that model normal behavior and identify deviations, Isolation Forest is based on the principle of
isolation. The core idea is that anomalous instances are more susceptible to isolation—they tend to
be located further away from dense clusters of points and can therefore be separated more quickly
through recursive partitioning of the feature space. The model consists of an ensemble of isolation
trees (iTrees). The construction of each isolation tree involves the random selection of a feature,
followed by the choice of a split value within the domain of that feature. This recursive
partitioning continues until the instance is fully isolated in a leaf node, with the resulting isolation
depth serving as an empirical indicator of separability—where shorter path lengths typically
correspond to anomalous instances due to their greater isolation from the general data distribution.</p>
      <p>
        The overall structure of the algorithm can be summarized as follows: the isolation forest builds
a tree-based structure capable of efficiently isolating each individual instance [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Due to the
sensitivity of this method to isolation characteristics, outliers tend to appear closer to the root of
the tree, while normal points are usually isolated deeper in the tree. This recursive isolation
mechanism underpins the effectiveness of the approach, and the resulting structure is referred to as
an isolation tree or iTree. An isolation forest consists of an ensemble of iTrees built from the
dataset, where anomalies are identified as those instances with a significantly shorter average path
length across the trees.
      </p>
      <p>The mathematical formalization of the Isolation Forest algorithm can be introduced as follows.
In the first step, the expected path length for isolating a data point x in a tree constructed from n
observations is computed:</p>
      <p>where ℎ() denotes the path length required to isolate the point x, and  is a correction constant
(the Euler’s constant). Following this, the anomaly score is calculated as:</p>
      <p>E [h( x) ]≈ log2 ( n)+ γ,</p>
      <p>− E [h( x)]
s ( x , n)= 2 c (n) ,
where () represents the average path length, which, for a binary search tree, can be
approximated as follows:</p>
      <p>2 ( n− 1)
c ( n)= 2 H ( n− 1)− , (6)</p>
      <p>n
where () denotes the i-th harmonic number, defined as the sum of the reciprocals of the first i
positive integers.</p>
      <p>The interpretation of results obtained using the Isolation Forest algorithm is based on the
anomaly score s ( x), which ranges from 0 to 1. Values approaching 1 (i.e., s ( x ) → 1) indicate a
high likelihood that the observation is anomalous, suggesting that the behavior of the
corresponding instance significantly deviates from the norm observed in the overall dataset.
Conversely, values of s ( x) &lt; 0.5 are indicative of typical, non-anomalous behavior consistent with
the majority of the observations. Scores near 0.5 require further investigation, as they may
represent borderline or potentially risky cases that have not yet escalated into overt violations.</p>
      <p>
        Due to the limited availability of actual corporate data for insider threat research, this study
utilizes the publicly accessible CERT Insider Threat Dataset [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Although the CERT dataset does
not originate from real-world organizational environments, it is synthetically generated to support
the development and evaluation of insider threat detection models and the analysis of personnel
security risks. Despite its artificial nature, the dataset incorporates behavioral patterns that closely
resemble real employee activities within an organizational information system. The CERT Insider
Threat Tools include logs of employee interactions with computer systems, complemented by
selected organizational attributes. The dataset structure consists of relational tables containing
attributes such as user identifiers, event timestamps, and descriptions of observed actions. For this
study, version R4.2 of the CERT dataset was selected, as it provides a sufficient number of
observations for both normal users and insider threat cases. Specifically, this version includes data
on approximately 1000 employees, among whom about 70 are labeled as potential insider threats to
organizational personnel security.
      </p>
      <p>The first step of the analysis involved constructing a focused sample of users for more detailed
investigation. Since the full dataset is quite large and contains more information than necessary for
initial model development, a decision was made to reduce its size while preserving a balanced
distribution between classes. Specifically, 20 out of the 70 users flagged as insider threats were
selected, along with 30 users from the remaining 930 who exhibited no suspicious behavior. This
resulted in a working sample of 50 employees, with insiders making up 40% of the group. Class
balance was a key factor in building this sample. In the original dataset, insider cases were
relatively rare accounting for less than 10% of all observations. Using the original distribution
would have led to a strong class imbalance, which could hinder the model’s ability to detect insider
threats effectively. By selecting a more balanced subset, we aimed to improve the model’s capacity
to learn from and identify unusual behavioral patterns.</p>
      <p>A key innovation in this study is the enhancement of existing behavioral data through the
integration of additional features generated by large language models. This approach aims to
increase the informativeness of features relevant to personnel security analysis, particularly in
relation to employees’ web activity. The analysis focuses on two primary risk scenarios:
1.</p>
      <p>Visiting job search websites, which may be linked to attempts to take sensitive data when
changing jobs.</p>
      <p>Accessing websites that could be used to share confidential company information without
permission.</p>
      <p>Traditional threat classification methods often rely on fixed lists of domains (such as job search
or data leak sites), but these are inflexible and cannot detect new or unlisted threats. To address
this, we propose a content-based approach that uses the text of visited websites to determine their
type. Based on the content and url fields in the CERT dataset, we created two binary indicators,
which were evaluated using LLM prompts to classify the intent behind each web visit. This
LLMdriven feature enrichment allows for more nuanced and adaptive threat identification beyond rigid
rule-based systems.</p>
      <p>
        Given the large volume of data it was essential to select a fast and resource-efficient model for
generating LLM-based responses. For this purpose, the Phi-3.5-mini-instruct model [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] was
chosen. With fewer than 4 million parameters, this lightweight model can be deployed locally on a
personal computer, enabling high-speed processing without significant loss in output quality. To
further enhance processing throughput, the vLLM library [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], a state-of-the-art tool for optimized
text generation, was integrated into the pipeline. In total, more than one million web activity
records were processed within approximately five hours using GPU resources provided by Google
Colab A100 [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. As a result, two new variables were generated based on the model’s responses
indicating whether the user was likely searching for a job and whether the user’s activity
suggested potential information leakage. One of the key steps in feature engineering for modeling
personnel security threats was the segmentation of employee activity based on time, distinguishing
between actions occurring during standard working hours and those outside of them. To support
this, an analysis of user activity distribution across the 24-hour period was conducted, allowing for
the identification of core working hours (Figure 2). Based on this, two separate groups of variables
were created: one representing activity within working hours and the other representing activity
beyond that period. This approach enabled differentiation of behavioral patterns and contributed to
improved accuracy in anomaly detection. Particular attention was given to the development of
novel variables not previously used in insider threat research. For example, the variable
num_site_jobs_1_growth_total_40_days was introduced to capture the maximum increase in visits
to job search websites over the past 40 days. This feature was motivated by the assumption that
seeking new employment is typically a prolonged process and not limited to a short time window,
which may significantly elevate the risk of insider activity. Another important variable,
num_connect_after_job_search, was designed to reflect the multiplicative interaction between the
increase in the number of connected devices and the growth in job search activity over the same
40-day period. The rationale behind this feature is to detect a specific behavioral pattern in which
an employee first intensifies job search efforts and subsequently connects additional devices to
transfer or store sensitive corporate data, indicating high likelihood of data leakage. The final
dataset comprised of 5 input advanced variables over 15000 instances, where each instance
represents the behavioral profile of a specific user on a given day.
      </p>
      <p>The complete dataset was then split into training and test sets. Model training was performed
on the training subset, while evaluation was conducted using the test set. Following best practices,
the training data was composed primarily of normal users, whereas the test set included a higher
proportion of users exhibiting anomalous or risky behavior. Specifically, the split was performed at
the user level: 19 users were assigned to the test set, 68\% of whom were identified as insiders,
while 31 users were included in the training set, of which only 22\% were labeled as insiders. The
distribution of data instances was similarly stratified, with approximately 65\% of records allocated
to the training set.</p>
      <p>The next logical step in the study involved selecting appropriate evaluation metrics to assess the
performance of the developed models. Three commonly used classification metrics from machine
learning practice were chosen: Detection Rate (DR), Accuracy, and F1 Score.</p>
      <p>The primary and most critical metric is the Detection Rate, which reflects the proportion of true
insider threats correctly identified by the model out of the total number of actual insiders. This
metric is particularly important in the context of personnel security, where the main objective is to
minimize the number of undetected threats. The formula for calculating the Detection Rate is as
follows:</p>
      <p>Number of detected insiders
D R=</p>
      <p>Total number of actual insiders
The second evaluation metric is Balanced Accuracy, which reflects the overall correctness of the
model’s predictions. It measures the proportion of all correctly classified instances relative to the
total number of observations. The formula for calculating Balanced Accuracy is:</p>
      <p>Balanced Accuracy=</p>
      <p>T P + TN
T P + F P + TN + F N
,
where: TP—correctly identified insiders, TN—correctly identified non-insiders, FP—normal
employees incorrectly labeled as insiders, FN—actual insiders missed by the model.</p>
      <p>The third key evaluation metric is the F1 Score, which represents the harmonic mean of
Precision and Recall. This metric provides a more balanced assessment of model performance,
particularly in scenarios where the dataset exhibits significant class imbalance. The F1 Score is
especially useful when both false positives and false negatives carry important consequences. The
formula for calculating the F1 Score is:</p>
      <p>F 1=
2 P R ,</p>
      <p>P + R
R =</p>
      <p>T P
T P + F N
where P denotes precision and calculated as follow:
Accordingly, the evaluation of these metrics provides a multifaceted understanding of the
model’s performance, capturing both its ability to promptly detect insider threats and its overall
classification accuracy with respect to employee activity.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental results</title>
      <p>Based on the engineered features, anomaly detection models were trained to identify potential
insider threats. The evaluation was conducted at two levels: individual daily activity records and
aggregated user-level behavior. This section presents the results of model performance at the daily
level, where each data point represents a single day of activity for a specific user. The evaluation
metrics are summarized in Table 2.</p>
      <p>As shown in Table 2, both models demonstrate relatively low F1 Scores, suggesting that
accurately identifying anomalous behavior based solely on daily patterns remains challenging. This
limitation is likely due to the subtle and context-dependent nature of insider activity, which may
not manifest clearly within a single day’s behavioral footprint.</p>
      <p>Nevertheless, the Balanced Accuracy values are moderately higher, particularly for the Isolation
Forest model, which outperforms LOF by more than two percentage points. Balanced Accuracy,
which accounts for both sensitivity and specificity in imbalanced datasets, indicates that Isolation
Forest has a slightly better ability to distinguish between normal and anomalous classes. This result
highlights its relative robustness in handling behavioral data where insider activity represents a
minority class.</p>
      <p>A substantially different outcome was observed when the analysis shifted from daily-level
activity to user-level behavior. In this approach, all behavioral data were aggregated per user,
allowing for a more holistic representation of activity patterns. The results of this evaluation are
presented in Table 3. As shown in Table 3, the Isolation Forest model achieved outstanding
performance across all metrics, including a 100% Detection Rate, meaning that it correctly
identified all insider threats in the test set.
(8)
(9)
(10)
(11)
Moreover, it achieved perfect Balanced Accuracy and F1 Score, indicating a high degree of
precision and recall with no false positives or negatives. This level of performance suggests that,
when user activity is considered in aggregate rather than on a daily basis, distinct behavioral
patterns become more detectable and easier to classify. Although the LOF model also achieved a
100% Detection Rate, its Balanced Accuracy (58%) and F1 Score (84%) were substantially lower than
those of Isolation Forest. This indicates that while LOF was able to detect all insiders, it
misclassified a larger portion of non-insiders, resulting in reduced precision and overall
classification stability.</p>
      <p>It is important to emphasize that in insider threat detection tasks, timeliness of detection is as
critical as the detection itself. The earlier a threat is identified, the greater the likelihood of
implementing preventive measures before confidential information is leaked or organizational
harm occurs. To illustrate this aspect, the modeling results for an individual user are presented in
Figure 3, which shows the temporal dynamics of several key features used by the Isolation Forest
model for user “KRL0501” over the analyzed period.</p>
      <p>Figure 3 reveals several noteworthy behavioral signals. The purple line shows intermittent
spikes, suggesting periods where the user connected additional devices shortly after job search
activity which is a potentially suspicious pattern. Red vertical dotted lines indicate the actual
periods of insider activity, as defined by the original dataset authors. Overall, the user’s activity
pattern appears irregular, characterized by alternating periods of high engagement and inactivity.
Especially concerning are the timeframes where elevated job search behavior coincides with
increased access to suspicious files or tools. This convergence of risk indicators may signal a
heightened probability of insider threat—for instance, a user actively seeking new employment
while accessing or exfiltrating sensitive resources. Visualizations of this kind are critical tools for
insider threat detection systems, as they capture the interplay between behavioral indicators and
provide actionable insights into user intent and risk escalation over time. Such patterns can help
security teams prioritize monitoring efforts and trigger early interventions.</p>
      <p>To further assess the effectiveness of the proposed insider threat detection model, an additional
analysis was conducted to examine the timing of the model’s first alert relative to the actual onset
of insider activity. Given the critical importance of early detection in mitigating risks such as data
leaks or internal sabotage, a key objective of this analysis was to evaluate how promptly the model
responds to anomalous behavioral patterns. A comparative study was carried out, measuring the
time between the model’s first detection of suspicious activity and the verified beginning of insider
behavior. The findings are summarized below:



6 insiders were flagged either on the exact day of their first suspicious activity or several
days prior, indicating the model’s high sensitivity and its capacity to detect emerging
threats at a very early stage.
4 users were identified within the first 31 days after the onset of insider activity, confirming
the model’s effectiveness under medium- and short-term monitoring scenarios.</p>
      <p>All insiders were detected before the end of their malicious activity period, demonstrating
the model’s reliability in completing the detection-response cycle before significant harm
was done.</p>
      <p>These results provide strong evidence that the model not only performs well in detecting
anomalies but also offers timely intervention, which is essential for proactive threat mitigation.
Therefore, the developed approach can be considered a robust and practical solution for real-time
monitoring of employee activity in enterprise environments. It is well-suited for supporting timely,
evidence-based decision-making in the context of personnel security.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This study introduced a data-driven approach to insider threat detection by combining
unsupervised machine learning algorithms such as Isolation Forest and Local Outlier Factor with
behavioral data and novel features generated through LLMs. Using the CERT R4.2 dataset, the
models were evaluated at both daily and user-aggregated levels. While detection performance on
daily records was limited, the user-level analysis yielded highly promising results, with the
Isolation Forest model achieving 100% detection rate, perfect balanced accuracy, and no false
positives. A key innovation was the integration of LLM-based content classification to dynamically
assess web activity associated with job search and data leak risks. These enriched features
improved the model’s ability to capture complex behavioral patterns indicative of insider threats.
Time-to-detection analysis further demonstrated that the model could identify risks before or
shortly after the start of malicious activity, supporting its use for early intervention. Overall, the
results confirm the high effectiveness, robustness, and practical utility of the proposed approach in
real-world personnel security contexts.</p>
      <p>Based on the outcomes of this study, the following directions are recommended for practical
implementation and broader adaptation of the proposed model:
</p>
      <p>Integration into real-world monitoring systems. The results support the effective
integration of the Isolation Forest model into corporate employee monitoring platforms. Its
ability to detect suspicious behavior in a timely manner makes it a valuable tool for
mitigating the risk of internal data leaks.</p>
      <p>Strengthening internal control and risk management. Real-time detection of insider threats
enables organizations not only to respond to current risks but also to proactively prevent
future incidents. The model can be used as a key part of internal control systems, helping
quickly detect unusual behavior that may point to data misuse, access abuse, or attempts to
leak confidential information.</p>
      <p>Adaptation across sectors and organizations. Due to its flexibility, the Isolation Forest
model can be adapted for use in various industries, including finance, government and
technology. Its deployment can be tailored to the specific operational needs and risk
profiles of different organizations, making it a scalable solution for enhancing personnel
security in high-stakes environments.</p>
      <p>In conclusion, the proposed approach demonstrates high potential as a practical tool for
realtime monitoring and early detection of insider threats. It offers a balanced combination of detection
accuracy, timeliness, and operational applicability, making it well-suited for deployment in
enterprise security systems operating in dynamic, data-rich environments.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>While preparing this work, the authors used the AI programs Grammarly Pro to correct text
grammar and Strike Plagiarism to search for possible plagiarism. After using this tool, the authors
reviewed and edited the content as needed and took full responsibility for the publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Deloitte</surname>
          </string-name>
          , Managing Risk in Digital Transformation,
          <year>2018</year>
          . https://www2.deloitte.com/content/ dam/Deloitte/za/Documents/risk/za_managing_risk_in_digital_transformation_112018.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>IBM</given-names>
            <surname>Security</surname>
          </string-name>
          ,
          <article-title>Cost of a Data Breach: A Million-Dollar Race to Detect</article-title>
          and Respond,
          <year>2024</year>
          . https://www.ibm.com/reports/data-breach
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Thales</given-names>
            <surname>Group</surname>
          </string-name>
          ,
          <source>Thales Data Threat Report: Global Edition</source>
          ,
          <year>2025</year>
          . https://cpl.thalesgroup.com
          <article-title>/ data-threat-report</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V. G.</given-names>
            <surname>Goulart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. B.</given-names>
            <surname>Liboni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. O.</given-names>
            <surname>Cezarino</surname>
          </string-name>
          ,
          <article-title>Balancing Skills in the Digital Transformation Era: The Future of Jobs and the Role of Higher Education</article-title>
          ,
          <source>Industry and Higher Education</source>
          <volume>36</volume>
          (
          <year>2022</year>
          )
          <fpage>118</fpage>
          -
          <lpage>127</lpage>
          . doi:
          <volume>10</volume>
          .1177/09504222211029796
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Buhas</surname>
          </string-name>
          , et al.,
          <article-title>Using Machine Learning Techniques to Increase the Effectiveness of Cybersecurity</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems</source>
          , vol.
          <volume>3188</volume>
          , no.
          <issue>2</issue>
          (
          <year>2021</year>
          )
          <fpage>273</fpage>
          -
          <lpage>281</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhebka</surname>
          </string-name>
          , et al.,
          <article-title>Methodology for Predicting Failures in a Smart Home based on Machine Learning Methods</article-title>
          ,
          <source>in: Workshop on Cybersecurity Providing in Information and Telecommunication Systems, CPITS</source>
          , vol.
          <volume>3654</volume>
          (
          <year>2024</year>
          )
          <fpage>322</fpage>
          -
          <lpage>332</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
             
            <surname>Adamantis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
             
            <surname>Sokolov</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
           Skladannyi,
          <article-title>Evaluation of State-of-the-Art Machine Learning Smart Contract Vulnerability Detection Method, Advances in Computer Science for Engineering and Education VII, vol</article-title>
          .
          <volume>242</volume>
          (
          <year>2025</year>
          )
          <fpage>53</fpage>
          -
          <lpage>65</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -84228-
          <issue>3</issue>
          _
          <fpage>5</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>O.</given-names>
            <surname>Milov</surname>
          </string-name>
          et al.,
          <article-title>Development of Methodology for Modeling the Interaction of Antagonistic Agents in Cybersecurity Systems</article-title>
          ,
          <string-name>
            <surname>Eastern-European</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <source>Enterp. Technol. 2</source>
          .
          <issue>9</issue>
          (
          <issue>98</issue>
          ) (
          <year>2019</year>
          )
          <fpage>56</fpage>
          -
          <lpage>66</lpage>
          . doi:
          <volume>10</volume>
          .15587/
          <fpage>1729</fpage>
          -
          <lpage>4061</lpage>
          .
          <year>2019</year>
          .164730
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Raissi-Dehkordi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Carr</surname>
          </string-name>
          ,
          <article-title>A Multi-Perspective Approach to Insider Threat Detection, in: 2011 MILCOM Military Communications Conf</article-title>
          .,
          <year>2011</year>
          ,
          <fpage>1164</fpage>
          -
          <lpage>1169</lpage>
          . doi:
          <volume>10</volume>
          .1109/MILCOM.
          <year>2011</year>
          . 6127457
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Rashid</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Agrafiotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Nurse</surname>
          </string-name>
          ,
          <string-name>
            <surname>A New</surname>
          </string-name>
          <article-title>Take on Detecting Insider Threats: Exploring the Use of Hidden Markov Models</article-title>
          ,
          <source>in: Proceedings of the 8th ACM CCS Int. Workshop on Managing Insider Security Threats</source>
          , MIST'16,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2016</year>
          ,
          <fpage>47</fpage>
          -
          <lpage>56</lpage>
          . doi:
          <volume>10</volume>
          .1145/2995959.2995964
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Salem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hershkop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Stolfo</surname>
          </string-name>
          ,
          <article-title>System Level User Behavior Biometrics using Fisher Features and Gaussian Mixture Models</article-title>
          ,
          <source>in: 2013 IEEE Security and Privacy Workshops</source>
          ,
          <year>2013</year>
          ,
          <fpage>52</fpage>
          -
          <lpage>59</lpage>
          . doi:
          <volume>10</volume>
          .1109/SPW.
          <year>2013</year>
          .33
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Gavai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sricharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gunning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rolleston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hanley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <article-title>Detecting Insider Threat from Enterprise Social and Online Activity Data</article-title>
          ,
          <source>in: Proceedings of the 7th ACM CCS Int. Workshop on Managing Insider Security Threats</source>
          , MIST'15,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2015</year>
          ,
          <fpage>13</fpage>
          -
          <lpage>20</lpage>
          . doi:
          <volume>10</volume>
          .1145/2808783.2808784
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Zincir-Heywood, Anomaly Detection for Insider Threats using Unsupervised Ensembles</article-title>
          ,
          <source>IEEE Transactions on Network and Service Management</source>
          <volume>18</volume>
          (
          <year>2021</year>
          )
          <fpage>1152</fpage>
          -
          <lpage>1164</lpage>
          . doi:
          <volume>10</volume>
          .1109/TNSM.
          <year>2021</year>
          .3071928
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T. A</given-names>
            .
            <surname>Al-Shehari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rosaci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Al-Razgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Alfakih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kadrie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nawaz</surname>
          </string-name>
          ,
          <article-title>Enhancing Insider Threat Detection in Imbalanced Cybersecurity Settings using the Densitybased Local Outlier Factor Algorithm</article-title>
          ,
          <source>IEEE Access 12</source>
          (
          <year>2024</year>
          )
          <fpage>34820</fpage>
          -
          <lpage>34834</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2024</year>
          .3373694
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Song</surname>
          </string-name>
          , L. Ma, J. Zheng,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Audit-llm: Multi-agent collaboration for log-based insider threat detection</article-title>
          ,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2408.08902
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhuravchak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tolkachova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piskozub</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dudykevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Korshun</surname>
          </string-name>
          , Monitoring Ransomware with Berkeley Packet Filter,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems</source>
          ,
          <volume>3550</volume>
          ,
          <year>2023</year>
          ,
          <fpage>95</fpage>
          -
          <lpage>106</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Brydinskyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Khoma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sabodashko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Podpora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Khoma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Konovalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kostiak</surname>
          </string-name>
          ,
          <article-title>Comparison of Modern Deep Learning Models for Speaker Verification</article-title>
          ,
          <source>Applied Sciences (Switzerland)</source>
          <volume>14</volume>
          (
          <issue>4</issue>
          ) (
          <year>2024</year>
          )
          <fpage>1329</fpage>
          -
          <lpage>1</lpage>
          -
          <fpage>1329</fpage>
          -12.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>O.</given-names>
            <surname>Deineka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Harasymchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Partyka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Obshta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Korshun</surname>
          </string-name>
          ,
          <article-title>Designing Data Classification and Secure Store Policy According to SOC 2 Type II</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems</source>
          ,
          <volume>3654</volume>
          ,
          <year>2024</year>
          ,
          <fpage>398</fpage>
          -
          <lpage>409</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>O.</given-names>
            <surname>Deineka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Harasymchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Partyka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Obshta</surname>
          </string-name>
          ,
          <article-title>Application of LLM for Assessing the Effectiveness and Potential Risks of the Information Classification System According to SOC 2 Type II</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems</source>
          ,
          <volume>3991</volume>
          ,
          <year>2025</year>
          ,
          <fpage>215</fpage>
          -
          <lpage>232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>V.</given-names>
             
            <surname>Chouksey</surname>
          </string-name>
          ,
          <source>Automatic Selection of Outlier Detection Techniques, Master's Thesis</source>
          , Eindhoven University of Technology, Netherlands,
          <year>2018</year>
          . https://pure.tue.nl/ws/portalfiles/ portal/109406381/CSE663_Vishal_Chouksey_31_aug.pdf
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>B.</given-names>
             
            <surname>Lindauer</surname>
          </string-name>
          , Insider Threat Test Dataset,
          <year>2020</year>
          . https://kilthub.cmu.edu/articles/dataset/ Insider_Threat_Test_Dataset/12841247
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
             
            <surname>Abdin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
             
            <surname>Zhou</surname>
          </string-name>
          , Phi-3
          <source>Technical Report: A Highly Capable Language Model Locally on Your Phone</source>
          ,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2404.14219
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>W.</given-names>
             
            <surname>Kwon</surname>
          </string-name>
          , et al.,
          <source>Efficient Memory Management for Large Language Model Serving with Pagedattention</source>
          ,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2309.06180
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Google</surname>
            ,
            <given-names>Google</given-names>
          </string-name>
          <string-name>
            <surname>Colaboratory</surname>
          </string-name>
          ,
          <year>2025</year>
          . https://colab.google/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>