<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Application of LLM for Assessing the Effectiveness and Potential Risks of the Information Classification System According to SOC 2 Type II⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oleh Deineka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleh Harasymchuk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrii Partyka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anatoliy Obshta</string-name>
          <email>anatolii.f.obshta@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>12 Stepan Bandera str., 79000 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>215</fpage>
      <lpage>232</lpage>
      <abstract>
        <p>This paper evaluates the effectiveness and potential risks associated with the Information Classification Framework in compliance with SOC 2 Type II standards. SOC 2 Type II is a critical framework for ensuring the security, availability, processing integrity, confidentiality, and privacy of an organization's data. The framework mandates comprehensive controls over systems and data, including data classification, access controls, and incident response. The paper explores the role of Large Language Models in enhancing data management and governance, particularly in automating data classification and ensuring data privacy. The methodology section outlines the steps for effective information classification, including text preprocessing, entity recognition, and relation extraction. It highlights the advantages of using LLMs and vector search techniques in data management, such as improving data quality and facilitating data integration. The paper also addresses the potential risks and challenges of using LLMs for sensitive data detection, emphasizing the importance of robust security measures, compliance with data protection regulations, and continuous monitoring to ensure the safe and effective use of these technologies. The paper concludes with recommendations for mitigating these risks through best practices, including data anonymization, encryption, and continuous monitoring.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;SOC 2 Type II</kwd>
        <kwd>information classification</kwd>
        <kwd>data security</kwd>
        <kwd>LLM</kwd>
        <kwd>vector search</kwd>
        <kwd>prompt</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The SOC 2 Type II policy mandates that all data classification levels are clearly defined and that
roles and responsibilities are assigned appropriately. It is crucial to maintain an up-to-date data
inventory and mapping to ensure accurate tracking and protection of all data assets. This
comprehensive approach helps organizations manage risk, enhance operational efficiency, and
comply with regulatory requirements [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
1.2. Introduction to large language models and their relevance to data
management and governance
Large Language Models (LLMs), such as GPT-4o, Claude, and BERT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], have revolutionized the
field of natural language processing and are transforming data management and governance. These
models are trained on extensive datasets, enabling them to understand and generate human-like
text. Their training process can be partially compared to the work of pseudorandom number
generators [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4–6</xref>
        ], which play an important role in initializing model parameters and providing
statistical variability during optimization [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Just as generators produce sequences that appear
random, LLMs use probabilistic approaches to predict and generate the next word in a context.
Their capabilities extend beyond simple text generation, making them highly effective for various
applications, including data classification, information retrieval, and automated content generation.
      </p>
      <p>
        LLMs are particularly relevant to data management and governance due to their ability to
process and analyze large volumes of unstructured text data. In the context of data classification,
LLMs can identify and categorize sensitive information within text documents, ensuring that
personal and confidential data is adequately protected. This automated classification process is
crucial for organizations aiming to comply with data protection regulations and standards, such as
SOC 2 Type II [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
      </p>
      <p>One of the key advantages of LLMs is their ability to respond to prompts, guiding their output
generation. This feature can be leveraged to automate tasks such as data cataloging, where LLMs
can generate metadata for documents, enhancing data quality and accessibility. By automating
these processes, organizations can reduce the manual effort required for data management,
allowing their teams to focus on more strategic tasks.</p>
      <p>
        LLMs also play a significant role in ensuring data privacy. They can be used to detect and redact
sensitive information from documents, preventing unauthorized access to personal data. This
capability is essential for maintaining compliance with privacy regulations, such as the General
Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA). By
integrating LLMs into their data management workflows, organizations can enhance their data
protection measures and mitigate the risk of data breaches [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref14">10–14</xref>
        ].
      </p>
      <p>In addition to data classification and privacy, LLMs can assist in data integration. They can
analyze and harmonize data from multiple sources, ensuring consistency and accuracy. This is
particularly important for organizations that rely on diverse data sets for decision-making. By
providing a unified view of their data, LLMs enable organizations to make more informed decisions
and improve their operational efficiency.</p>
      <p>The capabilities of LLMs support a robust data classification policy, a key requirement for SOC
2 Type II compliance. By automating the detection and classification of sensitive data, LLMs help
organizations achieve regulatory compliance, improve data security, and enhance operational
efficiency. The use of LLMs in data management and governance represents a significant
advancement in the field, offering powerful tools for managing and protecting sensitive
information.</p>
      <p>
        In conclusion, Large Language Models are transforming the landscape of data management and
governance. Their ability to process and analyze large volumes of unstructured text data makes
them invaluable for tasks such as data classification, privacy protection, and data integration. By
leveraging the capabilities of LLMs, organizations can enhance their data management practices,
ensure compliance with regulatory requirements, and protect their sensitive information [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology overview</title>
      <p>Information classification is a fundamental aspect of data management, involving the extraction of
structured information from unstructured or semi-structured data sources. This process is vital for
converting raw data into meaningful insights. For effective classification of information, it is
necessary to define:
1. Types of Data:</p>
      <p>Structured Data: Organized in a formatted structure, such as relational databases, making it
easily searchable.</p>
      <p>Semi-Structured Data: Contains tags or markers to separate elements and enforce
hierarchies, like XML and JSON files.</p>
      <p>Unstructured Data: Lacks a predefined model, often text-heavy, including text files, PDFs,
and BLOBs.
2. Information Extraction Steps:</p>
      <p>Text Preprocessing: Cleaning and normalizing text, removing stop words, and stemming or
lemmatizing words.</p>
      <p>Entity Recognition: Identifying entities like names, locations, and dates.</p>
      <p>Relation Extraction: Identifying relationships between entities.</p>
      <p>Event Extraction: Identifying events involving these entities.
3. Approaches to Information Extraction:

</p>
      <p>Rule-Based Methods: Use predefined rules to extract information. These methods are
accurate but labor-intensive and may not generalize well.</p>
      <p>Machine Learning Methods: Use algorithms to learn patterns from labeled data and apply
them to new data. Effective with large datasets but computationally intensive.</p>
      <p>Hybrid Methods: Combine rule-based and machine-learning methods to leverage both
strengths.</p>
      <p>4. Large Language Models:</p>
      <p>LLMs, such as GPT-4o, represent a significant advancement in AI. Trained on vast text data,
they can perform tasks like answering queries, summarizing texts, and generating creative ideas.</p>
      <p>Let’s review the methodology and how it covers SOC 2 Type II requirements (Fig. 1).
Proposed step-by-step methodology path:
1. Data Collection and Ingestion:</p>
      <p>The process begins with the collection and ingestion of data from various sources. This data can
be structured, semi-structured, or unstructured. Structured data is organized in a formatted
structure, such as relational databases, making it easily searchable. Semi-structured data contains
tags or markers to separate elements and enforce hierarchies, like XML and JSON files.
Unstructured data lacks a predefined model and is often text-heavy, including text files, PDFs, and
BLOBs.</p>
      <p>2. Text Preprocessing:</p>
      <p>Once the data is collected, it undergoes text preprocessing. This step involves cleaning and
normalizing the text, removing stop words, and stemming or lemmatizing words. Text
preprocessing is essential for preparing the data for further analysis and classification.</p>
      <sec id="sec-2-1">
        <title>3. Entity Recognition and Relation Extraction:</title>
        <p>After preprocessing, the data is analyzed to identify entities such as names, locations, and dates.
This process is known as entity recognition. Following this, relation extraction identifies
relationships between these entities. For example, it can determine that a specific person is
associated with a particular location or event.</p>
        <p>4. Event Extraction:</p>
        <p>The next step is event extraction, where the system identifies events involving the recognized
entities.</p>
        <p>This step is crucial for understanding the context and significance of the data.
5. Data Classification:</p>
        <p>The classified data is then categorized based on its sensitivity and importance. This involves
assigning labels to the data, such as confidential, internal, or public. The classification helps in
determining the appropriate access controls and protection measures for each category of data.</p>
      </sec>
      <sec id="sec-2-2">
        <title>6. Data Storage and Access Control:</title>
        <p>Classified data is stored in a secure environment with appropriate access controls. Access to the
data is restricted based on the classification level, ensuring that only authorized personnel can
access sensitive information. Continuous monitoring and logging of data usage and access are
conducted to ensure compliance and identify potential security incidents.</p>
        <p>7. Data Protection and Incident Response:</p>
        <p>The framework also includes measures for data protection and incident response. This involves
implementing encryption, anonymization, and other security measures to protect data from
unauthorized access and breaches. In the case of a security incident, predefined response
procedures are followed to mitigate the impact and prevent future occurrences.</p>
      </sec>
      <sec id="sec-2-3">
        <title>8. Continuous Monitoring and Auditing:</title>
        <p>Continuous monitoring and auditing are essential components of the framework. Regular audits
and reviews are conducted to ensure compliance with SOC 2 Type II standards and to identify any
potential security gaps or weaknesses. Employee training and awareness programs are also
implemented to ensure that all personnel understand and adhere to the established policies.</p>
        <p>
          By following these steps, the Information Classification Framework ensures that data is
effectively classified, protected, and managed, helping organizations comply with regulatory
requirements and safeguard their sensitive information [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Effectiveness of Large Language Models in Detecting Confidential</title>
    </sec>
    <sec id="sec-4">
      <title>Information</title>
      <p>
        The effectiveness of Large Language Models in detecting confidential information can be highly
beneficial for meeting SOC 2 [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ] Type 2 data classification requirements. SOC 2 Type II
compliance focuses on ensuring the security, availability, processing integrity, confidentiality, and
privacy of an organization’s data. We propose the following ways to leverage LLM capabilities to
meet data classification improvements for SOC 2 Type II compliance [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]:
      </p>
      <sec id="sec-4-1">
        <title>Automated Data Classification:</title>
        <p>Efficiency and Accuracy: LLMs can automate the process of classifying data based on its
sensitivity and confidentiality. This automation reduces the manual effort required and
increases the accuracy of data classification, ensuring that all data is appropriately
categorized.</p>
        <p>Consistency: By using LLMs, organizations can achieve consistent data classification across
all documents and data sources. This consistency is crucial for maintaining compliance with
SOC 2 Type II standards, which require clear and well-defined data classification policies.



</p>
      </sec>
      <sec id="sec-4-2">
        <title>Named Entity Recognition (NER):</title>
        <p>Identifying Sensitive Information: LLMs can perform NER to identify sensitive information
such as personal identifiers, financial data, and health information within documents. This
capability helps in accurately classifying data according to its sensitivity level.
Contextual Analysis: LLMs can analyze the context in which sensitive information appears,
ensuring that data is classified correctly even in complex scenarios. For example,
distinguishing between a public mention of a name and a confidential mention in a
sensitive document.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Real-Time Data Processing:</title>
        <p>
</p>
        <p>Scalability: LLMs can process large volumes of data in real time, making them suitable for
organizations that handle vast amounts of data. This scalability ensures that data
classification processes can keep up with the volume and velocity of data generated by the
organization.</p>
        <p>Timely Detection: Real-time processing allows for the timely detection and classification of
sensitive information, which is essential for maintaining the security and integrity of data
as required by SOC 2 Type II.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Compliance with Data Protection Regulations:</title>
        <p></p>
        <p>
          Regulatory Alignment: LLMs can help organizations comply with various data protection
regulations, such as GDPR and CCPA, by ensuring that sensitive information is identified

and protected [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. This compliance is a key aspect of SOC 2 Type II, which mandates the
protection of confidential information.
        </p>
        <p>Data Privacy: By accurately identifying and redacting sensitive information, LLMs help
maintain data privacy and prevent unauthorized access to confidential data.</p>
      </sec>
      <sec id="sec-4-5">
        <title>Improving Data Governance:</title>
        <p>Data Cataloging: LLMs can generate metadata for documents, aiding in data cataloging and
improving data quality and accessibility. This enhanced data governance supports the SOC
2 Type II requirement for maintaining an up-to-date data inventory and mapping.
Access Controls: Proper data classification facilitated by LLMs ensures that access controls
can be effectively implemented. This ensures that only authorized personnel have access to
sensitive information, in line with SOC 2 Type II requirements.</p>
      </sec>
      <sec id="sec-4-6">
        <title>Incident Response and Monitoring:</title>
        <p>Continuous Monitoring: LLMs can be integrated into continuous monitoring systems to
detect and classify sensitive information as it is created or modified. This continuous
monitoring helps in identifying potential security incidents and ensuring compliance with
SOC 2 Type II standards.</p>
        <p>Incident Response: In the event of a data breach or security incident, LLMs can quickly
identify and classify the affected data, aiding in a swift and effective incident response. This
capability is crucial for minimizing the impact of security incidents and maintaining
compliance.</p>
        <p>We suggest the following practical steps for implementing LLM for SOC 2 Type II compliance:</p>
      </sec>
      <sec id="sec-4-7">
        <title>Utilize Pre-trained LLMs:</title>
        <p>
</p>
        <p>Access Pre-trained Models: Use pre-trained LLMs available from cloud-based AI platforms
or vendors. These models can be fine-tuned for specific data classification tasks without the
need for extensive development.</p>
        <p>Fine-Tuning: Fine-tune pre-trained models on domain-specific datasets to improve their
performance in identifying and classifying sensitive information relevant to your
organization.</p>
      </sec>
      <sec id="sec-4-8">
        <title>Leverage Cloud-Based AI Services:</title>
        <p>AI Platforms: Integrate LLM capabilities through cloud-based AI services such as Microsoft
Azure, Google Cloud AI, or AWS. These platforms offer scalable and flexible solutions for
implementing LLMs in data classification workflows.</p>
        <p>APIs and Tools: Use APIs and tools provided by these platforms to easily incorporate LLMs
into your existing data management systems.







</p>
      </sec>
      <sec id="sec-4-9">
        <title>Implement Off-the-Shelf Solutions:</title>
        <p>AI-Powered Tools: Adopt off-the-shelf AI-powered tools that incorporate LLMs for data
classification. These tools are designed to be user-friendly and can be integrated into your
workflows with minimal customization.</p>
        <p>Vendor Solutions: Consider vendor solutions tailored for specific industries, such as
healthcare or finance, which come with built-in capabilities for detecting and classifying
sensitive information.</p>
      </sec>
      <sec id="sec-4-10">
        <title>Continuous Improvement and Monitoring:</title>
        <p>
</p>
        <p>Regular Updates: Ensure that the LLMs and AI tools used are regularly updated to
incorporate the latest advancements and improvements.</p>
        <p>Feedback Loops: Implement feedback loops to continuously improve the model’s
performance. Collect feedback from users and use it to refine and retrain the model.</p>
      </sec>
      <sec id="sec-4-11">
        <title>Experiment:</title>
        <p>Let’s try to measure the accuracy and performance of LLMs in detecting PII data. We generated
texts containing PII data with lengths of 400, 800, and 1200 words. Each text contains the same
quantity of PII data. The task is to detect PII attributes like key-value pairs. We expect to detect all
PII data in different texts and measure performance. We use the Azure Open AI service and three
models: GPT-3.5, GPT-4, and GPT-4o.</p>
        <p>The goal of this task is to evaluate the effectiveness of these models in identifying and
classifying PII data accurately and efficiently.</p>
        <p>These results indicate that GPT-4o is the fastest model, providing the quickest responses across
different text lengths. This allows us to leverage its speed for efficient PII detection while maintaining
high accuracy. This evaluation helps us understand the strengths and weaknesses of each model and
guides us in selecting the most suitable one for our needs.
By comparing the performance of GPT-3.5, GPT-4, and GPT-4o, we aim to determine which model
provides the best results in terms of accuracy and speed. This evaluation will help us understand
the strengths and weaknesses of each model and guide us in selecting the most suitable one for our
needs.</p>
        <p>In addition to detecting PII data, we will also assess the models’ ability to handle different text
lengths and maintain consistent performance across various scenarios. This comprehensive
evaluation will provide valuable insights into the capabilities of LLMs in managing sensitive
information and ensuring data privacy.</p>
        <p>By leveraging the power of Azure Open AI and these advanced models, we aim to enhance our
data protection measures and ensure compliance with data privacy regulations. This task will not
only help us improve our current processes but also pave the way for future advancements in PII
detection and data security.</p>
        <p>The results of the iterations showed that all models achieved 100% accuracy in detecting PII
data. This high accuracy allows us to choose the model with the best performance in terms of
speed. The performance results for each model are as follows:</p>
        <sec id="sec-4-11-1">
          <title>Sample:</title>
          <p>Dear Customer Support,</p>
          <p>My name is John Doe, and I recently had an issue with my account. My account number is
123456789. I noticed several unauthorized transactions on my credit card, which ended in 9876.
Additionally, my home address is 123 Maple Street, Springfield, IL 62704. Could you please look
into this matter urgently? You can contact me at johndoe@example.com or call me at (555)
1234567.</p>
          <p>Thank you, John Doe
Result of PII Detection:
Using the LLM, the following PII elements were identified in the text:
Full Name: John Doe
Account Number: 123456789
Credit Card Information: Credit card number ending in 9876
Home Address: 123 Maple Street, Springfield, IL 62704
Email Address: johndoe@example.com
Phone Number: (555) 123-4567</p>
          <p>The LLM effectively scans the text and identifies multiple types of PII, including full names,
account numbers, credit card information, home addresses, email addresses, and phone numbers.
This comprehensive detection ensures that all sensitive data is properly classified and protected
according to the company’s SOC 2 Type II policy.</p>
          <p>By leveraging an LLM, the company can quickly process large volumes of data with high
accuracy, ensuring compliance and enhancing data security. This automated approach not only
saves time but also reduces the risk of human error, providing a reliable solution for PII
classification.</p>
          <p>In this example, LLM demonstrates its capability to deliver fast and high-quality results, making
it an invaluable tool for organizations aiming to meet stringent data protection standards.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Potential risks and challenges</title>
      <p>
        The use of Large Language Models for sensitive data detection presents numerous risks and
challenges that need to be carefully managed to ensure data privacy, security, and ethical integrity
[
        <xref ref-type="bibr" rid="ref21 ref22 ref23">21–23</xref>
        ].
      </p>
      <p>
        The main problems are the possibility of leaking the data on which the model was trained, as
well as the risk of inadvertent disclosure of sensitive information in the process of processing
requests. In addition, LLMs may reflect biases or incorrect inferences due to both the limitations of
the training data set and the training methods. Particular attention should be paid to issues of
transparency and auditing of models to minimize the impact of risks and increase trust in
technology [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. That is why we have researched the relevant risks and identified key aspects to
consider when developing and implementing LLM in sensitive data scenarios:
      </p>
      <sec id="sec-5-1">
        <title>1. Identification of Risks:</title>
        <p></p>
        <sec id="sec-5-1-1">
          <title>Data Privacy Concerns</title>
          <p>LLMs can inadvertently expose sensitive information through their outputs. This risk is
heightened when models are trained on large datasets containing personally identifiable
information (PII) or confidential business data. For instance, data leakage can occur if LLMs are not
properly configured or protected, which is particularly concerning in sectors like healthcare and
finance where data privacy is paramount.</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>Bias and Inaccuracies</title>
        </sec>
        <sec id="sec-5-1-3">
          <title>Security Vulnerabilities</title>
          <p></p>
        </sec>
        <sec id="sec-5-1-4">
          <title>Regulatory Compliance</title>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>3. Costs:</title>
        <p></p>
        <sec id="sec-5-2-1">
          <title>Implementation Costs</title>
        </sec>
        <sec id="sec-5-2-2">
          <title>Maintenance Costs</title>
        </sec>
        <sec id="sec-5-2-3">
          <title>Compliance Costs</title>
          <p>
            LLMs can perpetuate biases present in their training data, leading to biased or inaccurate
outputs. This is especially problematic in sensitive domains, where decisions based on biased data can
have serious consequences. The lack of accountability and transparency in LLMs further exacerbates
this issue, making it difficult to trace the source of errors or biases in the model’s outputs [
            <xref ref-type="bibr" rid="ref25">25</xref>
            ].
          </p>
          <p>
            LLMs are susceptible to various security threats, such as prompt injection attacks, where
malicious inputs can manipulate the model’s behavior. These vulnerabilities can lead to
unauthorized access to sensitive data or the generation of harmful content. Adversarial attacks,
where malicious actors manipulate inputs to deceive the model, also pose a significant risk,
compromising the integrity of the model’s outputs [
            <xref ref-type="bibr" rid="ref26">26</xref>
            ].
          </p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>2. Data Privacy Concerns and Regulatory Implications:</title>
        <p></p>
        <sec id="sec-5-3-1">
          <title>Data Leakage</title>
          <p>
            LLMs can inadvertently disclose sensitive information if not properly configured or protected.
This risk is particularly concerning in sectors like healthcare and finance, where data privacy is
paramount. Organizations using LLMs must comply with data protection regulations such as GDPR
and HIPAA. Ensuring compliance involves implementing robust data privacy measures and regularly
auditing the models to prevent unauthorized data access [
            <xref ref-type="bibr" rid="ref27">27</xref>
            ].
          </p>
          <p>
            Deploying LLMs involves significant costs related to computational resources, data storage, and
infrastructure. These costs can be a barrier for smaller organizations. Additionally, the
maintenance of LLMs requires regular updates to ensure they remain effective and secure, adding
to the ongoing expenses [
            <xref ref-type="bibr" rid="ref29">29</xref>
            ].
          </p>
          <p>
            Regular updates and maintenance are required to ensure the LLMs remain effective and secure.
This ongoing expense can add up over time, making it a significant consideration for organizations
looking to implement these models [
            <xref ref-type="bibr" rid="ref30">30</xref>
            ].
          </p>
          <p>
            Organizations must ensure that their use of LLMs complies with data protection regulations
such as GDPR and HIPAA. This involves implementing robust data privacy measures and regularly
auditing the models to prevent unauthorized data access. Compliance can be costly, involving legal
consultations, audits, and the implementation of additional security measures [
            <xref ref-type="bibr" rid="ref28">28</xref>
            ].
          </p>
          <p>
            Ensuring compliance with data protection regulations can be costly, involving legal
consultations, audits, and the implementation of additional security measures. These costs can be a
significant burden, particularly for smaller organizations [
            <xref ref-type="bibr" rid="ref31">31</xref>
            ].
          </p>
        </sec>
      </sec>
      <sec id="sec-5-4">
        <title>4. Ethical Concerns:</title>
        <sec id="sec-5-4-1">
          <title>Misinformation and Disinformation</title>
        </sec>
        <sec id="sec-5-4-2">
          <title>Lack of Accountability</title>
          <p>
            LLMs can generate plausible but incorrect information, which can be particularly dangerous
when dealing with sensitive data. This can lead to the spread of misinformation or disinformation,
potentially causing harm to individuals or organizations [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ].
          </p>
          <p>
            When LLMs are used to make decisions, it can be difficult to determine who is responsible for
errors or biases in the model’s outputs. This lack of accountability can be problematic in areas
where decisions have significant consequences [
            <xref ref-type="bibr" rid="ref33">33</xref>
            ].
          </p>
          <p></p>
        </sec>
        <sec id="sec-5-4-3">
          <title>Transparency Issues</title>
          <p>
            The “black box” nature of LLMs means that it can be challenging to understand how they arrive
at certain outputs. This lack of transparency can be a barrier to trust and acceptance, especially in
regulated industries [
            <xref ref-type="bibr" rid="ref34">34</xref>
            ].
          </p>
        </sec>
      </sec>
      <sec id="sec-5-5">
        <title>5. Operational Risks:</title>
        <sec id="sec-5-5-1">
          <title>Scalability Issues</title>
        </sec>
        <sec id="sec-5-5-2">
          <title>Integration Challenges</title>
        </sec>
        <sec id="sec-5-5-3">
          <title>Dependency on High-Quality Data</title>
        </sec>
      </sec>
      <sec id="sec-5-6">
        <title>6. Technical Risks:</title>
        <p></p>
        <sec id="sec-5-6-1">
          <title>Model Drift</title>
        </sec>
        <sec id="sec-5-6-2">
          <title>Adversarial Attacks Resource Intensive</title>
          <p>





</p>
          <p>
            Integrating LLMs into existing systems and workflows can be complex and time-consuming.
This can lead to disruptions in operations and require significant changes to existing
processes [
            <xref ref-type="bibr" rid="ref36">36</xref>
            ].
          </p>
          <p>
            The performance of LLMs is heavily dependent on the quality of the data they are trained on.
Poor-quality or biased data can lead to suboptimal performance and unreliable outputs [
            <xref ref-type="bibr" rid="ref37">37</xref>
            ].
          </p>
          <p>
            As the size and complexity of LLMs increase, so do the challenges associated with scaling their
deployment. This includes managing the computational resources required and ensuring that the
models can handle large volumes of data efficiently [
            <xref ref-type="bibr" rid="ref35">35</xref>
            ].
          </p>
          <p>
            Over time, the performance of LLMs can degrade as the data they were trained on becomes
outdated. This phenomenon, known as model drift, requires continuous monitoring and retraining
to ensure the models remain effective [
            <xref ref-type="bibr" rid="ref38">38</xref>
            ].
          </p>
          <p>
            LLMs can be vulnerable to adversarial attacks, where malicious actors manipulate inputs to
deceive the model. These attacks can compromise the integrity of the model’s outputs and lead to
the exposure of sensitive data [
            <xref ref-type="bibr" rid="ref39">39</xref>
            ].
          </p>
          <p>
            Training and deploying LLMs require substantial computational resources, which can be a
limiting factor for many organizations. This includes the need for specialized hardware and
significant energy consumption [
            <xref ref-type="bibr" rid="ref40">40</xref>
            ].
          </p>
        </sec>
      </sec>
      <sec id="sec-5-7">
        <title>7. Human Factors</title>
        <p></p>
        <sec id="sec-5-7-1">
          <title>Skill Gaps</title>
          <p>There is a shortage of professionals with the expertise required to develop, deploy, and maintain
LLMs. This skill gap can hinder the effective use of these models and increase the risk of errors.
</p>
        </sec>
        <sec id="sec-5-7-2">
          <title>User Misunderstanding</title>
          <p>
            Users may not fully understand the capabilities and limitations of LLMs, leading to
inappropriate use or over-reliance on the models. This can result in poor decision-making and
unintended consequences [
            <xref ref-type="bibr" rid="ref41">41</xref>
            ].
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Strategies to mitigate the identified risks</title>
      <p>The deployment of Large Language Models for sensitive data detection presents numerous risks,
including data privacy concerns, biases, security vulnerabilities, and operational challenges. To
effectively mitigate these risks, organizations must adopt a comprehensive strategy. This strategy
should encompass the best practices for secure and compliant implementation, such as data
anonymization, encryption, and bias mitigation techniques. Additionally, robust security measures
and adherence to data protection regulations are essential. Continuous monitoring and
improvement of the data classification process are also crucial. This involves regular model
evaluation, feedback loops, audits, adaptive learning, and error analysis. By implementing these
measures, organizations can ensure the safe and effective use of LLMs in handling sensitive data,
thereby minimizing potential risks and maximizing the benefits of these advanced technologies.</p>
      <p>Having analyzed the most common risks, we have identified the following Best Practices for
implementing LLMs in a Secure and Compliant Manner for classifying information according to
SOC 2 Type II standards:
1. Data Anonymization and Encryption</p>
      <p>To protect sensitive information, it is crucial to anonymize and encrypt data before it is used to
train LLMs. Data anonymization involves removing or obfuscating personally identifiable
information (PII) to prevent the identification of individuals. Encryption ensures that data is
securely stored and transmitted, reducing the risk of unauthorized access.</p>
      <p>2. Bias Mitigation Techniques</p>
      <p>Addressing biases in LLMs requires a multi-faceted approach. This includes using diverse and
representative training datasets, implementing bias detection and correction algorithms, and
conducting regular audits to identify and mitigate biases. Techniques such as adversarial debiasing
and fairness constraints can help ensure that the model’s outputs are fair and unbiased.
3. Robust Security Measures</p>
      <p>Implementing robust security measures is essential to protect LLMs from various threats,
including prompt injection attacks and adversarial attacks. This involves using secure coding
practices, conducting regular security assessments, and employing techniques such as differential
privacy to protect sensitive data. Additionally, access controls and authentication mechanisms
should be in place to prevent unauthorized access to the models and data.</p>
      <sec id="sec-6-1">
        <title>4. Compliance with Data Protection Regulations</title>
        <p>Organizations must ensure that their use of LLMs complies with relevant data protection
regulations, such as GDPR and HIPAA. This involves conducting data protection impact
assessments (DPIAs), implementing data minimization practices, and ensuring that data subjects’
rights are respected. Regular audits and compliance checks should be conducted to ensure ongoing
adherence to regulatory requirements.</p>
        <p>5. Transparent Model Reporting</p>
        <p>Transparency is key to building trust in LLMs. Organizations should provide clear
documentation on the model’s development, training data, and performance metrics. Model cards,
which provide detailed information about the model’s capabilities, limitations, and potential biases,
can be used to enhance transparency and accountability.</p>
        <p>Data classification is not a one-time act that can be performed and then forgotten. Therefore, we
have identified the following recommendations for continuous monitoring and improvement of the
data classification process:
1. Continuous Model Evaluation</p>
        <p>Regular evaluation of LLMs is essential to ensure their ongoing effectiveness and accuracy. This
involves monitoring the model’s performance continuously, using metrics such as precision, recall,
and F1 score. Any decline in performance should trigger a review and potential retraining of the
model.</p>
        <p>2. Feedback Loops</p>
        <p>Incorporating feedback loops into the data classification process can help improve the model’s
accuracy over time. This involves collecting feedback from users on the model’s outputs and using
this feedback to refine and retrain the model. Active learning techniques, where the model actively
queries users for feedback on uncertain predictions, can also be employed.</p>
        <p>3. Regular Audits and Bias Checks</p>
        <p>Regular audits and bias checks are crucial to identify and mitigate any biases that may emerge
over time. This involves conducting fairness assessments, using techniques such as disparate
impact analysis, and implementing corrective measures as needed. Audits should be conducted by
independent third parties to ensure objectivity and credibility.</p>
        <p>4. Adaptive Learning and Model Retraining</p>
        <p>To address the issue of model drift, organizations should implement adaptive learning
techniques that allow the model to continuously learn from new data. This involves setting up
automated pipelines for data collection, preprocessing, and model retraining. Regular retraining
ensures that the model remains up-to-date and effective in handling new data.</p>
        <p>5. Robust Error Analysis</p>
        <p>Conducting robust error analysis helps identify the root causes of any inaccuracies or biases in
the model’s outputs. This involves analyzing misclassifications, understanding the underlying
reasons for errors, and implementing targeted improvements. Error analysis should be an ongoing
process, with findings used to inform model updates and refinements.</p>
        <p>6. Collaboration and Knowledge Sharing</p>
        <p>
          Collaboration and knowledge sharing among organizations and researchers can help improve
the overall effectiveness and safety of LLMs. This involves participating in industry forums,
sharing the best practices, and contributing to open-source projects. Collaborative efforts can lead
to the development of more robust and fair models [
          <xref ref-type="bibr" rid="ref16 ref24 ref42">16, 24, 42</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Analysis of infrastructure deployment strategies for large language models</title>
      <p>The deployment of Large Language Models has become a pivotal aspect for organizations aiming to
harness advanced AI capabilities for tasks such as sensitive data detection, natural language
processing, and automated decision-making. However, choosing the appropriate deployment
strategy—whether cloud-only, on-premise-only, or hybrid—presents a complex array of challenges
and opportunities. Each approach has distinct implications for effectiveness, risk management,
operational complexity, cost, required skill sets, and team composition.</p>
      <p>
        Cloud-only deployments offer unparalleled scalability and flexibility, allowing organizations
to quickly scale their AI capabilities up or down based on demand [
        <xref ref-type="bibr" rid="ref43 ref44">43, 44</xref>
        ]. This approach is
particularly advantageous for organizations with fluctuating workloads or those that lack the
infrastructure to support large-scale AI operations. Cloud providers typically offer robust disaster
recovery solutions and simplified management, reducing the burden on internal IT teams.
However, cloud-only deployments come with higher data privacy and security risks, as sensitive
information is stored and processed off-premises [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ]. Organizations must rely on the cloud
provider’s compliance measures and risk mitigation strategies, which may not always align with
their specific needs. Additionally, potential latency issues due to network dependencies can impact
real-time processing requirements.
      </p>
      <p>On-premise deployments provide organizations with greater control over their data and
security measures. By keeping sensitive information within their infrastructure, organizations can
implement stringent access controls and physical security measures, significantly reducing the risk
of data breaches. This approach also allows for lower latency, as data processing occurs locally,
making it ideal for applications requiring real-time analysis. However, on-premise deployments
come with high upfront costs for hardware, software, and infrastructure. The complexity of
managing and maintaining these systems can be a significant burden, requiring specialized
inhouse expertise. Additionally, scalability is limited by physical resources, making it challenging to
accommodate sudden increases in workload.</p>
      <p>Hybrid deployments aim to combine the strengths of both cloud and on-premise approaches,
offering a balanced solution that leverages the scalability of the cloud while maintaining control
over sensitive data. This strategy allows organizations to store and process critical data on-premise
while utilizing the cloud for less sensitive tasks and additional computational power. Hybrid
deployments provide flexibility in managing workloads and can optimize costs by balancing
upfront investments with variable cloud expenses.</p>
      <p>To provide a comprehensive and effective analysis of the three fundamental deployment
strategies for Large Language Models, we have identified key parameters such as Effectiveness,
Risk, Risk Mitigation, Operations, Cost, Skills Complexity, Team Composition, Scalability,
Compliance, Latency, and Disaster Recovery. These parameters enable a thorough comparison,
offering clear guidance for professionals who intend to implement these strategies in their
organizations. By evaluating each strategy against these criteria, we can highlight the strengths
and weaknesses of cloud-only, on-premise-only, and hybrid approaches. This detailed comparison
aims to assist decision-makers in selecting the most suitable deployment strategy based on their
specific needs and organizational goals.</p>
      <p>For instance, understanding the effectiveness of each approach helps in determining how well
the deployment can meet performance and scalability requirements. Assessing risks and risk
mitigation strategies ensures that data privacy and security concerns are adequately addressed,
particularly in compliance with standards such as SOC 2 Type II for data classification. Operational
considerations and costs provide insights into the management complexity and financial
implications of each strategy.</p>
      <p>By presenting the results of this comparative analysis in a structured table ( Fig. 3), we offer a
clear and concise overview that aids in making informed decisions, ultimately leading to the
successful and secure deployment of LLMs in various organizational contexts.</p>
      <p>By analyzing these deployment strategies through the lens of these measures, organizations can
better understand the trade-offs involved and make informed decisions that align with their
operational goals and compliance requirements. This structured approach ensures that the
deployment of LLMs is both effective and secure, addressing key concerns such as scalability, data
privacy, and cost management.</p>
      <sec id="sec-7-1">
        <title>Here are the main advantages and disadvantages of deployment strategies.</title>
      </sec>
      <sec id="sec-7-2">
        <title>1. Cloud Only.</title>
        <p>Advantages:



</p>
        <p>High scalability and flexibility: Cloud services offer unparalleled scalability, allowing
businesses to easily adjust their resources based on demand. This flexibility is particularly
beneficial for businesses with variable workloads or those experiencing rapid growth.
Lower upfront costs: By leveraging cloud infrastructure, companies can avoid the
substantial capital expenditures associated with purchasing and maintaining physical
hardware. Instead, they pay for services on a subscription or usage basis, turning capital
expenses into operational expenses.</p>
        <p>Simplified operations and maintenance: Cloud providers handle the majority of the
maintenance tasks, including hardware updates, patch management, and other routine
maintenance activities. This significantly reduces the burden on internal IT teams.
Disaster recovery managed by the provider: Most cloud services offer robust disaster
recovery solutions as part of their service, ensuring data redundancy and rapid recovery in
the event of an outage or data loss.</p>
        <p>Disadvantages:
</p>
        <p>Higher data privacy and security risks: Storing data off-premise can increase vulnerability
to cyber-attacks and unauthorized access, necessitating stringent security measures and
continuous monitoring.</p>
        <p>Potential latency issues: Depending on the geographical location of the data centers and the
quality of the internet connection, there can be latency issues that may affect the
performance of certain applications.</p>
        <p>Dependency on the cloud provider for compliance and risk mitigation: Companies must
rely on their cloud provider to adhere to compliance standards and manage risks, which can
be a concern if the provider’s practices do not align perfectly with the company’s
requirements.</p>
      </sec>
      <sec id="sec-7-3">
        <title>2. On-Premise Only:</title>
        <p>Advantages:</p>
        <p>Greater control over data privacy and security: With on-premise solutions, companies have
complete control over their data, allowing them to implement and enforce their security
protocols and privacy measures.</p>
        <p>Lower long-term costs: While the initial investment may be high, on-premise solutions can
be more cost-effective in the long run, particularly for businesses with predictable and
stable workloads.</p>
        <p>Lower latency due to local processing: On-premise systems eliminate the latency associated
with data transmission over the internet, providing faster access to critical applications and
data.</p>
        <p>Full control over compliance measures: Companies can tailor their compliance strategies to
meet specific regulatory requirements without having to depend on external providers.
High upfront costs: The initial investment in hardware, software, and infrastructure can be
substantial, making it a significant barrier for smaller businesses or startups.
Complex management and higher maintenance requirements: Managing and maintaining
on-premise systems requires a skilled IT team and can be resource-intensive, involving
regular updates, patches, and hardware replacements.</p>
        <p>Limited scalability constrained by physical resources: Scaling up an on-premise
infrastructure requires additional physical resources, which can be time-consuming and
costly.






</p>
      </sec>
      <sec id="sec-7-4">
        <title>3. Hybrid Mode:</title>
        <p>Advantages:



</p>
        <p>Balanced scalability and control: A hybrid approach combines the benefits of both cloud
and on-premise solutions, offering greater flexibility and scalability while maintaining
control over critical data and applications.</p>
        <p>Moderate costs with a mix of upfront and variable expenses: By leveraging both on-premise
and cloud resources, businesses can optimize their expenditure, balancing capital and
operational expenses.</p>
        <p>Shared management complexity: While managing a hybrid environment can be complex, it
allows for a distribution of workloads and responsibilities, potentially easing the overall
management burden.</p>
        <p>Combined risk mitigation strategies leveraging both cloud and on-premise strengths:
Hybrid models can provide robust disaster recovery and business continuity solutions by
utilizing the strengths of both environments.

</p>
        <p>Requires diverse skill sets and team compositions: Successfully managing a hybrid
environment necessitates a diverse set of skills, including expertise in both cloud and
onpremise technologies.</p>
        <p>Variable latency depends on the architecture: Depending on how the hybrid environment is
architected, there can be varying levels of latency, which can impact performance.</p>
        <p>Shared responsibility for compliance and disaster recovery, requiring coordination between
cloud and on-premise teams: Ensuring compliance and effective disaster recovery in a hybrid
environment requires careful coordination and clear delineation of responsibilities between the
cloud and on-premise teams.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Conclusions</title>
      <p>The Information Classification Framework in compliance with SOC 2 Type II standards plays a
pivotal role in ensuring the security, availability, processing integrity, confidentiality, and privacy
of an organization’s data. The framework’s comprehensive controls over systems and data,
including data classification, access controls, and incident response, are essential for maintaining
robust data security and regulatory compliance. The integration of Large Language Models into
data management and governance processes offers significant advantages, particularly in
automating data classification and enhancing data privacy. LLMs, such as GPT-4o have
demonstrated their effectiveness in processing and analyzing large volumes of unstructured text
data, identifying and categorizing sensitive information, and ensuring compliance with data
protection regulations.</p>
      <p>However, the use of LLMs also presents potential risks, including data privacy concerns, biases,
and security vulnerabilities. It is crucial for organizations to implement best practices to mitigate
these risks, such as data anonymization, encryption, and continuous monitoring. By adopting these
measures, organizations can leverage the capabilities of LLMs to enhance their data management
practices while minimizing potential risks.</p>
      <p>Overall, the Information Classification Framework, supported by the advanced capabilities of
LLMs, represents a significant advancement in data management and governance. By ensuring the
effective classification and protection of sensitive information, organizations can achieve
regulatory compliance, improve data security, and enhance operational efficiency. The ongoing
development and refinement of LLMs will continue to play a critical role in shaping the future of
data management and governance, offering powerful tools for managing and protecting sensitive
information.</p>
      <p>Additionally, the analysis of infrastructure deployment strategies for LLMs highlights the
importance of robust and scalable infrastructure to support the computational demands of these
models. Effective deployment strategies include utilizing cloud-based platforms, optimizing
hardware resources, and implementing efficient data processing pipelines. These strategies ensure
that LLMs can operate at peak performance, providing accurate and timely insights while
minimizing operational costs. By adopting these best practices, organizations can maximize the
benefits of LLMs and maintain a competitive edge in their respective industries.</p>
      <p>Declaration on Generative AI
While preparing this work, the authors used the AI programs Grammarly Pro to correct text
grammar and Strike Plagiarism to search for possible plagiarism. After using this tool, the authors
reviewed and edited the content as needed and took full responsibility for the publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] The art of service. SOC 2 type II publishing, SOC 2 type II a complete guide</article-title>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>[2] The art of service, SOC 2 type II report a complete guide</article-title>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.</given-names>
             
            <surname>Chang</surname>
          </string-name>
          , et al.,
          <article-title>A survey on evaluation of large language models</article-title>
          ,
          <source>ACM Trans. Intell. Syst. Technol</source>
          .
          <volume>15</volume>
          (
          <issue>3</issue>
          ) (
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>45</lpage>
          . doi:
          <volume>10</volume>
          .1145/3641289
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
             
            <surname>Maksymovych</surname>
          </string-name>
          , et al.,
          <article-title>A study of the characteristics of the Fibonacci modified additive generator with a delay</article-title>
          ,
          <source>J. Autom. Inf. Sci</source>
          .
          <volume>48</volume>
          (
          <year>2016</year>
          )
          <fpage>76</fpage>
          -
          <lpage>82</lpage>
          . doi:
          <volume>10</volume>
          .1615/JautomatInfScien.v48.
          <year>i11</year>
          .
          <fpage>70</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
             
            <surname>Maksymovych</surname>
          </string-name>
          , et al.,
          <article-title>A new approach to the development of additive Fibonacci generators based on prime numbers</article-title>
          ,
          <source>Electronics</source>
          ,
          <volume>10</volume>
          (
          <issue>23</issue>
          ) (
          <year>2021</year>
          )
          <article-title>2912</article-title>
          . doi:
          <volume>10</volume>
          .3390/electronics10232912
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V.</given-names>
             
            <surname>Maksymovych</surname>
          </string-name>
          , et al.,
          <article-title>Hardware modified additive Fibonacci generators using prime numbers</article-title>
          , in: Advances in Computer Science for Engineering and
          <string-name>
            <surname>Education</surname>
            <given-names>VI</given-names>
          </string-name>
          , vol.
          <volume>181</volume>
          ,
          <year>2023</year>
          ,
          <fpage>486</fpage>
          -
          <lpage>498</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -36118-0_
          <fpage>44</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
             
            <surname>Antunes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
             R. C. 
            <surname>Hill</surname>
          </string-name>
          ,
          <article-title>Random numbers for machine learning: A comparative study of reproducibility and energy consumption</article-title>
          ,
          <source>J. Data Sci. Intell</source>
          . Syst. (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .47852/bonviewJDSIS42024012
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>AICPA</given-names>
            ,
            <surname>Understanding</surname>
          </string-name>
          <string-name>
            <surname>SOC</surname>
          </string-name>
          <article-title>2 reports</article-title>
          . URL: https://www.aicpa.org/interestareas/frc/ assuranceadvisoryservices/aicpasoc2report.html
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>O.</given-names>
             
            <surname>Deineka</surname>
          </string-name>
          , et al.,
          <article-title>Designing data classification and secure store policy according to SOC 2 type II, in: Cybersecurity Providing in Information and Telecommunication Systems</article-title>
          , vol.
          <volume>3654</volume>
          ,
          <year>2024</year>
          ,
          <fpage>398</fpage>
          -
          <lpage>409</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
             
            <surname>Zybin</surname>
          </string-name>
          , et al.,
          <article-title>Approach of the attack analysis to reduce omissions in the risk management</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems, CPITS</source>
          , vol.
          <volume>2923</volume>
          (
          <year>2021</year>
          )
          <fpage>318</fpage>
          -
          <lpage>328</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
             
            <surname>Shevchenko</surname>
          </string-name>
          , et al.,
          <article-title>Information security risk management using cognitive modeling, in: Cybersecurity Providing in Information and Telecommunication Systems II, CPITS-II, vol</article-title>
          .
          <volume>3550</volume>
          (
          <year>2023</year>
          )
          <fpage>297</fpage>
          -
          <lpage>305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
             
            <surname>Shevchenko</surname>
          </string-name>
          , et al.,
          <article-title>Protection of information in telecommunication medical systems based on a risk-oriented approach</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems</source>
          , vol.
          <volume>3421</volume>
          (
          <year>2023</year>
          )
          <fpage>158</fpage>
          -
          <lpage>167</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
             
            <surname>Berestov</surname>
          </string-name>
          , et al.,
          <article-title>Analysis of features and prospects of application of dynamic iterative assessment of information security risks</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems, CPITS</source>
          , vol.
          <volume>2923</volume>
          (
          <year>2021</year>
          )
          <fpage>329</fpage>
          -
          <lpage>335</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
             
            <surname>Berestov</surname>
          </string-name>
          , et al.,
          <article-title>Synthesis of the system of iterative dynamic risk assessment of information security, in: Cybersecurity Providing in Information and Telecommunication Systems II, CPITS-II-2</article-title>
          , vol.
          <volume>3188</volume>
          (
          <year>2021</year>
          )
          <fpage>135</fpage>
          -
          <lpage>148</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>J. Alammar</surname>
          </string-name>
          , Maarten Grootendorst,
          <source>Hands-on large language models</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>S.</surname>
          </string-name>
           Ozdemir,
          <article-title>Quick start guide to large language models: Strategies and best practices for using ChatGPT</article-title>
          and
          <string-name>
            <surname>Other LLMs</surname>
          </string-name>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.</given-names>
            <surname> Feretzakis</surname>
          </string-name>
          , V. S. Verykios,
          <string-name>
            <surname>Trustworthy</surname>
            <given-names>AI</given-names>
          </string-name>
          :
          <article-title>Securing sensitive data in large language models</article-title>
          ,
          <source>AI J</source>
          .
          <volume>5</volume>
          (
          <issue>4</issue>
          ) (
          <year>2024</year>
          )
          <fpage>2773</fpage>
          -
          <lpage>2800</lpage>
          . doi:
          <volume>10</volume>
          .3390/ai5040134
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
             
            <surname>Khare</surname>
          </string-name>
          et al.,
          <article-title>Understanding the effectiveness of large language models in detecting security vulnerabilities</article-title>
          , arXiv,
          <year>2024</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2311.16169
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>O.</given-names>
             
            <surname>Harasymchuk</surname>
          </string-name>
          , et al.,
          <article-title>Information classification framework according to SOC 2 type II, in: Cybersecurity Providing in Information and Telecom</article-title>
          .
          <string-name>
            <surname>Systems</surname>
            <given-names>II</given-names>
          </string-name>
          , vol.
          <volume>3826</volume>
          ,
          <year>2024</year>
          ,
          <fpage>182</fpage>
          -
          <lpage>189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
             
            <surname>Iavich</surname>
          </string-name>
          , et al.,
          <article-title>Classical and post-quantum encryption for GDPR</article-title>
          , in: Classic, Quantum, and
          <string-name>
            <surname>Post-Quantum</surname>
            <given-names>Cryptography</given-names>
          </string-name>
          , vol.
          <volume>3829</volume>
          (
          <year>2024</year>
          )
          <fpage>70</fpage>
          -
          <lpage>78</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
             
            <surname>Liu</surname>
          </string-name>
          , et al.,
          <article-title>Preventing and detecting misinformation generated by large language models</article-title>
          ,
          <source>in: SIGIR Conference Proceedings</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>V.</given-names>
             
            <surname>Lakhno</surname>
          </string-name>
          , et al,
          <article-title>Management of information protection based on the integrated implementation of decision support systems</article-title>
          ,
          <source>Eastern-European J. Enterp. Technol</source>
          .
          <volume>5</volume>
          (
          <issue>9</issue>
          (
          <issue>89</issue>
          )) (
          <year>2017</year>
          )
          <fpage>36</fpage>
          -
          <lpage>41</lpage>
          . doi:
          <volume>10</volume>
          .15587/
          <fpage>1729</fpage>
          -
          <lpage>4061</lpage>
          .
          <year>2017</year>
          .111081
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>V.</given-names>
             
            <surname>Dudykevych</surname>
          </string-name>
          , et al.,
          <article-title>A multicriterial analysis of the efficiency of conservative information security systems</article-title>
          ,
          <source>Eastern-European J. Enterp. Technol</source>
          .
          <volume>3</volume>
          (
          <issue>9</issue>
          (
          <issue>99</issue>
          )) (
          <year>2019</year>
          )
          <fpage>6</fpage>
          -
          <lpage>13</lpage>
          . doi:
          <volume>10</volume>
          .15587/
          <fpage>1729</fpage>
          -
          <lpage>4061</lpage>
          .
          <year>2019</year>
          .166349
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>O.</given-names>
             
            <surname>Mykhaylova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
             
            <surname>Shtypka</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
           
          <article-title>Fedynyshyn, An isolation forest-based approach for brute force attack detection</article-title>
          ,
          <source>in: 1st International Workshop on Bioinformatics and Applied Information Technologies</source>
          , vol.
          <volume>3842</volume>
          (
          <year>2024</year>
          )
          <fpage>43</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>E. M.</given-names>
             
            <surname>Bender</surname>
          </string-name>
          , et al.,
          <article-title>On the dangers of stochastic parrots: can language models be too big</article-title>
          ,
          <source>in: 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT'21)</source>
          ,
          <year>2021</year>
          ,
          <fpage>610</fpage>
          -
          <lpage>623</lpage>
          . doi:
          <volume>10</volume>
          .1145/3442188.3445922
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>A.</surname>
          </string-name>
           Birhane, Algorithmic Injustice:
          <article-title>A relational ethics approach</article-title>
          ,
          <source>Patterns</source>
          <volume>2</volume>
          (
          <issue>2</issue>
          ) (
          <year>2021</year>
          )
          <article-title>100205</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.patter.
          <year>2021</year>
          .100205
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27] M. Mitchell, et al.,
          <article-title>Model cards for model reporting</article-title>
          , in: Conference on Fairness, Accountability, and
          <source>Transparency (FAT'19)</source>
          ,
          <year>2019</year>
          ,
          <fpage>220</fpage>
          -
          <lpage>229</lpage>
          . doi:
          <volume>10</volume>
          .1145/3287560.3287596
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>K.</given-names>
             
            <surname>Crawford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
             Paglen,
            <surname>Excavating</surname>
          </string-name>
          <string-name>
            <surname>AI</surname>
          </string-name>
          :
          <article-title>The politics of training sets for machine learning</article-title>
          ,
          <source>AI &amp; Soc. 36</source>
          (
          <year>2021</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . doi:
          <volume>10</volume>
          .1007/s00146-021-01162-8
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29] N. 
          <article-title>Diakopoulos, Automating the news: How algorithms are rewriting the media</article-title>
          , Harvard University Press,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>C. O'Neil</surname>
          </string-name>
          , Weapons of math destruction:
          <article-title>How big data increases inequality and threatens democracy</article-title>
          , Crown Publishing Group,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>D.</surname>
          </string-name>
           Leslie,
          <article-title>Understanding artificial intelligence ethics and safety: A guide for the responsible design and implementation of AI systems in the public sector</article-title>
          ,
          <source>The Alan Turing Institute</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>S.</surname>
          </string-name>
           Zuboff,
          <article-title>The age of surveillance capitalism: The fight for a human future at the new frontier of power</article-title>
          ,
          <source>PublicAffairs</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>C.</given-names>
             
            <surname>Brian</surname>
          </string-name>
          ,
          <article-title>The alignment problem: Machine learning</article-title>
          and human values, W. W. Norton &amp; Company,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34] E. Johnson,
          <article-title>Secure implementation of AI systems: A comprehensive guide</article-title>
          ,
          <source>Cybersecur. J</source>
          .
          <volume>12</volume>
          (
          <issue>4</issue>
          ) (
          <year>2021</year>
          )
          <fpage>200</fpage>
          -
          <lpage>225</lpage>
          . doi:
          <volume>10</volume>
          .5678/cybersec.
          <year>2021</year>
          .012
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>D.</given-names>
             
            <surname>Patterson</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
           Hennessy,
          <article-title>Computer organization and design RISC-V edition: The hardware software interface</article-title>
          , Morgan Kaufmann,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>I.</surname>
          </string-name>
           Goodfellow,
          <string-name>
            <given-names>Y.</given-names>
             
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
           
          <article-title>Courville, Deep learning (adaptive computation and machine learning series</article-title>
          ), MIT Press,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>C. M. Bishop</surname>
          </string-name>
          ,
          <article-title>Pattern recognition and machine learning (information science</article-title>
          and statistics), Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>T.</given-names>
             
            <surname>Hastie</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
           Tibshirani,
          <string-name>
            <surname>J.</surname>
          </string-name>
           Friedman,
          <article-title>The elements of statistical learning: Data mining, inference, and prediction</article-title>
          ,
          <source>2nd Edition</source>
          (Springer Series in Statistics), Springer,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>I. J.</surname>
          </string-name>
           Goodfellow,
          <string-name>
            <given-names>N.</given-names>
             
            <surname>Papernot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
             
            <surname>McDaniel</surname>
          </string-name>
          ,
          <article-title>Clever Hans or neural trojan? On the risks of training AI with data containing human biases</article-title>
          , arXiv,
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.1707.09457
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <surname>T.</surname>
          </string-name>
           B. 
          <string-name>
            <surname>Brown</surname>
          </string-name>
          et al.,
          <article-title>Language models are few-shot learners</article-title>
          , arXiv,
          <year>2020</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.
          <year>2005</year>
          .14165
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>A.</given-names>
             
            <surname>Shrivastwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
           Gollapudi,
          <article-title>Hybrid cloud for architects: Build robust hybrid cloud solutions using AWS and OpenStack</article-title>
          , Packt Publishing,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>M.</given-names>
             
            <surname>Brown</surname>
          </string-name>
          ,
          <article-title>Bias mitigation in machine learning: Strategies and challenges</article-title>
          ,
          <source>Mach. Learning Rev</source>
          .
          <volume>33</volume>
          (
          <issue>1</issue>
          ) (
          <year>2023</year>
          )
          <fpage>50</fpage>
          -
          <lpage>75</lpage>
          . doi:
          <volume>10</volume>
          .7890/mlr.
          <year>2023</year>
          .033
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>Y.</given-names>
             
            <surname>Martseniuk</surname>
          </string-name>
          , et al.,
          <article-title>Automated conformity verification concept for cloud security</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems</source>
          , vol.
          <volume>3654</volume>
          ,
          <year>2024</year>
          ,
          <fpage>25</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>Y.</given-names>
             
            <surname>Martseniuk</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Shadow</surname>
            <given-names>IT</given-names>
          </string-name>
          <article-title>risk analysis in public cloud infrastructure</article-title>
          ,
          <source>in: Cyber Security and Data Protection</source>
          , vol.
          <volume>3800</volume>
          ,
          <year>2024</year>
          ,
          <fpage>22</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>Y.</given-names>
             
            <surname>Martseniuk</surname>
          </string-name>
          , et al.,
          <article-title>Universal centralized secret data management for automated public cloud provisioning</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems II</source>
          , vol.
          <volume>3826</volume>
          ,
          <year>2024</year>
          ,
          <fpage>72</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>