<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>KATH: A no-coding data processing aid for genetic researchers*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kajus Cerniauskas</string-name>
          <email>kajus.cerniauskas@ktu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dainius Kirsnauskas</string-name>
          <email>dainius.kirsnauskas@ktu.edu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paulius Preiksa</string-name>
          <email>paulius.preiksa@ktu.edu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gabrielius Salyga</string-name>
          <email>gabrielius.salyga@ktu.edu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nojus Sajauskas</string-name>
          <email>nojus.sajauskas@ktu.edu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Junius Vaitkus</string-name>
          <email>junius.vaitkus@ktu.edu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gerda Zemaitaityte</string-name>
          <email>gerda.zemaitaityte@ktu.edu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kazimieras Bagdonas</string-name>
          <email>kazimieras.bagdonas@ktu.edu</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kaunas University of Technology</institution>
          ,
          <addr-line>Kaunas</addr-line>
          ,
          <country country="LT">Lithuania</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents KATH - a No-Coding Data Processing system designed to assist genetic researchers at Harvard University in their work on mutations in the human genome related to eyesight pathologies. We aim to deploy a no-coding solution that enables R &amp; D researchers who are untrained in programming skills to process data from open-source gene databases with advanced DNA processing algorithms by applying a Large Language Model (LLM) based instruction interpreter. The paper presents an overview of the design system and the project's current status.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;No-coding</kwd>
        <kwd>AI</kwd>
        <kwd>LLM</kwd>
        <kwd>DNA</kwd>
        <kwd>Data analysis</kwd>
        <kwd>Data processing</kwd>
        <kwd>Open Source Database</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The exponential decrease in price for human DNA sequencing has reduced the cost from $3.8
billion for the Human Genome Project to under $1,000 currently, with expectations of reaching a $100
price tag shortly [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Such development enables geneticists to acquire vast amounts of genetic data for
their research activities. However, the rapid price decrease and rapid increase in available data need
to meet the increased capability to analyze this data, as training geneticists is a high-cost and
longterm endeavor. Further, by devoting their time to studying genetics, these world-class specialists in
genetics need to gain the IT skills to use the most sophisticated tools, such as Artificial Intelligence
(AI), which computer scientists often develop without considering the broader research community's
ability to use them. The limited number of IT specialists trained in software development and genetics
are known as Bioinformaticians and are highly sought after in industry and research institutions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
These experts' skills and limited time are best used for developing advanced software tools for DNA
data processing and analysis. However, in day-to-day research and development (R&amp;D) activities,
many tasks of relatively simple or moderate complexity require expertise in IT or computer science.
Currently, these tasks are being performed by geneticists themselves, which in the best-case scenario
ends up taking significant time from their primary functions, may stall the R&amp;D activity until a
bioinformatician colleague can allocate time to solve the issue, or, in the worst case, the skill gap may
become prohibitive to proceed with the intended R&amp;D activity.
      </p>
      <p>
        Due to the recent developments in AI and specifically the introduction of Large Language Models
(LLM) [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] that are capable of processing human speech, new opportunities have emerged to apply
the no-coding [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] paradigm in order to aid geneticists in their R&amp;D activities, by filling in the
aforementioned IT skill gap with specialized AI systems. LLM is a rapidly evolving technology that
has seen significant improvements in both proprietary and open-source models. Some models are
trained for general purposes and can understand human-produced text, speech, and even video. In
contrast, others are trained for specialized purposes, e.g., aiding in writing computer software [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ].
      </p>
      <p>The ability of geneticists to rapidly perform DNA analysis and effectively employ state-of-the-art
software tools can significantly benefit the scientific community and society as the discoveries in
genetics and pathologies of gene expressions directly influence the development of new medicine and
therapies.</p>
      <p>
        Currently, there are open-source scientific and medical databases of human DNA mutations and
associated pathologies that are being amended daily by researchers and physicians [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. These
opensource tools provide new opportunities for scientific and medical discoveries, with the existing
bottleneck of geneticists and bioinformaticians having sufficient time and skills required to perform
data analysis. The most prominent software tools for DNA data analysis focus on sequence alignment
or data mining. For example, Ensembl [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] gives access to genomes of various organisms and provides
data like gene annotations, genetic variation ratio, and more. Additionally, there is a possibility of
integrating other existing analysis tools like AlphaMissense [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] or REVEL [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] to make results more
precise. Another project might be the ENCODE framework, which provides information on where
and when the genes become active in cells—however, this software and similar software hinges on
expertise in genetics.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>2.1.</p>
      <sec id="sec-2-1">
        <title>Creation of the Database of genetic information</title>
        <p>
          The first step in creating the database is aggregating the data. It can be downloaded or
webscraped. Web scraping is the extraction of information from a website through computer software.
The software can use Hypertext Transfer Protocol (HTTP) methods or simulate a browser. The
program manipulates the captured unstructured web data into a demanded structure and moves it
into a database [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>
          The first stage is called the “fetching stage.” Requesting the website via Uniform Resource Locator
(URL) returns its Hypertext markup language (HTML). “Curl” and “wget” command line tools or
Python’s request library, Perl’s Mechanize module, or Java’s Apache HttpClient can make the
requests [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. To interact with a webpage's Document Object Model (DOM), "Selenium" is one of the
tools [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. The extraction stage utilizes regular expressions, HTML parsing libraries, and XML Path
Language (XPath) queries. After downloading a web page, the scraper uses the following tools:
Python's regular expressions library, BeautifulSoup library, and Lxml library [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. During the final
stage, the transformation stage, the remaining data becomes structured. Python's Pandas library is
one of the tools that manipulates and transforms the data [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>
          Another crucial step after initial data gathering would be modeling the database. It is known that
gathered data from different sources will likely contain different properties or attributes. The data
needs to be normalized to store this different data in a database. This phase involves analyzing
gathered data and identifying different entities, attributes, and relations within the data. This allows
us to create a base blueprint of our database model based on how the data structure would look. Upon
establishing the conceptual model, the next step involves its formal representation using established
modeling tools and methodologies, such as the Unified Modeling Language (UML), a widely adopted
standard for creating and displaying system architectures and data structures [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Entities within the
model are frequently connected through relationships, prompting the creation of Entity-Relationship
(ER) diagrams to visualize relations.
        </p>
        <p>2.2.</p>
      </sec>
      <sec id="sec-2-2">
        <title>No-coding technologies</title>
        <p>
          No-coding technologies provide a new perspective on using complex software for people needing
extensive coding knowledge or even informatics. These tools use artificial intelligence, particularly
large language models (LLMs), to provide users with a simple user interface (UI) that is intuitive and
easy to use yet utilizes powerful capabilities of complex algorithms. This enables people, especially
professional researchers without experience in coding, to focus and allocate more time to their
specific work rather than learning the complex intricacies of coding. Because of that, people can
accelerate the product's development, allowing for a broader range of individuals to collaborate and
contribute. No-coding technologies can be used in software development, data analysis, research, or
business specifics and in any field that uses complex software. This approach revolutionizes tasks and
improves efficiency and innovation across all industries. Many businesspeople see this technology as
the next step in the industry's future [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
        <p>In genetic research, no-coding technologies present a solution for not being able to use a specific
tool for an algorithm that requires some expertise in coding to execute. As mentioned before,
geneticists often need bioinformaticians' help to set up and use specific tools. Because of this,
researchers are wasting their time learning how to set up software, while bioinformaticians complete
a trivial but time-consuming task for them. This is a significant inconvenience that no coding
technologies can solve. By making a universal tool and user-friendly UI, researchers can provide data
and instructions that the software can interpret and provide the requested output.
2.3.</p>
      </sec>
      <sec id="sec-2-3">
        <title>AI - LLMs</title>
        <p>
          A Large Language Model (LLM) is an (AI) program trained on large amounts of data and is used to
recognize and generate text. LLMs are built on machine learning. They use neural networks and their
transformer model [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>
          Large Language Models (LLMs) provide a great way to skip the knowledge gaps to deliver results.
Current models already demonstrate general intelligence across various fields [19]. They show
remarkable capabilities, matching or exceeding expert's performance in the domain. The LLMs
significantly augment the professional's ability to forecast results in various fields[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The improved
fine-tuned models have even greater accuracy [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. One possible usage case is to act as a bridge
between genetics and Python script writing. The LLMs can hide the programming of the algorithms
with a no-coding approach using visual programming methods [20]. As of 2024 03 25, the up-to-date
list of available open-source LLMs is presented in a GitHub repository [21].
Nov-23
Aug-23
Jul-23
Jun-23
May-23
May-23
Apr-23
Mar-23
May-23
7
1.1
7
7
16
1.1-15
59k (Samples)
15k (Samples)
44,000k (Samples)
Dolphin
        </p>
        <sec id="sec-2-3-1">
          <title>DeciCoder-1B</title>
        </sec>
        <sec id="sec-2-3-2">
          <title>CodeGen2.5</title>
        </sec>
        <sec id="sec-2-3-3">
          <title>XGen-7B</title>
        </sec>
        <sec id="sec-2-3-4">
          <title>StarCoder</title>
        </sec>
        <sec id="sec-2-3-5">
          <title>MPT-7B-Instruct data bricks-dolly-15k OIG</title>
        </sec>
        <sec id="sec-2-3-6">
          <title>StarChat Alpha</title>
          <p>2.4.</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>DNA data processing algorithms (tools)</title>
        <p>Tensor processing units lack the memory to process genomes efficiently; therefore, researchers
implemented an algorithm for DNA sequence alignment that can be used within quantum simulation
to address performance problems [22]. This paper [23] analyzes major ML algorithms for data mining
and reviews current DNA sequence alignment, classification, clustering, and data mining
applications. Another research [24] explores the features of DNA-binding proteins using extraction
methods. Although this technique is already precedent, in other words, the tool outperforms many
algorithms of that kind in the UniSwiss dataset. The ENCODE consortium platform [25] aims to
narrow the data variations by supplying the pipelines since experimental labs use different protocols
for carrying out the research. DNA methylation analysis helps to switch repetitive genes off without
altering the DNA. This can be done with software like RnReads[26] and DunedinPoAm [27]; the
latter, however, is focused on predicting biological age.</p>
        <p>
          In the following table, we analyzed tools for scoring DNAs. Geneticists from Harvard provided a
list of tools as possible integration options with KATH. We analyzed them for compatibility with our
system's architecture and their overall purpose [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11, 28, 29, 30, 31, 32</xref>
          ].
        </p>
        <p>CADD is a tool for scoring the deleteriousness of single-nucleotide
variants, multi-nucleotide substitutions, and insertion/deletion variants in
the human genome.</p>
        <p>REVEL is an ensemble method for predicting the pathogenicity of
missense variants based on a combination of scores from 13 individual
tools.</p>
        <p>SpliceAI is a deep-learning-based tool used to score variants.</p>
        <p>Pangolin is a deep-learning-based method for predicting splice site
strengths.
32768
2048
2048
8192
8192
NA
NA
NA
8192</p>
        <p>Runs
locally</p>
        <p>No
Yes
Yes
Yes</p>
        <sec id="sec-2-4-1">
          <title>Metadome</title>
          <p>EVE is a model for predicting the clinical significance of human variants
based on sequences of diverse organisms across evolution.</p>
          <p>MetaDome analyses the mutation tolerance at each position in a human
protein.</p>
          <p>AlphaMissens AI model that predicts whether genetic mutations in proteins are likely to
be harmless or disease-causing.
No
Yes</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The proposed no-coding system architecture</title>
      <p>The proposed system comprises four major components: user interface, AI, DNA analysis,
processing and generation tools, and data storing and collection modules. The structure of the KATH
system is presented in Figure 1. All interactions and their directions of subsystems are shown below
in the diagram. The general flow is represented, although the final implementation architecture
might differ from the provided diagram.</p>
      <p>A user using a user interface (UI) can write his request to LLM, which is pre-configured with a
prime prompt. As the most simple implementation and lightweight solution, a web graphic interface
was proposed as a UI. The user interface should allow the user to see output from analysis tools,
represent data in the most convenient way, and interact with data.</p>
      <p>A simple and understandable interface should not restrict users' access to the implementation of
KATH and its code but provide the most efficient way to interact with the system. It saves time for
geneticists and also allows more advanced users to update the system according to their needs.
3.2.</p>
      <sec id="sec-3-1">
        <title>LLMs and prompt priming</title>
      </sec>
      <sec id="sec-3-2">
        <title>3.2.1. Prompt priming</title>
        <p>Our model will have a predefined set of functions, ranging from data retrieval from databases to
complex data manipulation operations such as merging, frequency calculations, and mutation
generation. These functions serve as the foundational building blocks for completing user requests,
which often require a combination of multiple predefined functions. To reduce errors and efficiently
complete user requirements, it is essential to ensure these functions' precise and efficient use. This is
where prompt priming plays a crucial role.</p>
        <p>Prompt priming serves as a mechanism to acquaint the LLM with the set of functions and to guide
its usage in accordance with the task at hand. Given the infinite range of prompts users may provide,
equipping the model to comprehend and execute requests is crucial. We take a methodical approach
to achieve this.</p>
        <p>Initially, the user's request is deconstructed into discrete steps, each corresponding to using a
predefined function within our model. This breakdown ensures that a specific function can execute
each part of the user's request. Furthermore, the order in which these functions must be used to
achieve the desired outcome is determined.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.2.2. User assistant LLM</title>
        <p>User assistant LLM is introduced to simplify nonspecialists' IT work with KATH. Its purpose is to
familiarize users with the system's basic functionality and assist in writing and checking the
correctness of written requests. It decreases the threshold of entry for new users and minimizes
possible errors from both users and the script-generating tool.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.2.3. Script generating LLM</title>
        <p>It converts user requests to the list of instructions to be executed: where to get data, what tool to
use to process it, and so on. This module entirely relies on the prime prompt. Therefore, the quality of
the result depends not only on a well-trained model but also on correct and understandable LLM
prompts that can "explain" to LLM what the output result should look like.</p>
        <p>3.3.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Tool integration</title>
      </sec>
      <sec id="sec-3-6">
        <title>3.3.1. Mutation generator</title>
        <p>As there are many different types of gene mutations, we must provide a tool for generating
different mutations by specifying a category of mutation and parameters. These mutations include
many alterations, ranging from single nucleotide changes to structural adjustments.</p>
        <p>Our tool will provide functionality to generate point mutations, deletions, inversions, insertions,
and duplications. A point mutation involves changing the specific nucleotide within the sequence,
possibly causing a cell to produce a different amino acid. Deletion, insertion, inversion, and
duplication involve a structural change of a gene by adding, removing, or rearranging nucleotide
sequences, potentially causing the frame to shift and affecting the splicing procedure, leading to
completely different protein production.</p>
        <p>As the user selects a type of mutation, our tool prompts them to provide corresponding
parameters. These parameters include the location or region of mutation on a specific gene, amino
acid change, and some information on genetic background.</p>
        <p>3.3.2. Tools API</p>
        <p>API facilitates the integration of tools represented by a predefined set of various functions. It
introduces scalability, allowing us to add functionality to our toolkit and modularity, encapsulating
each tool's functionality within defined interfaces. API allows standardized communication protocols
between tools, ensuring seamless data and results transactions.</p>
        <p>Initially, the API accepts input data in a standardized format, ensuring compatibility across
different tools. Upon receiving input data and the name of the target tool, the API identifies the
corresponding tool within our model. The API initiates the execution of the requested function
within the specified tool, using provided data as parameters. Lastly, it waits for the results and
outputs them for the user.</p>
        <p>3.4.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Data collection and refactoring</title>
        <p>All databases have an official website where it is possible to download the data about gene
mutations by pressing a button. However, each database has a different way of downloading data
about gene mutations. LOVD has a static Uniform Resource Locator (URL) for all genes. Data can be
retrieved using Python script. Some links are accessible to anyone, and others have restricted
permission to download them. ClinVar and GnomAd have dynamic URLs, meaning we cannot apply
the same solution as for LOVD.</p>
        <p>For this reason, we get this data by executing JavaScript code on the pages to press the "download"
button. The last approach also provides a universal way of downloading data for any website as long
as pages do not change their hypertext markup language (HTML) structure. However, it is always
better to search for another way to retrieve data for the system's time efficiency.</p>
        <p>Due to structural differences in formats across databases, refactoring to the universal structure
was required. According to the preferences of Harvard scientists, LOVD's database format became
the basis for the system. The main tasks while merging were to solve the inconsistency of data in
LOVD and extract relevant information required by geneticists.</p>
        <p>Data about mutation positions is mixed with the usage of new and old coding notations. For some
mutations, databases contain protein positions instead of cDNA positions. Tools in the refactoring
package are applied to align data from other databases with data from LOVD, solving the mentioned
problems.</p>
        <p>The system's design was driven by a decision-making methodology prioritizing user accessibility,
flexibility, and scalability.</p>
        <p>One of the critical decisions was incorporating large language models (LLMs). LLMs can
understand and generate human-like text, making them well-suited for natural language processing
tasks. By leveraging LLMs, the system could provide a user-friendly interface that allows geneticists
to interact with the system using natural language, eliminating the need for complex coding or
command-line interfaces.</p>
        <p>The decision to employ two distinct LLMs, the User Assistant LLM and the Script Generating LLM,
was strategic: the User Assistant LLM serves as a conversational interface, assisting users in
formulating their requests and providing explanations or clarifications when needed. This
component was designed to make the system more accessible to non-technical users by allowing
them to communicate their requirements naturally and intuitively. The Script Generating LLM, on
the other hand, was chosen to translate the user's requests into executable instructions, bridging the
gap between natural language and programmatic instructions.</p>
        <p>A Prompt Priming Module was incorporated to ensure the effectiveness of the Script Generating
LLM. This module plays a crucial role in preparing the user's request for the LLM, ensuring that the
LLM receives the request in a format that it can effectively understand and process. The quality of the
output generated by the Script Generating LLM heavily depends on the effectiveness of the prompt
priming process, and this decision aimed to optimize the LLM's performance.</p>
        <p>Throughout the design process, decisions were made to create a user-friendly, flexible, and
scalable system.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>The KATH system architecture was developed to facilitate the no-coding approach and meet the
needs of geneticists performing R&amp;D activities. Specifications and system requirements have been
analyzed, and the minimal viable prototype for the KATH system has been defined.</p>
      <p>The analysis of existing public DNA data processing tools has been carried out, and candidates for
integration into minimal viable prototypes have been selected.</p>
      <p>The analysis of existing public gene mutation and pathogenicity databases was performed. An API
was implemented to automate the periodic retrieval and refactoring of the data from LOVD,
gnomAD, and Clinvar databases. Due to the unique format of LOVD data, the additional function to
parse the downloaded data was integrated into the data collection pipeline. The most significant data
for Harvard scientists from Clinvar and gnomAd is extracted and refactored. Data from databases is
merged on values from cDNA mutations' coding variables.</p>
      <p>An analysis of candidate LLMs for the UI and script generation functionalities has been
performed, and candidate models were preselected for the testing and integration stages of the
project. We evaluated the effectiveness of different large language models (LLMs) - OLMO, Dolphin,
ChatGPT, and DeciCoder-1b - in converting prompts from scientists into Python code for processing
genetic data. The LLMs were tasked with generating code to download data from genetic databases,
merge data, convert to VCF format, and process with the SpliceAI tool.</p>
      <p>OLMO struggled to understand some prompts and provided incomplete or incorrect responses.
Dolphin successfully generated the expected code for all prompts. ChatGPT also exhibited an
adequate understanding, providing correct code. However, DeciCoder-1b failed to provide any
answers and, therefore, was not included in the results table.</p>
      <p>The results highlighted LLMs' varying capabilities in handling code generation tasks from natural
language prompts. While some models, like Dolphin and ChatGPT, performed reasonably well,
others, like OLMO and DeciCoder-1b, exhibited limitations or failed to generate the expected output.
print "store_database('lovd')" I don't understand... what do you mean? Print a
and nothing more random example and nothing more?</p>
      <p>pass "lovd" argument to ok, that's all right, thank you! The end
"store_database(database_name)"
function and print result</p>
      <p>download data from clinvar
and lovd. Merge data, convert to
vcf format and pass to SpliceAI.</p>
      <p>We downloaded the data from ClinVar and LovD,
merged them using the merge_lovd_with_clinvar
and merge_lovd_with_gnomad functions
respectively, and converted the...</p>
      <p>download data from clinvar store_database('clinvar')
and lovd. Merge data, convert to merge_lovd_with_clinvar()
vcf format and pass to SpliceAI process_with_splice_ai()
store_database('lovd')
convert_to_vcf()
download data from clinvar processed_data =
and lovd. Merge data, convert to process_with_splice_ai(convert_to_vcf(merge_lo
vcf format and pass to SpliceAI. vd_with_clinvar(store_database('clinvar'),
Display results. store_database('lovd'))))</p>
      <p>processed_data.display()
ChatGPT</p>
      <p>Download
GNOMAD
data</p>
      <p>from store_database('gnomad')
download data from clinvar Hey, here is your code
and lovd. Merge data, convert to store_database('clinvar')
vcf format and pass to SpliceAI store_database('lovd')
merge_lovd_with_clinvar()
convert_to_vcf()
process_with_splice_ai()
download data from clinvar store_database('clinvar')
and lovd. Merge data, convert to store_database('lovd')
vcf format and pass to SpliceAI. merge_lovd_with_clinvar()
Display results. convert_to_vcf()
processed_data = process_with_splice_ai()
processed_data.display()</p>
    </sec>
    <sec id="sec-5">
      <title>5. Future Work</title>
      <p>The following steps in realizing the minimal viable prototype of the KATH system involve
improving and generalizing the database module. We seek to implement the conversion of disparate
notations of the gene location in the genome and the merging of different public databases into a
homogeneous solution. The unified relational database management system is also planned to
facilitate higher performance and scalability.</p>
      <p>The AI module is scheduled to be implemented during Q2 of 2024. To evaluate the performance of
selected candidates, an extensive test campaign will be conducted for both user assistance AI based on
a general LLM and script-generating AI based on code-generating LLM. The prompt priming module
will be integrated into the subsystem, and subsystem-level validation tests will be conducted.</p>
      <p>For the minimal viable prototype, the selected DNA analysis tools will be integrated with the
KATH system, which will be tested at Massachusetts Eye and Ear, Harvard University. The minimum
viable prototype will be focused on the Eyes shut homolog (EYS) gene data. The user interface will be
implemented in a chat-based system that will enable researchers to express their desired results in
text and receive the results of the DNA analysis as textual output and as files stored in the database.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>The paper describes a proposed architecture for a no-coding data processing system for genetic
researchers. The KATH system consists of a UI designed for a user with no significant programming
skills or IT training that enables them to express their desired function colloquially. The user's input
is provided to the first LLM, which is used to convert it to a concrete action plan and subsequently
converted to commands by a specialized LLM. These commands are provided to one or several
integrated DNA analysis tools. A data retrieval and database refactoring module is designed to
automatically update the local database and refactor the gathered data into a homogeneous data
structure. The data is aggregated and refactored from publicly available data on human genetic
mutations and associated pathologies obtained from open-access scientific databases such as LOVD,
gnomAD, and Clinvar. The refactored data and the obtained results from integrated open-source
analysis tools for genetic data, such as REVEL, Metadome, AlphaMissens, etc., are stored locally.</p>
      <p>We expect the minimum viable system to be deployed in Q3 2024. Further tests and development
are expected to occur in collaboration with academic partners from the US and Germany and the
leading industry R&amp;D partners from Lithuania.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We would like to acknowledge the assistance of associate professor Ph.D. Kinga M. Bujakowska,
Ph.D. Egle Galdikaite-Braziene and Ph.D. Riccardo Sangermano from Massachusetts Eye and Ear,
Harvard University, and express our sincere gratitude for their guidance and expert support.</p>
      <p>G. D. Varsamis et al., “Quantum gate algorithm for reference-guided DNA sequence
alignment,” Computational Biology and Chemistry, vol. 107, p. 107959, Dec. 2023, doi:
10.1016/j.compbiolchem.2023.107959.</p>
      <p>A. Yang, W. Zhang, J. Wang, K. Yang, Y. Han, and L. Zhang, “Review on the
Application of Machine Learning Algorithms in the Sequence Data Mining of DNA,” Front.
Bioeng. Biotechnol., vol. 8, Sep. 2020, doi: 10.3389/fbioe.2020.01032.</p>
      <p>A. Sun, H. Li, G. Dong, Y. Zhao, and D. Zhang, “DBPboost:A method of classification
of DNA-binding proteins based on improved differential evolution algorithm and feature
extraction,” Methods, vol. 223, pp. 56–64, Mar. 2024, doi: 10.1016/j.ymeth.2024.01.005.</p>
      <p>Y. Luo et al., “New developments on the Encyclopedia of DNA Elements (ENCODE)
data portal,” Nucleic Acids Research, vol. 48, no. D1, pp. D882–D889, Jan. 2020, doi:
10.1093/nar/gkz1062.</p>
      <p>F. Müller et al., “RnBeads 2.0: comprehensive analysis of DNA methylation data,”
Genome Biol, vol. 20, no. 1, p. 55, Mar. 2019, doi: 10.1186/s13059-019-1664-9.</p>
      <p>D. W. Belsky et al., “Quantification of the pace of biological aging in humans through
a blood test, the DunedinPoAm DNA methylation algorithm,” eLife, vol. 9, p. e54870, May
2020, doi: 10.7554/eLife.54870.</p>
      <p>P. Rentzsch, M. Schubach, J. Shendure, and M. Kircher, “CADD-Splice—improving
genome-wide variant effect prediction using deep learning-derived splice scores,” Genome
Med, vol. 13, no. 1, p. 31, Dec. 2021, doi: 10.1186/s13073-021-00835-9.</p>
      <p>K. Jaganathan et al., “Predicting Splicing from Primary Sequence with Deep
Learning,” Cell, vol. 176, no. 3, pp. 535-548.e24, Jan. 2019, doi: 10.1016/j.cell.2018.12.015.</p>
      <p>T. Zeng and Y. I. Li, “Predicting RNA splicing from DNA sequence using Pangolin,”
Genome Biology, vol. 23, no. 1, p. 103, Apr. 2022, doi: 10.1186/s13059-022-02664-4.</p>
      <p>“Evolutionary model of Variant Effect.” Accessed: Mar. 23, 2024. [Online]. Available:
https://evemodel.org/</p>
      <p>“MetaDome web server.” Accessed: Mar. 23, 2024. [Online]. Available:
https://stuart.radboudumc.nl/metadome/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[1] “DNA Sequencing Costs: Data.” Accessed: Mar. 24</source>
          ,
          <year>2024</year>
          . [Online]. Available: https://www.genome.gov/about-genomics/
          <article-title>fact-sheets/DNA-</article-title>
          <string-name>
            <surname>Sequencing-</surname>
          </string-name>
          Costs-Data
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rocha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Massarani</surname>
          </string-name>
          , S. J. de Souza,
          <article-title>and</article-title>
          <string-name>
            <surname>A. T.</surname>
          </string-name>
          R. de Vasconcelos, “
          <article-title>The past, present and future of genomics and bioinformatics: A survey of Brazilian scientists</article-title>
          ,”
          <source>Genet Mol Biol</source>
          , vol.
          <volume>45</volume>
          , no.
          <issue>2</issue>
          , p.
          <fpage>e20210354</fpage>
          , doi: 10.1590/
          <fpage>1678</fpage>
          -4685
          <string-name>
            <surname>-GMB-</surname>
          </string-name>
          2021-0354.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          et al., “Attention is All you Need,” in
          <source>Advances in Neural Information Processing Systems</source>
          , Curran Associates, Inc.,
          <year>2017</year>
          . Accessed: Mar.
          <volume>24</volume>
          ,
          <year>2024</year>
          . [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2017/hash/3f5ee243547dee91fbd053c1c4a8 45aa-
          <fpage>Abstract</fpage>
          .html
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Narasimhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salimans</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Sutskever</surname>
          </string-name>
          ,
          <article-title>"Improving Language Understanding by Generative Pre-Training."</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Faes</surname>
          </string-name>
          et al.,
          <article-title>“Automated deep learning design for medical image classification by health-care professionals with no coding experience: a feasibility study,” The Lancet Digital Health</article-title>
          , vol.
          <volume>1</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>e232</fpage>
          -
          <lpage>e242</lpage>
          ,
          <year>Sep</year>
          .
          <year>2019</year>
          , doi: 10.1016/S2589-
          <volume>7500</volume>
          (
          <issue>19</issue>
          )
          <fpage>30108</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T. D.</given-names>
            <surname>Ross</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Gopinath</surname>
          </string-name>
          , “
          <article-title>Chaining thoughts and LLMs to learn DNA structural biophysics</article-title>
          .
          <source>” arXiv, Mar. 02</source>
          ,
          <year>2024</year>
          . Accessed: Mar.
          <volume>20</volume>
          ,
          <year>2024</year>
          . [Online]. Available: http://arxiv.org/abs/2403.01332
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Schoenegger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Park</surname>
          </string-name>
          , E. Karger, and
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Tetlock</surname>
          </string-name>
          , “
          <string-name>
            <surname>AI-Augmented</surname>
            <given-names>Predictions</given-names>
          </string-name>
          :
          <article-title>LLM Assistants Improve Human Forecasting Accuracy</article-title>
          .” arXiv, Feb.
          <volume>12</volume>
          ,
          <year>2024</year>
          . Accessed: Mar.
          <volume>20</volume>
          ,
          <year>2024</year>
          . [Online]. Available: http://arxiv.org/abs/2402.07862
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Gunasekaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ramalakshmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rex Macedo Arokiaraj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Deepa</given-names>
            <surname>Kanmani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Venkatesan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Suresh Gnana</surname>
          </string-name>
          <string-name>
            <surname>Dhas</surname>
          </string-name>
          , “
          <article-title>Analysis of DNA Sequence Classification Using CNN and Hybrid Models,” Computational and Mathematical Methods in Medicine</article-title>
          , vol.
          <year>2021</year>
          , p.
          <fpage>e1835056</fpage>
          ,
          <string-name>
            <surname>Jul</surname>
          </string-name>
          .
          <year>2021</year>
          , doi: 10.1155/
          <year>2021</year>
          /1835056.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[9] “Ensembl genome browser 111.” Accessed: Mar. 24</source>
          ,
          <year>2024</year>
          . [Online]. Available: https://www.ensembl.org/index.html
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10] J. Cheng et al., “
          <article-title>Accurate proteome-wide missense variant effect prediction with AlphaMissense,” Science</article-title>
          , vol.
          <volume>381</volume>
          , no.
          <issue>6664</issue>
          , p.
          <fpage>eadg7492</fpage>
          ,
          <string-name>
            <surname>Sep</surname>
          </string-name>
          .
          <year>2023</year>
          , doi: 10.1126/science.adg7492.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>“</surname>
            <given-names>REVEL</given-names>
          </string-name>
          :
          <article-title>Rare Exome Variant Ensemble Learner</article-title>
          .”
          <source>Accessed: Mar. 23</source>
          ,
          <year>2024</year>
          . [Online]. Available: https://sites.google.com/site/revelgenomics/
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Khder</surname>
          </string-name>
          , “
          <article-title>Web Scraping or Web Crawling: State of Art, Techniques, Approaches</article-title>
          and Application,” IJASCA, vol.
          <volume>13</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>145</fpage>
          -
          <lpage>168</lpage>
          , Dec.
          <year>2021</year>
          , doi: 10.15849/IJASCA.211128.11.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13] “
          <article-title>Web scraping technologies in an API world</article-title>
          | Briefings in Bioinformatics | Oxford Academic.” Accessed: Mar.
          <volume>20</volume>
          ,
          <year>2024</year>
          . [Online]. Available: https://academic.oup.com/bib/article/15/5/788/2422275
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nyamathulla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Ratnababu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Shaik</surname>
          </string-name>
          , and
          <string-name>
            <surname>B. L. N</surname>
          </string-name>
          , “
          <article-title>A Review on Selenium Web Driver with Python,”</article-title>
          <source>Annals of the Romanian Society for Cell Biology</source>
          , pp.
          <fpage>16760</fpage>
          -
          <lpage>16768</lpage>
          , Jun.
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <article-title>“Algorithmic Enumeration of Ideal Classes for Quaternion Orders</article-title>
          |
          <source>SIAM Journal on Computing.” Accessed: Mar. 23</source>
          ,
          <year>2024</year>
          . [Online]. Available: https://epubs.siam.org/doi/10.1137/080734467
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Torre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Genero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Labiche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Elaasar</surname>
          </string-name>
          , “
          <article-title>How consistency is handled in model-driven software engineering and UML: an expert opinion survey</article-title>
          ,
          <source>” Software Qual J</source>
          , vol.
          <volume>31</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>54</lpage>
          , Mar.
          <year>2023</year>
          , doi: 10.1007/s11219-022-09585-2.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S. F. A.</given-names>
            <surname>Razak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. P.</given-names>
            <surname>Ernn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. I.</given-names>
            <surname>Yussoff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U. A.</given-names>
            <surname>Bukar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Yogarayan</surname>
          </string-name>
          , “
          <article-title>Enhancing Business Efficiency through Low-Code/No-Code Technology Adoption: Insights from an Extended UTAUT Model,”</article-title>
          <source>Journal of Human, Earth, and Future</source>
          , vol.
          <volume>5</volume>
          , no.
          <issue>1</issue>
          ,
          <string-name>
            <surname>Art</surname>
          </string-name>
          . no.
          <issue>1</issue>
          ,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          .
          <year>2024</year>
          , doi: 10.28991/HEF-2024
          <string-name>
            <surname>-</surname>
          </string-name>
          05-01-07.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <article-title>“Large language models (LLM) and ChatGPT: what will the impact on nuclear medicine be?</article-title>
          |
          <source>European Journal of Nuclear Medicine and Molecular Imaging.” Accessed: Mar. 21</source>
          ,
          <year>2024</year>
          . [Online]. Available: https://link.springer.com/article/10.1007/s00259-023- 06172-w S. Bubeck et al.,
          <source>“Sparks of Artificial General Intelligence</source>
          :
          <article-title>Early experiments with GPT-4</article-title>
          .” arXiv, Apr.
          <volume>13</volume>
          ,
          <year>2023</year>
          . Accessed: Mar.
          <volume>20</volume>
          ,
          <year>2024</year>
          . [Online]. Available: http://arxiv.org/abs/2303.12712
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cai</surname>
          </string-name>
          et al., “
          <string-name>
            <surname>Low-code</surname>
            <given-names>LLM</given-names>
          </string-name>
          :
          <article-title>Visual Programming over LLMs</article-title>
          .” arXiv, Apr.
          <volume>20</volume>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>Accessed: Mar. 20</source>
          ,
          <year>2024</year>
          . [Online]. Available: http://arxiv.org/abs/2304.08103 E. Yan, “eugeneyan/open-llms.
          <source>” Mar. 22</source>
          ,
          <year>2024</year>
          . Accessed: Mar.
          <volume>22</volume>
          ,
          <year>2024</year>
          . [Online].
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          Available: https://github.com/eugeneyan/open-llms [
          <volume>29</volume>
          ] [30] [31] [32]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>