Development and analysis of a parallel method for detecting chromosomal translocations and inversions in DNA sequences⋆ Lesia Mochurad a,∗,†, Renata Vladyka a,† a Lviv Polytechnic National University, 12 Bandera street, Lviv, 79013, Ukraine Abstract In the current context of the growing volume of genetic research and the need for the fast and accurate detection of chromosomal abnormalities, developing efficient algorithms that provide high performance, scalability, and accuracy in comparing DNA sequences is a critical task. In this study, the aim was to create and analyze a parallel algorithm for detecting chromosomal translocations and inversions in DNA sequences. For this purpose, we implemented an algorithm capable of comparing DNA sequences to detect genetic abnormalities, including translocations and inversions. The use of parallel computing made it possible to significantly improve the efficiency of the analysis, reducing the algorithm's execution time and increasing scalability. Particular attention was paid to how the algorithm uses multithreading and how it efficiently distributes the load between threads. To detect inversions, an algorithm was developed that compares two DNA sequences and detects possible changes in nucleotide sequences. To find translocations, the Needleman-Wunsch algorithm was applied, which uses parallel computing to optimally align genetic fragments. The results of the algorithms have shown high efficiency in detecting chromosomal mutations, which is important for genetic research and can be used for medical diagnosis. Keywords Translocation, inversion, DNA sequence, Needleman-Wunsch algorithm, nucleotides. 1. Introduction This paper is devoted to the development of an algorithm for preventing chromosomal translocations and inversions. Chromosomal translocations and inversions are complex structural changes in genetic materials that occur as a result of disruptions in the organization of chromosomes. Translocations involve moving a part of a chromosome to another, causing genetic sequences to be interrupted and reconnected, while inversions occur when a chromosome fragment rotates around its axis, changing the order of genes and other genetic elements. These structural variations can occur as a result of various errors during cell division. Incorrect joining or rotation of chromosome parts can lead to important changes in genetic information, which is often associated with the development of various genetic diseases and developmental disorders [1]. The main reasons why it is important to avoid chromosomal translocations and inversions: 1. A translocation between chromosomes 21 and 14 can lead to “familial” Down syndrome, where one of the parents may be a carrier of the translocation. This, in turn, can affect the likelihood of having a child with Down syndrome. 2. Gametes carrying defective chromosomes with inversion often develop into non-viable or miscarried organisms in the early stages of embryogenesis. 3. Robertson translocations can cause differences in the number of chromosomes between closely related species, which affects their fertility and genetic stability. 4. If gametes with an inverted chromosome develop into organisms, 50% of the gametes of these organisms may not be viable. However, the other 50% may be viable and maintain the mutation in the population. IDDM’24: 7th International Conference on Informatics & Data-Driven Medicine, November 14 - 16, 2024, Birmingham, UK ∗ Corresponding author. † These authors contributed equally. lesia.i.mochurad@lpnu.ua (L. Mochurad); renata.vladyka.shi.2022@lpnu.ua (R. Vladyka) 0000-0002-4957-1512 (L. Mochurad); 0009-0003-7965-044 (R. Vladyka) © 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). CEUR ceur-ws.org Workshop ISSN 1613-0073 Proceedings Thus, the diagnosis and prediction of chromosomal inversions and translocations are critical in genetic research and medical practice. Detection of these changes contributes to early diagnosis, which allows timely treatment and monitoring of patients. Predicting the risks of developing certain diseases based on the identified mutations also helps patients to realize their genetic risks and opportunities for prevention. To solve this problem, the Needleman-Wunsch algorithm is used [2]. However, one of the main problems with this algorithm is its high computational complexity; this algorithm has quadratic complexity. The latter makes it difficult to process large DNA sequences. The development of a parallel algorithm can reduce the execution time by simultaneously processing parts of the data [3], [4], [5]. Ensuring efficient parallelization requires balanced load distribution among threads to prevent idle time, which may involve optimizing data structures and algorithms. Proper synchronization is also essential to avoid conflicts when accessing shared resources, especially in algorithms with interdependent results. Maintaining result accuracy during parallel processing may require extra validation. Lastly, scalability is a key factor. It should be able to scale across different hardware platforms, from personal computers to supercomputer clusters, to achieve optimal results in different environments [6], [7]. Solving all the problems described above is an important step to improve the speed and accuracy of chromosomal abnormalities diagnosis, which, in turn, will increase the efficiency of genetic research and clinical practice. In the field of genetic research, there are many important aspects that relate to the study of genetic mutations, their diagnosis, and their impact on human health. Various articles offer valuable insights that emphasize the importance of such mutations in medicine, evolution, and biology. For example, paper [8] discusses the work of the Human Gene Mutation Database (HGMD), which collects and analyzes information about human genetic mutations. The authors describe the functions of the HGMD, its use in clinical diagnostics and research, as well as plans to expand the database, which include the integration of GTEx project data and the introduction of automated tools for mutation prediction. The advantages of the article are a detailed overview of the HGMD functions and an emphasis on its role in clinical practice, but the disadvantages are the complexity of automatic identification of new mutations and the lack of examples of real-world use of the HGMD in clinical cases. In [9], the authors analyze different types of genetic mutations and their impact on human health, including the occurrence of diseases due to mutations in genes or chromosomes. The advantages of this article are a clear description of the causes and consequences of genetic diseases, as well as a detailed overview of the types of mutations. However, there are also disadvantages, such as general statements that are not always supported by specific data, as well as a lack of details about research methods. Nevertheless, the article is useful for familiarizing oneself with genetic mutations and their impact on the human body. Paper [10] investigates the use of biodosimetry to assess radiation doses in US military personnel who participated in nuclear tests after World War II. The main goal of this study is to assess chromosomal aberrations, such as inversions and translocations, as a method of retrospective biodosimetry. The results show that inversions can be an effective way to establish radiation doses received decades ago and that their combination with translocations improves the accuracy of the estimate. The study also indicates the influence of age and smoking status on the frequency of aberrations, which is higher in older individuals and smokers. The PETI method is discussed in [11]. It effectively induces precise DNA recombinations in human cells. PETI successfully generates recombination and inversion mutations in endogenous genomes and can correct disease-related inversions and translocations, making it promising for disease modeling and therapeutic approaches. This method is characterized by high accuracy and flexibility compared to other genome editing methods such as Prime-Del and twinPE, but it is still limited in its use for episomal mutations and requires further research to confirm its results. The authors of [12] investigate the effectiveness of the third generation of sequencing analysis for detecting breakpoints in unbalanced chromosomal translocations. The use of Nanopore long reads allows for the detection of microdeletions, microinsertions, and other structural changes that complement translocations. This approach provides accurate genetic information necessary for diagnosis and treatment. However, the high cost and possible experimental failures point to the need for further research and optimization of methods. The paper [13] discusses the role of chromosomal rearrangements in the genetic differentiation and evolution of populations, in particular, using the example of the spiny frog with a chromosomal translocation polymorphism. The authors use whole-chromosome staining (WCP) and genetic analysis to confirm the common origin of the translocations. They found that translocated chromosomes have a higher level of genetic differentiation due to recombination suppression, which contributes to genetic diversity and population differentiation. The article also discusses the mechanisms of recombination suppression, such as the accumulation of repetitive sequences and the capture of adapted alleles. This review is rounded off by article [14], which analyzes the importance of chromosomal inversions in the plant kingdom and their role in plant evolution. The authors note that inversions are common in many groups of plants and are often associated with locally advantageous traits that promote adaptation and speciation. The article also discusses methods for detecting inversions, such as karyotyping, genetic mapping, and high-fidelity sequencing. The authors call for further research to better understand the origin, evolutionary role, and molecular mechanisms of inversions in plants. Thus, the analysis of scientific studies shows that existing approaches to diagnosing and predicting chromosomal inversions and translocations have their advantages, but also significant disadvantages. For example, the HGMD database, although providing useful information for clinical practice, faces problems with the automatic identification of new mutations. The use of biodosimetry to estimate radiation doses, in particular by analyzing chromosomal aberrations, has proven effective, but is dependent on factors such as age and smoking status. The PETI method demonstrates high accuracy in genetic editing, but requires additional research to confirm its widespread use. Chromosomal rearrangements confirm their role in the genetic differentiation of populations, but require a deeper study of the mechanisms of recombination suppression. The study of inversions in plants also emphasizes the need for further research into the evolutionary mechanisms of these changes. This confirms the importance of continuing research in this area. In particular, the need to develop a parallel algorithm for the diagnosis and prediction of chromosomal inversions and translocations is driven by the need for faster and more accurate detection of genetic abnormalities, which is critical for effective treatment and management of diseases. The main contribution of this paper:  A new parallel method for detecting chromosomal translocations and inversions using multithreading and parallel computing (OpenMP) is proposed. This method ensures optimal load distribution between threads, reducing processing time and increasing scalability.  The Needleman-Wunsch alignment method has been parallelized to compare sequences, which has increased the accuracy of translocation detection. This approach allows for more accurate identification of chromosomal abnormalities, outperforming traditional sequential alignment methods.  When using eight threads, the proposed method achieves speedups of up to 3.8 for translocations and 3.4 for inversions. This confirms its suitability for processing large amounts of data and the possibility of further optimization on multi-core or GPUs. 2. Methodic and materials 2.1. General scheme of the method for detecting chromosomal mutations The task is to develop a method for detecting mutations in DNA sequences. There are two sets of DNA for inversion and four for translocation: half of them are taken as a reference, i.e. without mutations, and the rest with mutations. This method should be able to detect two types of mutations: inversions and translocations. Let � and � be DNA sequences of length � and �, respectively, �=(�1, �2, ..., ��) and �=(�1, �2, ..., ��), where �i and �i take the values {A, C, T, G}. Sequence � is the same sequence �, but at a certain interval from i to i + k the sequence is inverted, i.e. �[i:i+k] ! �[i:i+k] and �[i:i+k] = reverse(�[i:i+k]). To search for inversions, the algorithm comparing two DNA sequences is parallelized. This algorithm detects mismatches in nucleotide sequences and determines the gaps where nucleotides in DNA molecules do not match. After that, the found gaps are checked for inversions. The result of this check is the determination of the gaps where inversions are detected. To search for translocations, we need to parallelize the Needleman-Wunsch algorithm, which consists in constructing an optimal alignment, using the optimal alignments of the initial fragments of the original sequences obtained in the previous steps. For two sequences � and � with elements �I (0