<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>ORCID:</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Tongue Twisters Detection in Ukrainian by Using TDA</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Iryna Yurchuk</string-name>
          <email>i.a.yurchuk@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olga Gurnik</string-name>
          <email>olga.gurnick@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aviation University”</institution>
          ,
          <addr-line>Metrobudivska str. 5-a, Kyiv, UA- 03065</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Taras Shevchenko National University of Kyiv</institution>
          ,
          <addr-line>Bohdan Hawrylyshyn str. 24, Kyiv, UA-04116</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>The current stage of development of the digital world requires solving many problems. In particular, the analysis of texts occupies one of the important places. The scope of such analysis is very wide, from the semantic analysis of resumes to the psychological portrait of a social network user based on his posts. The algorithm of tongue twisters detection in Ukrainian using the topological data analysis methods based on the consideration of a twister as a sample in the four-dimensional space with a construction simplicial structure on it, calculating its invariant and classifiers as machine learning methods to realize distinguishing a tongue twister from a simple narrative sentence is obtained by authors. A better rate of detection was obtained by using a support vector machine with a Gaussian RВF kernel. Ukrainian tongue twister, persistent homology, decision tree, support vector machine COLINS-2023: 7th International Conference on Computational Linguistics and Intelligent Systems, April 20-21, 2023, Kharkiv, Ukraine</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>A tongue twister is a syntactically short, correct phrase spoken in any language with especially
complicated articulation. The tongue twisters contain combinations of sounds that sound similar but
have different phonemes and are difficult to pronounce. A tongue twister is naive in content, simple,
even primitive text, which is built on intricate and difficult phrases and phrases. Often tongue twisters
contain rhymes and alliteration. There are several aspects that contribute to the need to study tongue
twisters as a cluster of texts in human speech:
</p>
      <p>They can be called hard talkers. Working on tongue twisters is very good, very useful diction
training. It is a mistake to think that it is necessary to achieve an extremely fast pace. It is not always
so. When dealing with hard talkers, as in all work on diction, each speaker requires an individual
approach.
 Speech therapists adhere to another name for this type of speech art - pure speaking. Phrases are
used to practice pronunciation and diction. Phrases are ideal for differentiating and training
automatic pronunciation. For the greater development of the speech apparatus, it is recommended to
pronounce colloquialisms with a change in tempo and volume, i.e. quietly, loudly or in a whisper,
slowly or quickly. But the main task when reading tongue twisters is to do it clearly and cleanly,
with deep, full articulation, regardless of the volume and speed of speech. Politicians, actors and
other public figures are well familiar with this type of literary creativity and the miracle that happens
to their articulation, diction and manner of speech after a few hard sessions with a good speech
therapist.
 Studying the connections between colloquialisms and the peculiarities of the language zone of
the brain, scientists found out why colloquialisms are so difficult to pronounce. The reason is the
close location of the groups of neurons needed to pronounce a set of sounds in colloquial speech.</p>
      <p>2023 Copyright for this paper by its authors.
Experts have identified two main mechanisms of the so-called “double attack” - a phenomenon when
the speaker tries to pronounce two sounds at the same time. The first type of error is when a person
tries to pronounce two sounds almost simultaneously, the second - when mixing sounds with a slight
delay, and then a vowel sound appears between consonants.</p>
      <p>Purpose of the work – to propose an algorithm and its realization for tongue twisters detection in the
Ukrainian language that provides the understanding of text shape aspects with the possibility to
implement it into machine learning.</p>
      <p>The aim of research – to propose topological invariants, the calculation of which will be informative
for understanding the nature of a tongue twister, to establish a set of data, the integration of which in a
certain method of machine learning has the ability to distinguish a tongue twister from a simple narrative
sentence.</p>
      <p>Major research objectives are:
 using methods of topological data analysis (TDA), propose an away to encode the Ukrainian
tongue twisters that allows to integrate it in machine learning;
 based on obtained results, propose an algorithm for detecting the tongue twisters among simple
narrative sentences.</p>
      <p>Practical tasks in which the algorithm of tongue twisters detection in Ukrainian can be used are the
following:
 rating complexity of existing texts for children who have speech problems. We analyze the
sentences of the text and check whether they have the feature of a tongue twister. The more tongue
twisters there are, the more difficult for such children a text is.
 artificial generation of a special type of text in the Ukrainian language. It may be senseless, which
is allowed by the nature of a tongue twister, but it is aimed at eliminating one or another problem,
for example, improving the speech of announcers or developing it in younger children.
 artificial generation of tongue twisters will make it possible to enrich the existing dataset (which
is currently consists of no more than 500 items for the Ukrainian language) in order to apply NLPT
for tongue twisters detection.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>In this section, we consider some existing works on analyzing text by its figurative shape, text
artificiality detection, classification, visualization and methods of machine learning with help of
topological data analysis (TDA). Below, the authors focus specifically on existing solutions using
topological data analysis.</p>
      <p>Note that TDA is a relatively new direction in data analysis, its main advantage is the ability to learn
the space characteristics from the point of view of its geometry (for details, see Section 3.3). However,
movement in this direction requires an understanding of the basic concepts and aspects of such a science
as topology, which often complicates and slows down the process of obtaining new results in this area.</p>
      <p>
        In today's world, social media produces a huge amount of content [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This content needs to be
processed to identify interesting topics, but this cannot be done manually. Various automated
approaches to topic detection have previously been proposed for use. Most of these methods use
document clustering and packet detection. For the most part, these approaches represent text functions
in standard n-dimensional Euclidean metric spaces. However, these methods have a drawback. It is
difficult to determine the subject when filtering noisy documents directly. The authors propose using
topology as a subject detection method based on Topological Data Analysis (TDA). This method
transforms the Euclidean feature space into a topological space. In this case, the forms of noisy
irrelevant documents are much easier to distinguish from thematically relevant documents. This
topological space is organized into a network according to the connectivity of points, that is, documents.
Therefore, we obtain competitive results by filtering data based on the size of connected components.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the authors proposed a method to compute the similarity between two users based on their
profile images using topological data analysis on Twitter. There were used converting of their captions
to numerical vectors usingWord2vec, calculating cosine distances of the caption vectors, and applying
them to the topological data analysis approach using mapper. Using k-means, hierarchical clustering
and topological data analysis, the analysis of the popularity of social media images and correlation
between image captions and the images’ popularity are obtained.
      </p>
      <p>
        The issue of distinguishing two texts is relevant, the distance between them can be used as a criterion,
that is, a metric. The authors [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] provide a first approach to a metric distance between the literary style
of these poets (Conceptismo and Culteranismo as the canon of the baroque Spanish literature) using
TDA techniques. At the beginning, skipgram was used, having the high-dimensional representation of
the words that compose the different sonnets of the dataset following with the Vietoris-Rips filtration
computation. The cosine distance as a measure of the similarity between words by the angle of their
vectors is used as a metric to compute the Vietoris-Rips filtration and applied in the word2vec algorithm.
      </p>
      <p>
        For artificial text detection, the experimental setup also highlights the applicability of the features
towards the TGM architecture, TGM’s size and the decoding method. Notably, the TDA-based
classifiers tend to be more robust towards unseen GPT style TGMs as opposed to the considered
baseline detectors, see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In addition, it is proposed three types of interpretable topological features
that can be derived from the attention maps of any transformer-based LM and shown that the features
capture surface and structural properties, lacking the semantic information.
      </p>
      <p>
        By Wlodek Zadrozny, see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the possibility of capturing ‘the notion of logical shape in text,’ using
TDA was researched. But the main result of this paper stated that there is no clear answer to the
question:”Can we find a circle in a circular argument?”
      </p>
      <p>
        Similarly to bag-of-words, there is persistence bag-of-words, see [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It is a novel and stable
vectorized representation that enables the seamless integration with machine learning. Comprehensive
experiments show that the new representation achieves state-of-the-art performance and beyond in
much less time than alternative approaches.
      </p>
      <p>
        Important aspects of text processing are text classification and visualization [
        <xref ref-type="bibr" rid="ref7 ref8 ref9">7-9</xref>
        ] by TDA. For
example, see [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the authors tried to identify the author of a given text based on its content. They
considered Hafezian poems and poems of Ferdowsi. As result, it proved that it is possible to divide the
number of Hafezian poems from poems of Ferdowsi in mixed data set.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <p>In this section, vectorizing the words, dataset and main terms of persistent homology are considered.</p>
      <p>
        We have to remark that there are several methods for obtaining vector representations for words:
unsupervised learning algorithm (GloVe, see [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]), semi-supervised sequence learning (BERT, ELMo,
ULMFiT, etc.), supervised learning algorithm (Word2Vec, etc) and construction of model
(bag-ofwords model, etc.). All of them are part of natural language processing (NLP), and the reader who is
interested in these techniques in more detail should study the classic textbooks on this topic. A review
of these techniques is not carried out within the scope of this article.
      </p>
      <p>The main disadvantage of all is the fact that the larger the sample (dataset), the better the results.
Moreover, the amount of training data is expressed in thousands of units. That is why the authors
propose the approach described in this section.
3.1.</p>
    </sec>
    <sec id="sec-4">
      <title>Principles of letter coding</title>
      <p>It is known that there are two classes of letters: consonants and vowels in Ukrainian. Every letter 
corresponds to a vector  ⃗( 1,  2,  3,  4), where:
  1 be the ordinal number of the letter in the text;
  2 be the ordinal number of the word in the text which contains letter  ;
  3 be the ordinal number of alphabetical ordered set of vowels. If a letter is consonant,  3 is
equal to a zero;</p>
      <p>  4 be the ordinal number of alphabetical ordered set of consonants. If a letter is vowel,  4 is
equal to a zero.</p>
      <p>For any tongue twister, there is a map into non negative real four-dimensional space ℝ4+.</p>
      <p>Since the mapping is carried out in a four-dimensional space, any visualization is complicated by
the human perception so it is necessary to reduce the dimensions.</p>
      <p>In Fig. 1 and Fig. 2, there are three projections in different spaces of the same tongue twister “Babyn
bib rozcviv u doshch, bude babi bib u borshch”. Common to all of them is the fact that the points are
grouped into certain clusters, which in the future can ensure the presence of a certain cycle in passing
from point (letter) to point (letter) in the space that is coded the tongue twister.
3.2.</p>
    </sec>
    <sec id="sec-5">
      <title>A Dataset</title>
      <p>
        For research, we propose a dataset that consists of 100 Ukrainian tongue twisters, see [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and
nontwisters. We have to remark that there is a repository of textual data in the Ukrainian language named
the UberText corpus. The authors do not use it seeing the texts are preprocessed for searching and their
vectorizations are based on methods (LexVec and GloVe) that do not take into account the shape of the
structure and their success depends on dataset dimension.
      </p>
      <p>In Fig. 3, there are two histograms of dataset parts used in this research. These histograms show
that most tongue twisters have up to 50 letters. They are short and aimed at the quick pronunciation of
sounds.
in the text and a histogram of the non-twisters dataset [right], where an axis X is the quantity of the
letters in the text</p>
      <p>This dataset contains 130 short texts, most of which are sentences. The non-twisters were chosen
arbitrarily, they are meaningful and are parts of some stories.
3.3.</p>
    </sec>
    <sec id="sec-6">
      <title>Persistent homologies</title>
      <p>
        Let consider the terms of computational topology. The exact mathematical definitions are in [
        <xref ref-type="bibr" rid="ref12 ref13">12,13</xref>
        ]
The topological invariant of a space the first Betti number ( 1=rank  1i,j) is the amount of cycles of the
space. For calculating this invariant we used the first persistent homology  1i,j, which is Im 
1
 , for 0≤
i&lt;j ≤ k+1, where  1
      </p>
      <p>, :  1i →  1 , i&lt;j, be a map. On other words,  1i,j= 1i /( 1j ∩  1i ), where  1i is 1-cycles
of    and  1j is 1-boundaries of    (a set {   } =1 of Vietoris-Rips complexes is the filtration for any

finite set { 1,  2, …,   }, where   &lt;  , i&lt;j.). There is a method of their calculation based on the matrices
algebra, the persistence barcode and the persistence diagrams. The 1-cycle is a 1-chain with empty
boundary. The group of 1-cycles is the kernel of the 1-th boundary homomorphism,  1=ker  1. The
1boundary is a 1-chain that is the boundary of a 2-chain. The group of 1-boundaries is the image of the
2-nd boundary homomorphism,  1=Im  1+1. A 1-chain is a formal sum of 1-simplices in a simplicial
complex K and its standard notation is c=∑     , where   is p-simplex in K and   is either 1 or 0.</p>
      <p>For next calculation, we used the GUDHI library, which is a generic open source C++ library with
Python interface, for Topological Data Analysis (TDA) and</p>
      <sec id="sec-6-1">
        <title>Higher Dimensional Geometry</title>
        <p>
          Understanding, see [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          In Fig. 4, there are two diagrams that are constructed for the tongue twister “Babyn bib rozcviv u
and  2 = 0 (Betti numbers), see [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], is topologically equivalent to a circle.
doshch, bude babi bib u borshch”. According to topological aspects, the space that has  0 = 1,  1 = 1
        </p>
        <p>So, if the maximum length of an edge of constructed Vietoris-Rips complex is equal to 0.6 and the
minimum persistence is equal to 0.36 then “Babyn bib rozcviv u doshch, bude babi bib u borshch” can
be considered as a circle.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>4. Algorithm and Experiments</title>
      <sec id="sec-7-1">
        <title>The authors propose the following algorithm:</title>
        <p>Step 1. Every letter of the text is coded according to Sec. 3.1. If the text consists of  symbols
(letters) { 1,  2, … ,   }, then it corresponds to a set  = {( 11,  21,  31,  41), … , ( 1 ,  2 ,  3 ,  4 ) }. In other
words, the text is considered as a cloud of the points in ℝ4+. We also normalize it by standard function
and map a cloud of the points into  4, where  4 = [0; 1]4 be a four-dimensional unit cube. Let denote
it by  ̃.</p>
        <p>Step 2. To construct on  ̃ the filtration    by Vietoris-Rips complexes, where { 1,  2, …,   } is a
finite set,   &lt;  , i&lt;j, and computing  1=rank  1i,j for every fixed   ,  = ̅1̅̅,̅̅. So, then we obtain a set
{ 11,  21, …,  1}.</p>
        <p>Step 3. For every text of dataset, we apply the previous steps and obtain a set  which consists of
vectors with  − coordinates. For set  some methods of classification can be used.</p>
        <p>In Fig.5, there is a pipeline of an algorithm. A coding corresponds to Step 1, Computation  1 – Step
2 and Classifier – Step 3. We remark the following:</p>
        <p> The output of Coding is a cloud of the points into  4, where  4 = [0; 1]4 be a four-dimensional
unit cube.</p>
        <p> The output of Computation  1 is N numbers of ordered sets of k positive numbers.
 The output of Classifier is “Yes or No” answer to “Is this text a tongue twister?”</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>5. Results and discussions</title>
      <p>We have to remark that the structure of the filtrations does not depend on the min_persistence
parameter.</p>
      <p>To construct a classifier we use the parameter test_size = [0.1,0.2, …, 0.9]. In Table 1, the accuracies
of considered classifiers are presented with values of considered parameters.</p>
      <p> Tongue twisters detection is stable with below average score with the following parameters –
Dataset 2 with random sate =1 by Decision tree classifier.</p>
      <p> Tongue twisters detection turned out to be the most unstable with the following parameters –
Dataset 2 with random sate =0 by Decision tree classifier.</p>
      <p>In Fig. 6 and Fig.7, there are two trees for classifying twisters. Both of them are realized on Dataset
2 with test size=10%.</p>
      <p>However, some differences are noticeable:
 For a tree in Fig. 6, there is a parameter max_depth=5. It regularized the model and reduced the
risk of overfitting accordingly. This effect is clearly visible in the accuracy value, it is higher than in
the solution in Fig.7.</p>
      <p> In Fig. 7, the proportion of nodes in which the Gini indicator is close to 0.5 is bigger, which is
not a good indicator for Decision tree classifier.</p>
      <p> In Fig. 7, a tree classifier contains many nodes with 1 text.</p>
      <p> Both trees contain a branch with length less than 20 letters. Moreover, the similarity of parameter
values for their longer classification of such short tongue twisters 80 and 40(45) should be noted. A
small tree (Fig. 6) has only two threshold values for the number of characters that are 35 and 50.
Whereas, a large tree (Fig. 7) generates new nodes with a step of 2-3 characters in the text, which makes
it very difficult for understanding.</p>
      <p>The authors remark that the application of the classical statistical approach with the formulation of
statistical hypotheses on this sample and determination of the level of significance is complicated. Since
the sample can not be expanded to increase the level of significance, due to the presence of established
expressions in the language that are tongue-in-cheek, it can only be expanded with the help of artificial
methods. For example, if you apply to this sample two-sample t-test to estimate the sample sizes of an
experimental group and a control group that are of equal size, then statistical power is equal to 0.25. On
the other hand, using the same assumptions, the presence of more than 1000 instances in the sample does
not guarantee 0.99, since models can have hundreds of thousands of data, but cannot have high training.</p>
      <p>That is why the selection of machine learning models and data mining is an art starting from the
formation of a dataset and ending with the selection of models and their parameters. There are no
theorems and statements that would clearly establish how and on what to teach the module in order to
obtain a high level of recognition.</p>
    </sec>
    <sec id="sec-9">
      <title>6. Conclusion</title>
      <p>An algorithm of tongue twisters detection in Ukrainian based on the consideration of a text as a
sample of points in four-dimensional space, calculation of Betti numbers as a topological invariant of
the existence of one-dimensional circles in it and implementation it into machine learning is proposed.
This approach makes it possible to see behind the text not only a set of points, but also the possibility
of generating a structure in space with a longer calculation of its invariants, the values of which make
it possible to see its geometry (a sentence is as "circle", "sphere" or "torus").</p>
      <p>In addition, there were realized two methods of machine learning as classifiers for detecting tongue
twisters among simple narrative sentences. According to research, more accuracy is provided by
Support Vector Machine Classifier with Gaussian RВF kernel.</p>
      <p>In further research, there are several ways to improve the accuracy of detection, as well as to improve
the process of encoding words or letters taking into account the analysis of the sounds, as well as the
complexity in the pronunciation of sounds and calculating of topological invariants that contain
cyclomaticity of higher orders.</p>
    </sec>
    <sec id="sec-10">
      <title>7. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Torres-Tramón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hromic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Heravi</surname>
          </string-name>
          ,
          <article-title>Topic detection in twitter using topology data analysis</article-title>
          ,
          <source>in: Proceedings of 15th International Conference</source>
          , ICWE 2015 Workshops,
          <string-name>
            <surname>NLPIT</surname>
          </string-name>
          , PEWET, SoWEMine, Rotterdam, The Netherlands,
          <year>2016</year>
          , pp.
          <fpage>186</fpage>
          -
          <lpage>197</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -24800- 4_
          <fpage>16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Almgren</surname>
          </string-name>
          ,
          <article-title>Employing topological data analysis on social networks data to improve information diffusion (Computer Science)</article-title>
          ,
          <source>Ph.D. thesis</source>
          , the school of engineering university of Bridgeport, Connecticut, USA,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Paluzo-Hidalgo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gonzalez-Diaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Gutierrez-Naranjo</surname>
          </string-name>
          ,
          <article-title>Towards a philological metric through a topological data analysis approach</article-title>
          ,
          <source>ArXiv</source>
          ,
          <year>2019</year>
          , abs/
          <year>1912</year>
          .09253.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Kushnareva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cherniavskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mikhailov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Artemova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Barannikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Piontkovskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Piontkovski</surname>
          </string-name>
          , E. Burnaev,
          <article-title>Artificial text detection via examining the topology of attention maps</article-title>
          ,
          <source>in: Proceedings of Conference on Empirical Methods in Natural Language Processing</source>
          , Punta Cana, Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>635</fpage>
          -
          <lpage>649</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .emnlpmain.
          <volume>50</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>W.</given-names>
            <surname>Zadrozny</surname>
          </string-name>
          ,
          <article-title>A note on argumentative topology: circularity and syllogisms as unsolved problems, arXiv - CS - Computation and</article-title>
          <string-name>
            <surname>Language</surname>
          </string-name>
          ,
          <year>2021</year>
          . doi:arxiv-
          <volume>2102</volume>
          .
          <fpage>03874</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Zielinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lipinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Juda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zeppelzauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dłotko</surname>
          </string-name>
          ,
          <article-title>Persistence bag-of-words for topological data analysis</article-title>
          ,
          <source>in: Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, Macao</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4489</fpage>
          -
          <lpage>4495</lpage>
          . doi:
          <volume>10</volume>
          .24963/ijcai.
          <year>2019</year>
          /624.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Elyasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Moghadam</surname>
          </string-name>
          ,
          <article-title>An introduction to a new text classification and visualization for natural language processing using topological data analysis, 2019, arXiv</article-title>
          . URL: https://arxiv.org/abs/
          <year>1906</year>
          .01726.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Yue</surname>
          </string-name>
          ,
          <article-title>Topological data analysis of two cases: text classification and business customer relationship management</article-title>
          ,
          <source>J. of Physics</source>
          ,
          <volume>1550</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .1088/
          <fpage>1742</fpage>
          -6596/1550/3/032081.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Sh</surname>
          </string-name>
          .
          <string-name>
            <surname>Gholizadeh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Savle</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Seyeditabari</surname>
          </string-name>
          , W. Zadrozny,
          <article-title>Topological data analysis in text classification: extracting features with additive information</article-title>
          , arXiv,
          <year>2020</year>
          . URL: https:// https://arxiv.org/abs/
          <year>2003</year>
          .13138.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>GloVe: Global Vectors for Word Representation</article-title>
          . URL: https://nlp.stanford.edu/projects/glove/.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[11] Top 100 Ukrainian tongue twisters by Leopolis published 17.11</source>
          .
          <year>2021</year>
          URL: https://lviv1256.com/lists/top-100
          <string-name>
            <surname>-</surname>
          </string-name>
          ukrajinskyh-skoromovok/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>G. Carlsson,</surname>
          </string-name>
          <article-title>Topology and data</article-title>
          ,
          <source>Bull.Amer.Math.Soc</source>
          ,
          <volume>46</volume>
          (
          <issue>2</issue>
          ) (
          <year>2009</year>
          ):
          <fpage>255</fpage>
          -
          <lpage>308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>I. Yurchuk</surname>
          </string-name>
          ,
          <article-title>Digital image segmentation based on the persistent homologies</article-title>
          ,
          <source>in: Proceedings of the 1st International Workshop on Information-Communication Technologies and Embedded Systems</source>
          , ICTES, Mykolaiv, Ukraine,
          <year>2019</year>
          , pp.
          <fpage>226</fpage>
          -
          <lpage>232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <article-title>The GUDHI library</article-title>
          . URL:https://gudhi.inria.fr/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>