<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Multifactor Method Applied for Authorship Attribution on the Phonological Level</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Start</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>Lviv 79013</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The multifactor method has been developed to enhance test validity of authorship attribution. The method is style based. An author of a text can be identified by three major factors: a style based factor, a topic based factor and an authorial style based factor. For each factor certain statistical parameters are determined. The statistical parameters are actual distributions of frequencies of occurrence of the researched language units. As the research is done on the phonological level, the language units are phonemes. To differentiate texts by different authors, the powerful statistical tests have been applied (the KolmogorovSmirnov's test, the chi-square test, the Student's t-test).</p>
      </abstract>
      <kwd-group>
        <kwd>Phoneme Group</kwd>
        <kwd>Multifactor Method</kwd>
        <kwd>Style Based Factor</kwd>
        <kwd>Topic Based Factor</kwd>
        <kwd>Authorial Style Based Factor</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Anonymous information in the Internet has always been an important problem for
researchers. The necessity to solve this problem is growing as the number of people using
this network is increasing. The anonymous texts occur in different areas of
communication. In certain cases the information doesn’t bother anyone and can be negligible.
But it does disturb when it threatens, harasses and is inappropriate. If the reader
considers the information personal and sensitive, the anonymous or given under a
pseudonym information must be studied. It should be noted that plagiarism detection is closely
connected with disputed authorship. It is particularly important to determine the real
author of a text of the same or similar content, but having several authors. In certain
cases it is necessary to establish the author of the text written a long time ago, not clear
when. The task of determining an author involves some text categorization and
classification. The choice of most optimal classifiers and feature sets depends on various
factors. The size of the text under study is of importance. The method efficient for the
short texts may be inefficient for long texts. Each language level and feature set on this
level have their specificity. The methods selected for a certain language level and
feature sets are sure to behave differently. The number of features in a set may be increased
or decreased, depending on the applied method. The problem of authorship attribution
implies inferring a certain style, a certain topic and an authorial style. A text by an
author can be of different complexity. The simplest case is when two studied texts are
of the same style, genre and topic. The most complex case is when the texts are of
different style, genre and topic. There also intermediate cases with greater or less
similarity or difference. The three mentioned factors lie in the basis of the developed
multifactor method. Each factor effect is determined by three efficient statistical methods
– the Kolmogorov-Smirnov’s test, the chi-square test and the Student’s t-test. The
purpose of the research is to enhance test validity of authorship attribution with the help of
the multifactor method. According to the results of recent research, different authorship
attribution approaches have been used. Thus, in the field of digital text forensics,
informal chat conversations have been researched. The algorithmic solutions have been
obtained with 72,7%, 75% accuracy [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. The problem of author identification in short
texts of Internet communication has been studied. In this research the temporal changes
of word usage are relevant [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The analysis of principle components for authorship
identification has been conducted in business systems research [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The multi sequence
word selection method has been chosen to determine the author of a text [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Method
of similar textual content selection based on thematic information retrieval has been
applied for an analysis of the text under study [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The quantitative methods have been
used to study lexical and stylistic peculiarities of a text [
        <xref ref-type="bibr" rid="ref10 ref7 ref8 ref9">7 – 10</xref>
        ]. The recurrent neural
networks have been used to model the flow of the text for authorship attribution. For
this study a large corpus has been recommended [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The unmasking approach has
been used in the forensic field for short texts – four pages. The accuracy is 75%, 80%
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Large candidate sets have been researched by machine learning techniques. This
is a novel approach, as the previous one studied a limited number of candidates [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
The Twitter site has been analyzed for stylometric features. The author for an
illegitimate text has been inferred [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Similarity-based methods have been used to consider
authorship attribution in the wild. Anonymous texts have been analyzed on the lexical
level [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In the investigation conducted by the support vector machine classifier, good
results have been obtained – around 95% on the feature set of bag of words [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In
comparison with the mentioned research, the novel approach in this study consists of
applying the proposed combination of the three methods: the Kolmogorov-Smirnov’s
test, the chi-square test and the Student’s t-test which have proved efficient in
authorship attribution. To maximize the accuracy, the language level with an unchangeable
number of elements has been chosen – the phonological level. The success rate is 95%,
97% and 98% [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Mathematical Support of Software System</title>
      <p>
        The Method Developed
The problem of authorship attribution is aimed at determining if two compared pieces
of text were written by a single author. Texts from poetry (G. Byron, T. Moore) and the
publicist style (B. Оbama, D. Trump, D. Webster, S. Logan) have been selected for
experiments. An author of a text can be revealed by three major factors: a style based
factor, a topic based factor and an authorial style based factor. The style based factor
consists of showing the difference between the two styles, the topic based factor – the
difference between the two texts on different topic, the authorial style based factor –
the difference between the two texts by different authors. The scheme is style – topic –
author. The average value of the three factor based values is considered general style
markedness of a text. The steps of the multifactor method algorithm are given below.
More detailed information regarding the steps of the Student’s t-test, the
KolmogorovSmirnov’s test and the chi-square test was presented in the previous research [
        <xref ref-type="bibr" rid="ref18 ref20">18, 20</xref>
        ].
      </p>
      <p>
        1) the Student’s t-test is performed for the texts from different styles, on different
topics and by different authors [
        <xref ref-type="bibr" rid="ref18 ref19 ref20">18, 19, 20</xref>
        ]:
If n,m   , the texts can be differentiated:
n,m 
      </p>
      <p>nm
n  m</p>
      <p>Dn,m 
nm</p>
      <p>sup Fn (z)  Fm (z) .</p>
      <p>n  m z
a a
t  x1  x2</p>
      <p>S
n1  n2 ,
n1  n2
Dn,m  sup Fn (z)  Fm (z) ,</p>
      <p>z
a a
where x1  x 2 is a difference of average frequencies for two samples for the fixed
group of consonants a, S is a variance, n is a sample size.</p>
      <p>
        2) the Kolmogorov-Smirnov’s test is performed for the texts from different styles,
on different topics and by different authors [
        <xref ref-type="bibr" rid="ref18 ref21 ref22 ref23">18, 21, 22, 23</xref>
        ]:
where Fn z and Fm z are two empirical distribution functions for n and m. Samples.
(1)
(2)
(3)
(4)
3) the chi-square test is performed for the texts from different styles, on different
topics and by different authors [
        <xref ref-type="bibr" rid="ref24 ref25 ref26 ref27">24, 25, 26, 27</xref>
        ]:
      </p>
      <p>s
ˆ n2  
i1 j1</p>
      <p>
k  i, j </p>
      <p>2
n j j </p>
      <p> k
n  ,  j   ij
j1
where  i, j is a realization number in j-th series, s is a number of consonant groups,
k is a number of samples, n j is a number of sample portions, n is a number of portions
for two samples. The texts can be differentiated if ˆ n2   12 ,(s1)(k1) .
a a
4) determining the style based factor value   1 for a phoneme group a: x1s  x 2 s (s
is a style);
a a
5) determining the topic based factor value   2 : x1t  x 2 t ( is a topic);
a a
6) determining the authorial style based factor value   3 ∶ x1а  x 2 а (а is an author);
7) determining the general style markedness:
8) the authorship attribution is calculated by the difference of values of the general
style markedness for two authors.
To develop the software for authorial differentiation, the Java programming language
has been used. The programming language is cross-platform and this is an advantage
of the chosen programming language. The developed program system realizes the
following algorithm (Fig. 1):</p>
      <p>The structure of the developed software has the following tabs: “Text”,
“Transcription Symbols”, “Consonant Phoneme Sample”, “Portion Division”, “Group Division”,
“Calculating Phonemes in Portions”, “Calculating Phonemes in Groups”, “Statistical
Test”, “Style Based Factor Value”, “Topic Based Factor Value”, “Authorial Style
Based Factor Value”, “General Style Markedness Values”, “Difference by General
Style Markedness Values”.</p>
      <p>The software classes are shown in the diagram s in Fig. 2.</p>
      <p>One of the advantages of this program system is its relative independence of the
transcription site on which an English text is transcribed. The transcription site is used
when the first experiments are made. The bag of words gets larger every time new texts
are processed. Therefore, it is advisable to process large samples. These are some short
documents in the forensic field. However, large samples are of interest when it is
necessary to characterize literary legacy of some author. Such problems are usually
researched in corpus linguistics. In this investigation, both short and long texts are
analyzed. The sample size is 50 000 phonemes and more.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results of the Study</title>
      <p>For the first experiment of author identification, two pieces of poetry have been
selected: one by G. Byron and another by T. Moore. According to the proposed scheme,
the two pieces of poetry must be analyzed in a comparison with some style having most
common bag of words. Evidently, this may be the conversational style in its literary
version which has few colloquial elements. This is particularly relevant in the
comparison with poems by romanticists who tried to use conversational elements. On the other
hand, it is necessary to compare the poems with another genre of the style of fiction. It
may be Byron’s emotive prose. These two samples are sure to have common language
units both being of fiction style. In the third stage of the study, the poems by the two
poets are compared. The multifactor method makes it possible to calculate the average
value of the values calculated for each mentioned above comparisons. The average
value is general style markedness (Table 1).</p>
      <sec id="sec-3-1">
        <title>Forming a sample of a transcribed text</title>
      </sec>
      <sec id="sec-3-2">
        <title>Performing interval division Calculating an average frequency value</title>
      </sec>
      <sec id="sec-3-3">
        <title>Performing theoretical normal distribution Calculating a theoretical frequency</title>
      </sec>
      <sec id="sec-3-4">
        <title>Performing the Pearson’s test for two texts</title>
      </sec>
      <sec id="sec-3-5">
        <title>Performing the Student’s t-test for two texts</title>
      </sec>
      <sec id="sec-3-6">
        <title>Performing the Kolmogorov-Smirnov’s test for two texts</title>
      </sec>
      <sec id="sec-3-7">
        <title>Calculating the style based factor value for two texts under study</title>
      </sec>
      <sec id="sec-3-8">
        <title>Calculating the topic based factor value for two texts under study</title>
      </sec>
      <sec id="sec-3-9">
        <title>Calculating the authorial style based factor value for two texts under study</title>
      </sec>
      <sec id="sec-3-10">
        <title>Determining general style markedness values</title>
      </sec>
      <sec id="sec-3-11">
        <title>Difference by general style markedness values End</title>
        <p>Having calculated the value of general style markedness for each poet, the author
identification test can be performed. The difference of the two authorial styles is calculated
by the difference of values general style markedness (Table 2). It is equal to 1.8.</p>
      </sec>
      <sec id="sec-3-12">
        <title>C ConsonantType</title>
        <p>m getNames()
1
1
C ConsonantUtils
m countConsonantTypes(String) Map&lt;ConsonantType, Long&gt;
m countConsonants(String) Map&lt;String, Long&gt;
m countConsonantTypes(Map&lt;String, Long&gt;)
Map&lt;ConsonantType, Long&gt;</p>
      </sec>
      <sec id="sec-3-13">
        <title>C ConsonantProcfssor</title>
        <p>m processConsonants(String, String)
MultiValueMap&lt;ConsonantType, Long&gt;
C PropertyUtils
m getltem(String)
m saveltem(String, String)
m builder()
+</p>
        <p>create
C Builder
m path(String) Builder
m build() PropertyUtils</p>
      </sec>
      <sec id="sec-3-14">
        <title>String</title>
        <p>void
Builder
ϟ NoSuchConsonantException
m NoSuchConsonantException()
m NoSuchConsonantException(char)
Comparison with the Comparison with Comparison with Value of general
1-st style the 2-nd style another author style markedness
GB-СLS – 16 GB-GBP – 17 GB-TM – 7 GB – 13.3
TM-СLS –15 TM-GBP –13 TM-GB –7 TM –11.5
In Table 1 the following designations are used: GB is Byron’s poetry, TM is Moore’s
poetry, GBP is Byron’s emotive prose, СLS is the conversational style.
For the second experiment four authors have been selected. They are: B. Оbama, D.
Trump, D. Webster, S. Logan. The pieces of writing represent the publicist style. In this
case the multifactor method involves comparison with the 1-st author, comparison with
the 2-nd author and comparison with the 3-rd author. The average value calculated from
the three values got in each comparison is the general style markedness by which the
author can be identified. Though the samples are of the same style, the topic varies from
sample to sample having common bag of words. The topic reflects international
political events all over the world. Therefore, the content is relatively homogeneous. This
relative homogeneity creates a problem of its own. The topic based factor affects the
final result of author identification. Sometimes it is rather difficult to draw a distinct
demarcation line between effect of the topic based factor and the authorial style based
factor. The case is easier with documents following strict standards of conveying
information. The same situation can hardly be observed in the publicist style. On the other
hand, the publicist style is the style in which the individual peculiarities of an author’s
manner of writing can be vividly revealed. The effect of the three factors mentioned is
expressed in the value of the general style markedness given in Table 3.
The value of general style markedness is the highest for Webster’s authorial style. It
equals to 16. The lowest value is for Obama’s writing style (14). But, as a significant
role is played here by the topic based factor, the authorial writing characteristics can be
different in another verbal content. Thus the essential information about author
identification can be got with the help of the multifactor method. The results of comparisons
of texts by different authors are shown in Table 4.
Table 4 shows that different effect of the topic and author based factors causes great
difference in a comparison Оbama-Webster, less difference – Оbama-Trump,
ОbamaLogan, still less difference – Trump-Webster, Webster-Logan and practically no
difference – Trump-Logan. The last pair of texts shows similarity of bag of words.</p>
        <p>Among the used statistical tests two are the most powerful. These are the
Kolmogorov-Smirnov’s test and the chi-square test. With the help of the former, the authorial
styles differ essentially in all eight phoneme groups. The latter is a little less powerful
– the difference in six of eight groups.</p>
        <p>The efficiency of the multifactor method may be analyzed for each of eight groups
of phonemes. Reduction of the number of phoneme groups makes the whole procedure
more economical. Consequently, it is necessary to analyze author-differentiating
capability for every group.</p>
        <p>The degree of author-differentiating capability of phoneme group depends on the
number of times the essential differences have been established. If the essential
difference has been revealed by three statistical tests, the phoneme group takes number 3, by
two statistical tests – number 2, by one statistical test – number 1. In Table 5, for labial
group, number 3 (Оbama-Logan) has been got once, number 1 – thrice. The group takes
the second degree of differentiation power.
Compared with labial group, dorsal group has higher degree of differentiating
capability (Table 6). Only in two pairs of texts one statistical test has proved efficient. More
differences have been obtained by two tests. The group takes the first degree of
differentiation power.</p>
        <p>The results of the conducted research have shown that the multifactor method is
efficient in authorship attribution. The established general style markedness of a text has
made it possible to classify each sample under study in accordance with three basic
factors – style, topic and author’s manner of writing. Taking into account the three
mentioned factors is particularly efficient when the compared samples represent
different style and topic. In this case it is impossible to characterize the authorial specificity
of writing because of the influence of style and topic factors. Having determined the
style based and topic based features, the authorial features can be identified.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>In order to single out particular features of an individual writing style, the style based
features and topic based features must be separated. To solve this task, the multifactor
method must be applied. In accordance with this method, the average value of the three
factor based values is calculated. The three factor based values involve: comparison
Оbama-Trump
Оbama-Webster
Оbama-Logan
Trump-Webster
Trump-Logan
Webster-Logan
Byron-Moore
1
1
3
2
2
2
2
with the text least marked by the style elements, comparison with the text of the same
topic, but different author and comparison with the text of the same style and topic, but
another author. The average value of these three values is general style markedness.
The author identification is calculated by the difference of general style markedness
values. The results show that the greatest difference is in the pair Obama – Webster (2),
less difference – in the pairs Obama – Trump, Obama – Logan (1.6), still less – in the
pairs Trump – Webster, Webster – Logan (0.4), the least in the pair Trump – Logan.
The test validity has been enhanced up to 95%, 97%.</p>
      <p>The developed software on the Java programming language has performed the
author identification procedure in a fewer number of consonant groups making it more
automated. The next step in our research will be concentrated on the other statistical
methods.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Halvani</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winter</surname>
          </string-name>
          , Ch.,
          <string-name>
            <surname>Graner</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Assessing the Applicability of Authorship Verification Methods</article-title>
          .
          <source>In: Proceedings of the 14th International Conference on Availability, Reliability and Security</source>
          , No.:
          <volume>38</volume>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . (
          <year>2019</year>
          ). https://doi.org/10.1145/3339252.3340508.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
          </string-name>
          , Sh.: Authorship Attribution:
          <article-title>What's Easy and What's Hard</article-title>
          ? In: Computer Science, (
          <year>2013</year>
          ). DOI:
          <volume>10</volume>
          .2139/ssrn.2274891.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Azarbonyad</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dehghani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marx</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamps</surname>
          </string-name>
          , J.:
          <article-title>Time-Aware Authorship Attribution for Short Text Streams</article-title>
          .
          <source>In: Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , New York, USA, pp.
          <fpage>727</fpage>
          -
          <lpage>730</lpage>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jamak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Can</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Principal Component Analysis for Authorship Attribution</article-title>
          .
          <source>In: Business Systems Research</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>49</fpage>
          -
          <lpage>56</lpage>
          . (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mubin</surname>
            ,
            <given-names>S. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rajesh</surname>
            <given-names>S. P.</given-names>
          </string-name>
          :
          <article-title>Authorship Identification with Multi Sequence Word Selection Method</article-title>
          .
          <source>In: Thermal Stresses-Advanced Theory and Applications</source>
          , pp.
          <fpage>653</fpage>
          -
          <lpage>661</lpage>
          , (
          <year>2019</year>
          ).
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -16657-1_
          <fpage>61</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Vysotska</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lytvyn</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalchuk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kubinska</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dilai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rusyn</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pohreliuk</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chyrun</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chyrun</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brodyak</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Method of similar textual content selection based on thematic information retrieval</article-title>
          .
          <source>In: CSIT, Proceedings of the XIVth Scientific and Technical Conference</source>
          , Lviv, pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kulchytskyi</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shandruk</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>The quantitative research of scientific texts at the symbolic level</article-title>
          . In:
          <article-title>Computational linguistics and intelligent systems</article-title>
          . Lviv: Lviv Polytechnic National University,
          <fpage>25</fpage>
          -
          <lpage>27</lpage>
          June, vol
          <volume>2</volume>
          , pp.
          <fpage>71</fpage>
          -
          <lpage>80</lpage>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Karamysheva</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nazarchuk</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fedoruk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Synonymic connections of cognitive verbs in English and Ukrainian languages: applied aspect</article-title>
          .
          <source>In: CSIT, Proceedings of the XIIIth Scientific and Technical Conference</source>
          . Lviv, pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Romanyshyn</surname>
          </string-name>
          , N.:
          <article-title>Application of computer technologies in conceptual analysis</article-title>
          .
          <source>In: CSIT, Proceedings of the XIIIth Scientific and Technical Conference</source>
          . Lviv, pp.
          <fpage>55</fpage>
          -
          <lpage>57</lpage>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Peleshchyshyn</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markovets</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vus</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albota</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Identifying specific roles of users of social networks and their influnce methods</article-title>
          .
          <source>In: CSIT, Proceedings of the XIIIth Scientific and Technical Conference</source>
          . Lviv, pp.
          <fpage>39</fpage>
          -
          <lpage>42</lpage>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bagnall</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Author Identification Using Multi-headed Recurrent Neural Networks</article-title>
          .
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          , Toulouse, France, pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Bevendorff</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Generalizing Unmasking for Short Texts</article-title>
          . In:
          <article-title>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</article-title>
          . Minneapolis, Minnesota, vol.
          <volume>1</volume>
          , pp.
          <fpage>654</fpage>
          -
          <lpage>659</lpage>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
          </string-name>
          , Sh.,
          <string-name>
            <surname>Winter</surname>
          </string-name>
          , Ya.:
          <source>The “Fundamental Problem” of Authorship Attribution</source>
          , vol.
          <volume>93</volume>
          , issue 3, pp.
          <fpage>284</fpage>
          -
          <lpage>291</lpage>
          . (
          <year>2012</year>
          ) DOI:
          <fpage>10</fpage>
          .1080/0013838X.
          <year>2012</year>
          .668794
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Bagnall</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Author Identification Using Multi-headed Recurrent Neural Networks</article-title>
          . In: Conference and
          <article-title>Labs of the Evaluation forum</article-title>
          , Toulouse, France, pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Bhargava</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mehndiratta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asawa</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Stylometric Analysis for Authorship Attribution on Twitter</article-title>
          .
          <source>In: Proceedings of the Second International Conference on Big Data Analytics</source>
          , vol.
          <volume>8302</volume>
          ,. pp.
          <fpage>37</fpage>
          -
          <lpage>47</lpage>
          . (
          <year>2013</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -03689-
          <issue>2</issue>
          _
          <fpage>3</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
          </string-name>
          , Sh.:
          <article-title>Authorship attribution in the wild</article-title>
          .
          <source>In: Language Resources and Evaluation</source>
          , vol.
          <volume>45</volume>
          , No.
          <volume>1</volume>
          , (
          <year>2011</year>
          ). URL: https://doi.org/10.1007/s10579-009- 9111-2
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Bozkurt</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Baghoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Uyar</surname>
          </string-name>
          , E.:
          <article-title>Authorship attribution</article-title>
          .
          <source>In: 22nd International Symposium on Computer and Information Sciences (ISCIS)</source>
          , pp.
          <fpage>158</fpage>
          -
          <lpage>162</lpage>
          . (
          <year>2007</year>
          ). DOI:
          <volume>10</volume>
          .1109/ISCIS.
          <year>2007</year>
          .
          <volume>4456854</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Khomytska</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teslyuk</surname>
          </string-name>
          , V.:
          <article-title>Statistical Models for Authorship Attribution</article-title>
          .
          <source>In: Advances in Intelligent Systems and Computing III / Natalia Shakhovska editor, Lviv,</source>
          . vol.
          <volume>1080</volume>
          . pp.
          <fpage>579</fpage>
          -
          <lpage>592</lpage>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Gomez</surname>
          </string-name>
          , P., C.:
          <article-title>Statistical Methods in Language and Linguistic Research</article-title>
          . Spain: Unibersity of Murcia, (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Khomytska</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teslyuk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kryvinska</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beregovskyi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The Nonparametric Method for Differentiation of Phonostatistical Structures of Authorial Style</article-title>
          .
          <source>In: Procedia Computer Science: Proceedings of the 10th International Conference on Emerging Ubiquitous Systems and Pervasive Networks, Coimbra, Portugal</source>
          , vol.
          <volume>160</volume>
          , pp.
          <fpage>38</fpage>
          -
          <lpage>45</lpage>
          . (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Kolmogorov</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          :
          <article-title>Мathematics аnd its Historical Development Edited by V. A</article-title>
          .
          <string-name>
            <surname>Uspensky</surname>
          </string-name>
          , Published by Nauka, Moscow (
          <year>1991</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Gnedenko</surname>
            ,
            <given-names>B. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolmogorov</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          :
          <article-title>Limit Distributions for Sums of Independent Variables Published by Addison-Wesley (</article-title>
          <year>1968</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Watanabe</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <source>Probability Theory and Mathematical Statistics</source>
          . Springer (
          <year>1988</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Gries</surname>
          </string-name>
          , Th. S.:
          <article-title>Statistics for Linguistics with R: A Practical Introduction (Trends in Linguistics: Studies &amp; Monographs)</article-title>
          ,
          <source>р. 348</source>
          . (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Rozanov</surname>
            ,
            <given-names>Iu. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silverman</surname>
            ,
            <given-names>R. A.</given-names>
          </string-name>
          :
          <article-title>Probability Theory: A Concise Course Dover Publications Inc</article-title>
          . (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Jorgensen</surname>
            ,
            <given-names>P.E.T.</given-names>
          </string-name>
          : Аnalysis and Рrobability. Springer (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Bhattacharya</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waymire</surname>
            ,
            <given-names>E. C.</given-names>
          </string-name>
          : A Basic Course in Probability Theory Springer; 2nd ed.
          <source>2016 edition, February</source>
          <volume>16</volume>
          , (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>