<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Method of fuzzy analysis of texts and their rubrics actualization</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>V V Borisov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M I Dli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>P Yu Kozlov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Engineering Department, The Branch of National Research University “Moscow Power Engineering Institute” in Smolensk</institution>
          ,
          <addr-line>Smolensk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>259</fpage>
      <lpage>263</lpage>
      <abstract>
        <p>The work deals with the offered method of fuzzy analysis of texts and their rubrics actualization. The method is oriented to analyze electronic nonstructural texts of not big size in the following conditions: first, nonstationary composition and the importance of the keywords of the rubric field, second, in the absence or weak stucturization of these texts, third, if there are grammar or syntaxes inaccuracy and errors. The offered method is based on the original approach to the identification of the degree of the texts words fuzzy correspondence according to the well-founded set of syntactical characteristics with subsequent finding the degrees of text documents fuzzy correspondence to all rubrics. The method also allows to carry out monitoring of changes and actualization of rubrics according to the results of checking the formulated conditions of rubric field changes for the following typical situations: formation of the additional rubrics on the “boundary” of the existing rubrics; rubrics division, creating new rubrics, rubrics exclusion, rubrics combining. The offered method allows to raise the accuracy of analysis and the quality of texts classification at the expense of using the fuzzy approach for the accounting of analysis conditions uncertainty and nonstationarity of thesaurus of these texts as well as at the expense of operational actualization of rubrics depending on the composition and importance of the rubrics key words.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
relatively small size of such texts;
such texts weak structuredness or no structuredness at all (no marking and fields for computer
processing);
presence of grammar and syntaxes inaccuracy and errors;
analysis conditions uncertainty and nonstationarity of composition and importance of rubric
field key words;
high degree of rubrics interdependency.</p>
      <p>
        These features put considerable limitations on the usage of traditional models and methods of
morphological, syntaxes and semantic analysis of the texts. However, famous models and methods of
knowledge acquisition from the text information take the requirements of operational rubric changes
into account not sufficiently, this leads to the growth of the number of errors because of the wrong
classification of the processing texts [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7">1–7</xref>
        ].
      </p>
      <p>Therefore, the actual problem is to make a method of fuzzy analysis of electronic nonstructural
texts and actualization of rubrics taking into account the detection of the following situations requiring
operational changes of the rubric field: the additional rubrics formation on the “boundary” of the
already exiting rubrics, rubrics division, creating new rubrics, rubrics exclusion, combining rubrics.</p>
      <p>The offered method of texts fuzzy analysis and rubrics actualization includes the following main
stages:</p>
      <p>Stage 1. Rubric tasks and texts presentation on the basis of the detected syntaxes characteristics.</p>
      <p>Stage 2. Texts analysis on the basis of the degree defining of their fuzzy correspondence to the
rubrics.</p>
      <p>Stage 3. Checking the rubric field changes conditions and rubrics of field actualization according to
the results of this checking.</p>
      <p>Let us consider the problems solving on the stages of the offered method in more details.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Rubric tasks and text presentation on the basis of the detected syntaxes characteristics</title>
      <p>On the basis of the preliminary texts analysis the initial rubric multitude is given:</p>
      <p>
        R  {Rj | j 1..J},
where for all j 1..J Rj   wjm , rjm | m 1..M j , w jm – m- word in the rubric R j , rjm [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] – the
degree of correspondence of the word w jm to rubric R j .
      </p>
      <p>
        For such texts presentation the «unification» of the set of the following syntaxes characteristics,
detected, for example, by analyzer LinkGrammar is done ([
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]):
      </p>
      <p>
        S  {sn | n 1..N}, then N  5 ,
where s1 – the root word or predicate; s2 – the subject; s3 – the adverbial modifier; s4 – the subject
under action; s5 –the predicate [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>The texts multitude is presented in the form of:</p>
      <p>
        SD  {SDk | k 1..K},
where SDk  { SDn(k) | n1..N}, SDn(k) – word multitude of k- text, corresponding the syntaxes
parameter sn.
3. Texts analysis on the basis of the identification of the degree of their fuzzy correspondence to
the rubrics
First, the degrees of fuzzy correspondence  jn (SDn(k) ) [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] relative to syntaxes characteristics SDn(k)
to all rubrics are determined:
1 L(nk)
j  J ,  jn (SDn(k) )  L(k) u(jpk) , n 1..N.
      </p>
      <p>n p1
where u(jpk) – the degree of correspondence of p-word from SDn(k) , primarily given for this word from
rubric R j .</p>
      <p>To determine the degree of the text fuzzy correspondence to the rubrics let us introduce the
parameter  (SDk , Rj ) characterizing the degree of text SDk fuzzy correspondence to rubric Rj:
 (SDk , Rj )  1 
1
N</p>
      <p>
        N
n1  Rj (Rjn )   Rj n (SDn(k) )2 ,
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. For the case under consideration R%j  1 / s1 , 1 / s2 , 1 / s3 , 1 / s4 , 1 / s5  , i.e.
j  J , 1%(SDk , Rj )  1 
1
N
      </p>
      <p>N
1   R n (SDn(k) )2 .
n1 j</p>
      <p>Text SDk refers, in the greatest degree, to that rubric Rl for which the degree of correspondence
is maximum:</p>
      <p>Rl : mj1a..xJ 1%(SDk , Rj ).</p>
    </sec>
    <sec id="sec-3">
      <title>4. Checking the conditions of rubric field changes and rubrics of field actualization in accordance with the results of this checking</title>
      <p>To check the conditions of changing of rubric field let us introduce additionally the following
parameters:
j  J , 0±,5 (SDk , Rj )  1 
1
N</p>
      <p>N
0,5   R n (SDn(k) )2 ,
n1 j
j  J</p>
      <p>j  J , 0%(SDk , Rj )  1  1%(SDk , Rj ),
where parameter 0±,5 (SDk , Rj ) characterizes the degree of uncertainty of text SDk referring to rubric
R j and parameter 0%(SDk , Rj ) characterizes the degree of text SDk discrepancy to rubric R j .</p>
      <p>This stage realization considers the calculation of parameters 1%(SDk , Rj ) , 0±,5 (SDk , Rj ) ,
0%(SDk , Rj ) for all texts and their analysis, according to the results of the analysis on the basis of the
conditions given bellow, the revision of composition and rubric field structure is performed.</p>
      <p>Let us consider the formulated conditions of detection and revision rules of composition and rubric
field structure for the following basic situations: additional rubric formation, rubric division, new
rubrics creation, rubric exclusion, rubric combining.</p>
      <sec id="sec-3-1">
        <title>4.1. Additional rubric formation</title>
        <p>The basis for the additional rubric formation on the «boundary» of the already existed rubrics Ri and
R j is the revealing of the considerable amount of texts (equal to the rubrics or more than the number
of rubrics), for every of which the following condition is valid:
  1%(SDk , Ri )      1%(SDk , Rj )   
 0±,5 (SDk , Ri ) 
 0±,5 (SDk , Rj ) </p>
        <p>
          
  0%(SDk , Ri )       0%(SDk , Rj )   
Rl  R, l  i  j : 1%(SDk , Rl )     0±,5 (SDk , Rl )    0%(SDk , Rl )  ,
where α and β – the upper and the lower boundary values (usually, α = 0.4 and β = 0.7 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]), defining
the reasonability of rubric field revision.
        </p>
        <p>When revealing the number of texts equal to the rubrics or more than the number of rubrics, for
which the above mentioned condition is performed, the conclusion about the reasonability of
additional «boundary» rubric is made.</p>
      </sec>
      <sec id="sec-3-2">
        <title>4.2. Rubric division</title>
        <p>The basis for the rubric division R j is the revealing of the considerable number of texts for each of
which the following condition is fulfilled:</p>
        <p>  1%(SDk , Rj )    0±,5 (SDk , Rj )     0%(SDk , Rj )   
Rl  R, l  j : 1%(SDk , Rl )    0±,5 (SDk , Rl )   0%(SDk , Rl )  ,
where j – the number of the divided rubric.</p>
      </sec>
      <sec id="sec-3-3">
        <title>4.3. New rubric creation</title>
        <p>The basis for the new rubric creation is the revealing of the considerable number of texts for each of
which the following condition is fulfilled:</p>
        <p>Rl  R : 1%(SDk , Rl )    0±,5 (SDk , Rl )   0%(SDk , Rl )  .</p>
      </sec>
      <sec id="sec-3-4">
        <title>4.4. Rubric exclusion</title>
        <p>The basis for the rubric exclusion is the revealing of the considerable number of texts for each of
which the following condition is fulfilled:</p>
        <p>1%(SDk , Rj )    0±,5 (SDk , Rj )   0%(SDk , Rj )  .</p>
      </sec>
      <sec id="sec-3-5">
        <title>4.5. Rubrics combining</title>
        <p>The basis for the rubric Ri and R j combining is the revealing of the considerable number of texts for
which the following condition is fulfilled:</p>
        <p>1%(SDk , Ri )   1%(SDk , Rj )  
 0±,5 (SDk , Ri )   0±,5 (SDk , Rj )  
 0%(SDk , Ri )     0%(SDk , Rj )   
Rl  R, l  i  j : 1%(SDk , Rl )     0±,5 (SDk , Rl )    0%(SDk , Rl )  ,
where Ri and R j – combining rubrics.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Experimental results</title>
      <p>
        The offered method was used in Administration of Smolensk region when automated analysis of
electronic nonstructural texts documents was performed, and it allowed to provide the operational
actualization of rubrics depending on the structure and parameters of the text documents in the
conditions of nonstationary composition of thesaurus and changes of the rubrics keywords importance.
Automated rubrication of 5062 massages received in 2016–2017 was performed through the internet
portal and by electronic mail. The analysis showed the presence of 17 different interconnected
rubrics, among them there are rubrics such as general issues of society and politics, separation of
powers and duties in Administration, social sphere, education, family, culture, housing and
communal service etc. The results of rubrication showed that rubrics dynamic accounting, when using
the probabalistic classification algorithm of text information as a basic tool of analysis [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ], allowed
to reduce the number of erroneously rubricated texts up to 13,3 % in general.
      </p>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <p>The offered method was used in Administration of Smolensk region when automated analysis of
electronic nonstructural texts documents was performed, and it allowed to provide the operational
actualization of rubrics depending on the structure and parameters of the text documents in the
conditions of nonstationary composition of thesaurus and changes of the rubrics keywords importance.
Eventually, the number of erroneously rubricated texts was managed to be reduced to 13.3 % on
average.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The work is conducted under support of the Russian Foundation for Basic Research (project
18-0100558).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Ageev</surname>
            <given-names>M S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dobrov</surname>
            <given-names>B V</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lukashevich N V 2008 Automatic Text</surname>
          </string-name>
          <article-title>Rubrication: Methods and Problems</article-title>
          , Scientific notes of Kazan State University. Vol
          <volume>150</volume>
          . No.
          <issue>4</issue>
          . pp
          <fpage>25</fpage>
          -
          <lpage>40</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Dumais</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Platt</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heckerman</surname>
            <given-names>D</given-names>
          </string-name>
          and
          <string-name>
            <surname>Sahami</surname>
            <given-names>M 1998</given-names>
          </string-name>
          <article-title>Inductive Learning Algorithms and Representations for Text Categorization</article-title>
          ,
          <source>Proc. Int. Conf. on Inform. and Knowledge Manage</source>
          . pp
          <fpage>148</fpage>
          -
          <lpage>155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Yang</surname>
            <given-names>Y</given-names>
          </string-name>
          and
          <string-name>
            <surname>Liu X 1999</surname>
          </string-name>
          <article-title>A Re-examination of Text Categorization Methods</article-title>
          ,
          <source>Proc. of Int ACM Conf. on Research and Development in Information Retrieval (SIGIR-99)</source>
          . pp
          <fpage>42</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Zaboleeva-Zotova A</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petrovsky</surname>
            <given-names>A B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orlova Yu</surname>
            <given-names>A</given-names>
          </string-name>
          and
          <string-name>
            <surname>Shitova</surname>
            <given-names>T A</given-names>
          </string-name>
          2016
          <source>Automated Analysis of News Texts Themes , Int. J. Information Content and Processing</source>
          . Vol.
          <volume>3</volume>
          . No.
          <issue>3</issue>
          . Pp 288-
          <fpage>299</fpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Kozlov</surname>
            <given-names>P Yu</given-names>
          </string-name>
          2017
          <source>Methods of Automated Analysis of Short Nonstructural Text Documents, Software products and systems No. 1</source>
          . pp
          <fpage>100</fpage>
          -
          <lpage>106</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Borisov</surname>
            <given-names>V V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dli</surname>
            <given-names>M I</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kozlov P Yu 2017</surname>
          </string-name>
          <article-title>Intellectual Methods of Nonstructural Texts (Smolensk: Universum)</article-title>
          .
          <source>p 156 ISBN 978-5-91412-364-9</source>
          (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Uchitelev</surname>
            <given-names>N V</given-names>
          </string-name>
          <year>2013</year>
          <article-title>Classification of Text Information with the Help of SVM, Information technologies and systems</article-title>
          .
          <source>No. 1</source>
          . pp
          <fpage>335</fpage>
          -
          <lpage>340</lpage>
          . (in Russian).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Sajadi</surname>
            <given-names>A</given-names>
          </string-name>
          and
          <string-name>
            <surname>Borujerdi M 2013 Machine Translation</surname>
          </string-name>
          <article-title>Based on Unification Link Grammar</article-title>
          ,
          <source>Journal of Artificial Intelligence</source>
          Review pp
          <fpage>109</fpage>
          -
          <lpage>132</lpage>
          . DOI:
          <volume>10</volume>
          .1007/s10462-011-9261-7.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Protasov</surname>
            <given-names>S Link</given-names>
          </string-name>
          <string-name>
            <surname>Grammar (Electronic materials)</surname>
          </string-name>
          . - http://sz.ru/parser/doc/ (Accessed July,
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Borisov</surname>
            <given-names>V V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fedulov A S and Zernov M M 2014</surname>
          </string-name>
          <article-title>The Base of the Fuzzy Sets Theory, The Base of Fuzzy Mathematics Series book 1 (Moscow: Hot line-Telecom). p 88 (in Russian)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Gimarov</surname>
            <given-names>V A</given-names>
          </string-name>
          <year>2004</year>
          <article-title>Methods and Automated Systems of Dynamic Classification of Complex Technogenic Objects, Synopsis of a thesis paper of Dr</article-title>
          .
          <source>Tech.Sc. (Moscow) (in Russian)</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>