<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Estimation by Phrases for the Closeness of a Topical Text to the Semantic Pattern without Paraphrasing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dmitry Mikhaylov</string-name>
          <email>Dmitry.Mikhaylov@novsu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gennady Emelyanov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Yaroslav-the-Wise Novgorod State University</institution>
          ,
          <addr-line>Velikii Novgorod</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>22</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>In this paper, the numerical estimation method for the closeness of a topical text to the most rational linguistic variant (i.e., semantic pattern or sense standard) of description the corresponding knowledge fragment without paraphrasing, is offered. As the analyzed texts the abstracts of scientific articles together with their titles are considered. The base for estimation of the closeness of a text to the semantic pattern is the splitting of words of each of its phrase into classes by the value of the TF-IDF metric relative to the corpus pre-formed by an expert. The paper considers two variants of estimation: relatively to the article title and the phrase closest to the semantic pattern.</p>
      </abstract>
      <kwd-group>
        <kwd>intelligent data analysis</kwd>
        <kwd>e-learning</kwd>
        <kwd>natural-language expression of expert knowledge</kwd>
        <kwd>human-computer interaction in education</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Development of e-learning significantly increases the qualitative requirements for
electronic training materials. The major requirement here may be formulated as the
sorting of information sources by degree of reflection of the most significant concepts
of the studied subject area at a maximal compactness and non-redundancy of
narration. Ideally, the information sources form a hierarchy at a top level of which will be
placed the start points for study.</p>
      <p>Essentially close problem is the construction and verification of thematic models of
major conferences with the finding of most relevant themes for a new participant [1].
Here the theme of a document is defined by its terms from the terminological
dictionary of the conference. The significance value of the term is expressed via its entropy
relatively to expert classification on a given level of hierarchy. A primary role here
plays a revelation of a set of text units and their relations necessary and sufficient to
represent a knowledge unit and satisfies the semantic pattern.</p>
      <p>The current work considers the possibility of applying the estimation offered in the
paper [2] for the closeness to a semantic pattern and based on the TF-IDF metric
without paraphrasing the phrases of the analyzed text. Herewith as the analyzed texts,
the abstracts of scientific articles together with their titles are considered. These parts
of the articles reflect the main content of each paper and the most important results
without unnecessary methodological details.
2</p>
      <p>The choice of estimating for the closeness to the pattern for
phrases and their groups</p>
      <p>Let D be a topical corpus of texts, selected by an expert. According to the
definition, TF-IDF is the product of term frequency (TF) and inverse document frequency
(IDF, [3]) and intended to reflect how important a word ti is to a document d  D .
We have used the classic case of term frequency, it is the number of times that the
word ti occurs in a document d divided by the total number of words in d . The IDF
metrics can be determined as idf ti , D  log D Di  , where Di  D is the number of
documents where the word ti appears at least once (i.e. tf ti , d   0 ).</p>
      <p>Let X be a descent-ordered sequence of TF-IDF values for words of the initial
phrase relatively to a document d from the corpus D .</p>
      <p>Let’s split X into clusters H1,, Hr using the algorithm offered by us in [4] and
close to FOREL class taxonomy algorithms [5]. Further in the current paper,
concerning to clustering of phrases and documents, we’ll have in mind this algorithm. As the
mass center of cluster Hi the arithmetic mean of all x j  H i like in [4] is taken. Also,
we note that for  i  j H i  H j   , and H1  H 2   H r  X . The rule to relate
elements of X to the same cluster is identical to the one used in the paper [4].</p>
      <p>Let firstX  be the first, lastX  be the last element of the X sequence, and
mcX  be the center of mass of X .</p>
      <p>
        Statement 1. The elements of X can be related to the same cluster, if
 mcX   firstX   mcX 
 4
 mcX   lastX   mcX 
 4
.
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
      </p>
      <p>
        The choice of denominators of right-hand sides of inequalities in formula (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) was
based on the assumption that the elements of the same cluster always have more
similarities than differences. To estimate the affinity of some phrase to the semantic pattern the
most important clusters obtained from the splitting of sequence X will be:
 the cluster H1 (the terms from the source phrase which are the most unique for the
analyzed text document);
 the “median” cluster Hr 2 which will host general vocabulary that ensures
periphrases and synonymous terms;
 the cluster Hr to which the terms that prevail in the corpus are corresponded.
      </p>
      <p>
        The estimation of the closeness of a separate phrase to the semantic pattern without
paraphrasing the natural-language description of a corresponding knowledge unit is
based on the following empirical consideration. First, the division of words into
general vocabulary and terms here should be expressed as much as possible. Another
important aspect is that the words in clusters H1,, Hr , formed by the TF-IDF of
words of the source phrase relative to a certain d  D , should be distributed more or
less evenly. But unlike the estimation from the paper [2], the number of resulted
clusters must be close to three as much as possible at a maximum of TF-IDF values for
words related to the cluster H1 . The latter requirement means the maximal relevance
of term words in phrases of selected documents to the formed corpus. Essentially, the
told above can be represented as the maximization of values
and, correspondingly,
val1  1 log10  H1 ,
val2  10 Hi , i1, r 2 ,r  ,
val3  H1 \ H r 2 \ H r lenX  .
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
      </p>
      <p>
        The logarithm in the denominator of (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) is taken from the sum of TF-IDF values
for words related to the cluster H1 by the value of this metric relative to document
d  D under consideration;  Hi , i  1, r 2 , r in formula (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) is the
root-meansquare deviation (RMSD) of number of elements in a cluster from H1, Hr 2, Hr ;
lenX  in the denominator of formula (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) is the length of X . In a case of  H1  0 ,
the value of val1 is assumed to be zero. If the number of TF-IDF-clusters obtained is
smaller than two, the values of H r 2 and H r are assumed to be zero. In a case of
only two TF-IDF-clusters obtained, the value of H r is assumed to be zero.
      </p>
      <p>
        Documents d  D are sorted by the descending product of estimations (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) and
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        ). As the numerical estimation of the closeness of an individual phrase to the pattern
the greatest of the resulting values herewith is taken.
      </p>
      <p>Let Ts be a group of phrases, first of which is the title of scientific article and
others represent its abstract. In the current paper, two variants for estimation of the
affinity of Ts to the semantic pattern are introduced. Both variants are equally assumed
the minimum of RMSD for the value of affinity to the pattern for all Tsi  Ts .</p>
      <p>The first one essentially corresponds to the order of selection of articles with the
analysis of title at first and assumes the maximal closeness to the standard for it, i.e.:
N1 Ts, D 
max  val1 Ts1 , d  val2 Ts1 , d  val3 Ts1 , d </p>
      <p>dD
.</p>
      <p>
        (
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
      </p>
      <p>
        Note, that estimation (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) does not imply the sorting of phrases Tsi  Ts by affinity
to the semantic pattern. Such a problem statement is the most adequate to requirement
general accepted in scientific periodicals to reflect in the title the content of the
article. Nevertheless, the a priori assumption of maximal closeness to the standard
exactly of the title of the article is not always performed in practice.
      </p>
      <p>
        Taking into account the mentioned above, in the second variant the maximum of
the found values of affinity to the standard for all phrases Tsi  Ts is used in the
numerator of formula (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ):
      </p>
      <p>N 2 Ts, D 
max  val1 Tsmax , d  val2 Tsmax , d  val3 Tsmax , d </p>
      <p>
        dD
 
σ  max val1 Tsi , d  val2 Tsi , d  val3 Tsi , d  , Tsi  Ts   1
 dD 
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
where Tsmax  Ts is the phrase for which the affinity to the sense standard is maximal.
To prevent a possible division by zero, a one is added to the RMSD value in the
denominator of each of the formulas (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) and (
        <xref ref-type="bibr" rid="ref6">6</xref>
        ).
      </p>
      <p>
        Statement 2. The maximal final rank in the collection will be designated to the
paper with the greatest value of estimation (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) related to the same cluster with the
value of estimation (
        <xref ref-type="bibr" rid="ref6">6</xref>
        ) for this paper according to the condition of Statement 1.
      </p>
      <p>
        Note. The correct applying of Statement 2 assumes relating to the same cluster
the value of estimation (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) for the article with the maximal final rank, and a
maximal value of estimation (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) in the collection for paper selection. At the absence
of an article meets this requirement, the maximal final rank will be designated to
the article with the highest value of estimation (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) in analyzed collection.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Experimental research</title>
      <p>To test the proposed estimations, as an expert-formed corpus D the variant from
experiments in the paper [2] was involved. It was formed from the following editions:
 Taurida journal of computer science theory and mathematics (3 papers);
 Proceedings of International conferences “Intelligent Information Processing”
IIP8 and IIP-9 of the years 2010 and 2012 (2 papers);
 Proceedings of the 15th All-Russian Conference with International Participation on</p>
      <p>Mathematical Methods for Pattern Recognition (MMPR-15, 2011, 1 paper);
 Proceedings of the Conference MMPR-13 (2007, 2 papers);
 Proceedings of the Conference MMPR-16 (2013, 14 papers);
 Proceedings of the Conference IIP-10 (2014, 2 papers);
 a scientific report prepared by the first author of the current paper in 2003.</p>
      <p>The scope of selected papers includes:
 mathematical methods for learning by precedents (K.V. Vorontsov, M.Yu.
Khachay, E.V. Djukova, N.G. Zagoruiko, Yu.Yu. Dyulicheva, I.E. Genrikhov, A.A.</p>
      <p>Ivakhnenko);
 methods and models of pattern recognition and forecasting (V.V. Mottl, O.S.</p>
      <p>Seredin, A.I. Tatarchuk, P.A. Turkov, M.A. Suvorov, A.I. Maysuradze);</p>
      <p>Here the number of words in corpus documents is varied from 218 to 6298, and the
number of phrases per document is varied between 9 and 587. Selection of articles
was made from:
 proceedings of the conference IIP-9 (2012), section “Theory and Methods of
Pattern Recognition and Classification” (14 articles);
 proceedings of the conference MMPR-14 (2009), section “Methods and Models of</p>
      <p>Pattern Recognition and Forecasting” (35 articles);
 proceedings of the conference MMPR-15, section “Theory and Methods of Pattern
Recognition and Classification” (18 articles) and “Statistical Learning Theory” (10
articles).</p>
      <p>The main criterion when choosing collections, as well as when selecting texts for
corpus D , was the most complete and evident division of words of the analyzed texts
into general vocabulary and terms.</p>
      <p>The software implementation (in Python 2.7) of the offered solutions and
experimental results are presented on the website of Yaroslav-the-Wise Novgorod State
University at http://www.novsu.ru/file/1504831.</p>
      <p>
        Taking into account the conclusions of [2] regarding the semantic context of terms,
the evaluation of estimations (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )–(
        <xref ref-type="bibr" rid="ref6">6</xref>
        ) was made without consideration of prepositions
and conjunctions. Text extraction from a PDF file was implemented using the
functions of the pdfinterp, converter, layout, and pdfpage classes as part of the PDFMiner
package [6]. For the correctness of formula recognition, as in [2], all formulas from
the analyzed documents here were translated by an expert manually into a format
close to that used in LaTeX. To select the boundaries of sentences in the text by
punctuation marks, the method sent_tokenize() of the tokenize class from the open-source
library NLTK [7] was used. Lemmatization of words was performed using the
morphological analyzer pymorphy2 [8]. If a word has more than one parsing variant when
determining its initial form (lemma), to calculate the TF-IDF measure, the closest one
issued by the n-gram tagger from the nltk4russian library [9] is taken.
      </p>
      <p>
        The experimental results represented further in the tables confirm the rule of “good
manners” of some periodicals on information science and computer engineering to
display in the title the name of method, model, algorithm presented by paper, as well
as the theoretical basis of the proposed solutions. For the collection “MMPR-15,
Statistical Learning Theory” the maximums of estimations (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) and (
        <xref ref-type="bibr" rid="ref6">6</xref>
        ) took place relative
to the same article, a similar result was reached for the collection “MMPR-15, Theory
and Methods of Pattern Recognition and Classification”.
      </p>
      <p>
        As can be seen from Tables 1 and 2, the values of estimations (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) and (
        <xref ref-type="bibr" rid="ref6">6</xref>
        ) for the
mentioned articles are coincided. So, according to the condition of Statement 2, the
papers “Принцип максимизации зазора для монотонного классификатора
ближайшего соседа” (The principle of gap maximization for nearest neighbor
monotonic classifier) by K.V. Vorontsov and G.A. Makhina, and “Полные решающие
деревья в задачах классификации по прецедентам” (Complete decision trees in
classification tasks by precedents) by I.E. Genrikhov and E.V. Djukova will have a
maximal final rank each in its collection.
      </p>
      <sec id="sec-2-1">
        <title>The maximum affinity to the sense standard for the title is achieved relative to the document Value of estimation (5)</title>
        <p>
          Value of estimation (
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
        </p>
        <sec id="sec-2-1-1">
          <title>Dvoenko, S.D., Pshenichny, D.O. Metrical correc</title>
          <p>tion of matrices of pairwise comparisons
(Метрическая коррекция матриц парных сравнений).</p>
          <p>In: MMPR-16 (2013)
0.0920
0.0920</p>
          <p>
            The result obtained for the collection “MMPR-14, Methods and Models of Pattern
Recognition and Forecasting” illustrates the case when an article with the greatest value
of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) in the collection has the value of estimation (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) not relates to the
same cluster with it. Indeed, for the article “Selection of support object set for robust
integral indicator construction” (Выбор опорного множества при построении
устойчивых интегральных индикаторов) by D.I. Melnikov, V.V. Strijov, E.Yu.
Andreeva and G. Edenharter the values of estimations (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) and (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) equal,
correspondingly, to 0.0129 and 0.1426, form two independent clusters according to the condition
of Statement 1. By this virtue, the maximal final rank in the collection will be
designated to the article by O.V. Barinova and D.P. Vetrov having the maximal value of
estimation (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) relatively to the considering collection.
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Phrase closest to the standard</title>
      </sec>
      <sec id="sec-2-3">
        <title>The maximum affinity to the sense standard for the phrase is achieved relative to the document</title>
        <p>
          The maximum affinity to the sense
standard for the phrase is achieved
relative to the document
Value of estimation (
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
        </p>
        <p>
          Value of estimation (
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
Phrase closest to the standard
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Title of the article</title>
        <sec id="sec-2-4-1">
          <title>IIP-9, Theory and Methods of Pattern Recognition and Classification Author(s) Dvoenko, S.D., Pshenichny, D.O. On negative eigenvalues removing from matrices of pairwise comparisons</title>
          <p>
            A similar situation also takes place for the collection “IIP-9, Theory and Methods
of Pattern Recognition and Classification”. Here the maximal value of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            )
equal to 0.1336 will be belonged to the article “Критерии точности
комбинаторных оценок обобщающей способности” (The exactness criteria of
combinatorial generalization bounds) by N.K. Zhivotovskiy and K.V. Vorontsov. The value of
estimation (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) here is equal to 0.0600 and related to the same cluster with the
maximal value equal to 0.0920 for this estimation in collection, but not lies in the same
cluster with the value of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) for this article. Therefore the maximal final
rank obtains the article by S.D. Dvoenko and D.O. Pshenichny having the greatest
value of estimation (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) in the considered collection.
          </p>
          <p>
            Since the title and phrases of the article abstract (by definition) represent a certain
single semantic image, it is entirely acceptable to swap with each other the
estimations (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) and (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) in Statement 2. In considered examples for both collections by
MMPR-15 conference, the maximal final ranks herewith will be designated to the
same articles. For the collection “MMPR-14, Methods and Models of Pattern
Recognition and Forecasting” the maximal rank here the article by O.V. Barinova and D.P.
Vetrov obtains again. Indeed, the maximal value of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) in this collection
will be for the article by D.I. Melnikov, V.V. Strijov, E.Yu. Andreeva and G.
Edenharter. But as we showed earlier, the values of estimations (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) and (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) for this article
are related to different clusters. Therefore, according to the condition of Statement 2,
the maximal final rank obtains the article having among the remaining articles
(except the article mentioned above) the maximal value of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) relating
to the same cluster with the value of estimation (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) for itself, i.e. the article by
O.V. Barinova and D.P. Vetrov.
          </p>
          <p>
            The single exclusion in the considered series of experiments will be the result for
collection “IIP-9, Theory and Methods of Pattern Recognition and Classification”.
As in the previous example, the maximal final rank in the collection may be
designated to the article by S.D. Dvoenko and D.O. Pshenichny as having the maximal value
of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) which relates to the same cluster with the value of estimation (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) for
this paper. But the value of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) for it does not relate to the same cluster
with the maximal value of this estimation in the collection. So, the maximal final rank
together with the maximal value of estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ) in this collection here obtains the
article by N.K. Zhivotovskiy and K.V. Vorontsov.
          </p>
          <p>
            It should be noted that both estimations, like previously proposed in [2], depend
essentially on the selection of a subject-oriented corpus D by the expert.
Nevertheless, the presented results confirm the hypothesis relative to the semantic load of title
for scientific paper on information science and computer engineering. For disputable
cases similar to the shown in the previous paragraph, depending on the subject area
it’s possible to give preference to the requirement of relating to the cluster of maximal
value of either estimation (
            <xref ref-type="bibr" rid="ref5">5</xref>
            ) or estimation (
            <xref ref-type="bibr" rid="ref6">6</xref>
            ).
          </p>
          <p>The main result of this paper is the proposed method for estimating the closeness of
a text to the semantic pattern relative to a topical text corpus.</p>
          <p>The effectiveness of the proposed method can be estimated by splitting of texts in
the collection into clusters by the value of used estimation for the closeness to a
pattern and the ratio of the number of texts assigned to the cluster of the highest
evaluation values to the total number of texts in the collection. So, on the material of
collections mentioned in Tables 1–3, we have at least a threefold reduction in the number of
documents that should be read first when studying a given subject area.</p>
          <p>Taking into account the evaluated degree of division of its words into general
vocabulary and terms, when a phrase is referred to as a “representative of the pattern”, it
is also of interest to reveal key combinations from words with the greatest TF-IDF
values. At disputable cases, the presence of key combinations in abstracts and titles
can be a basis for designating the final rank to the article. To identify the key
combination of words herewith it is necessary to enter into consideration the interpretation
of TF-IDF metrics which would estimate the number of simultaneous presence of all
words from analyzed combination in the phrases of separate document.</p>
          <p>The work was supported by the RFBR (project no. 19-01-00006).</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kuzmin</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aduenko</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strijov</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          :
          <article-title>Thematic Classification Using Expert Model for Major Conference Abstracts (in Russian)</article-title>
          .
          <source>Informational Technologies</source>
          <volume>6</volume>
          (
          <issue>214</issue>
          ),
          <fpage>22</fpage>
          -
          <lpage>26</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Emelyanov</surname>
            ,
            <given-names>G.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikhailov</surname>
            ,
            <given-names>D.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozlov</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          :
          <article-title>Relevance of a Set of Topical Texts to a Knowledge Unit and the Estimation of the Closeness of Linguistic Forms of Its Expression to a Semantic Pattern</article-title>
          .
          <source>Pattern Recognition and Image Analysis</source>
          <volume>28</volume>
          (
          <issue>4</issue>
          ),
          <fpage>771</fpage>
          -
          <lpage>782</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>K.S.:</given-names>
          </string-name>
          <article-title>A statistical interpretation of term specificity and its application in retrieval</article-title>
          .
          <source>Journal of Documentation</source>
          <volume>60</volume>
          (
          <issue>5</issue>
          ),
          <fpage>493</fpage>
          -
          <lpage>502</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Emelyanov</surname>
            ,
            <given-names>G.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikhaylov</surname>
            ,
            <given-names>D.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozlov</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          :
          <article-title>Formation of the representation of topical knowledge units in the problem of their estimation on the basis of open tests (in Russian)</article-title>
          .
          <source>Machine learning and data analysis 1 (8)</source>
          ,
          <fpage>1089</fpage>
          -
          <lpage>1106</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Zagoruiko</surname>
            ,
            <given-names>N.G.</given-names>
          </string-name>
          :
          <article-title>Applied Methods of Data and Knowledge Analysis (in Russian)</article-title>
          .
          <source>Institute of Mathematics SD RAS</source>
          ,
          <string-name>
            <surname>Novosibirsk</surname>
          </string-name>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>PDFMiner - Python PDF</surname>
          </string-name>
          parser and analyzer, https://euske.github.io/pdfminer/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Natural</given-names>
            <surname>Language</surname>
          </string-name>
          <string-name>
            <surname>Toolkit</surname>
          </string-name>
          , http://www.nltk.org
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Korobov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Morphological Analyzer and Generator for Russian and Ukrainian Languages</article-title>
          .
          <source>In: 4th International Conference on Analysis of Images, Social Networks and Texts</source>
          , pp.
          <fpage>320</fpage>
          -
          <lpage>332</lpage>
          . Springer (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Moskvina</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orlova</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Panicheva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrofanova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Development of the Core for Syntactic Parser for Russian based on NLTK libraries (in Russian)</article-title>
          .
          <source>In: Computational Linguistics and Digital Ontologies: Proceedings of the XIX International Joint Conference on Internet and Modern Society (IMS</source>
          <year>2016</year>
          ), pp.
          <fpage>44</fpage>
          -
          <lpage>54</lpage>
          . St.
          <string-name>
            <surname>Petersburg</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>