<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparative Research of Index Frequency - morphological Methods of Automatic Text Summarisation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vladimir Fomin vv_fomin@mail.ru</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Alexsander Osochkin</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Herzen State Pedagogical University of Russia Saint Petersburg</institution>
          ,
          <addr-line>Russian Federation</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Olga Yakovleva</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article considers the potential of frequency-morphological analysis implementation in index methods of automatic text summarisation. The main feature of the developed index method using frequency-morphological analysis is the consideration of the importance of parts of speech in the particular language. The evaluation of the effectiveness of automatic text summarisation of scientific and educational documents, fiction in Russian using various indexing methods is presented in the paper. Based on the experiments results, indexing methods were evaluated and quality ranked in automatic text summarisation algorithms, recommendations for their use were made.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Automatic summarisation is an automatic process that creates a result text from one or
more source texts that transmits most of the information in smaller size. [Brandow et al., 1995].
Today, there are many different methods of automatic summarisation [Clayton et al., 2011],
[Evdokimenko, 2013], [Al-Emran, 2017], [Salloum et al., 2017], [Sujit et al., 2013], [Shari, 2018],
but among all automatic summarisation methods, it is particularly worth highlighting indexing
methods, which are based on simple, well-proven frequency analysis methods [Sujit et al., 2013],
[Shari, 2018], [Molchanov, 2015], [Jansen, 2010] of text information on NL.</p>
      <p>In the field of text-mining and natural language processing NLP, frequency analysis is the
predominant method of text analysis, [Yogesh, 2014] but increasingly in scientific articles there
is a shift to more complex methods of text analysis, identification of the basis of sentences,
etc., using frequency-morphological analysis. The main problem of automatic summarisation is
the identification of the most significant parts of the text, which removed from the text would
save the integrity and reflect the main topic of the document. There is a quite large amount of
automatic summarisation methods developed over the past two decades, but all of them can be
conditionally divided into two groups: extraction methods and abstraction methods.</p>
      <p>Abstraction methods are automatic summarisation methods based on the creation of a new
text using new words and synonyms, consolidating the original text. These methods are of great
scientific interest, especially in the field of NLP, as they involve the use of complex semantic
analysis algorithms.</p>
      <p>Abstraction methods include three necessary steps:
1) creation of the text main idea, frequently used words and main topic identification, etc.;
2) indexing of words, phrases and other meaningful units;
3) indexing-based new text consolidation and synthesis.</p>
      <p>Extraction methods are a method of automatic summarisation text in which words with
low indexes are extracted from the text. The distinctive feature of this method is saving the
original text. The algorithm identifying the importance of sentences and words using indexing
methods, which allow ranking elements of the text: words, sentences, paragraphs. The majority
of industrial-scale automatic summarisation systems are implemented within the framework of
this approach [Yogesh, 2014], although these systems also have a number of problems.</p>
      <p>Regardless of the type of automatic summarisation method, each uses indexing of the text
internal content, in order to rank the text elements and save the most significant ones. Therefore,
the most important step for two kinds of automatic summarisation is the step of indexing the
text internal content. Despite the fact that this stage is key for summarisation method, there
is no general reliable indexing method, which would be effective in a large number of different
tasks. Therefore, there is a wide range of indexing algorithms; each of them includes an effective
method of text summarisation applying to the particular structure or text type.</p>
      <p>The first automatic summarisation methods were based solely on frequency or positional
analysis based on the analysis of each individual word or its position in the text. With the
development of text-mining and NLP, more sophisticated methods of semantic and linguistic
analysis began to be applied, scientists tried to apply text analysis with the higher
meaningful units as paragraphs, thematic parts, sentences and. etc. The main problem of methods
based on semantic and linguistic analysis is the lack of comparison of the sentences
importance [Baxendale et al., 1958], [Yogesh, 2014], [Shari, 2018], [Dragomir, 2012]. This feature
significantly reduces the quality of automatic summarisation, it refers in a big extent to the texts,
where the author mentions several topics as well as to the artistic texts.</p>
      <p>The main problem with indexing is that there is no consensus about what minimum unit
of text analysis is the best for auto summarisation. On the one hand, frequency methods of
text indexing are analysed at the level of unigrams, a separate word cut from the context, on
the other hand semantic and linguistic methods of indexing involve sufficiently large meaningful
units such as paragraphs, subsections, sentences, etc.</p>
      <p>Due to the same problem in two types of automatic summarisation indexing, we have chosen
indexing methods based on frequency analysis, as they are more universal, less dependent on
language specificity, there are ready-made software solutions for indexing documents and these
methods do not require specialized knowledge in the field of linguistics.</p>
      <p>In the science researchers of automatic summarisation [Sujit et al., 2013], [Shari, 2018],
[Molchanov, 2015], [Jansen, 2010], [Yogesh, 2014], [Tarasov, 2010], [Gambhir, 2016] it was
revealed that modern algorithms related to indexing of text including the frequency method are
based on uniform analysis of text, and ignore the analysis of higher organized units: phrases,
sentences, and paragraphs. When analysing higher units of text, new properties appear:
cohesion, coherence, and auto-somatization of individual paragraphs and text lines, etc. The use
of more highly organized units while indexing a document requires a transition to
frequencymorphological analysis, to identify and use special properties of phrases, sentences, paragraphs.
2.1</p>
      <sec id="sec-1-1">
        <title>Relevance and purpose of the article</title>
        <p>Information redundancy is a major problem in various information environments, where
huge amounts of semi-structured data in natural language are accumulated. This problem is
particularly relevant to information educational environments, because the provision of brief
reference material allows to speed up the process of searching for the necessary information,
which affects the level and quality of education in general. Nowadays there is no doubt that
intelligent search significantly increases efficiency in any information environment that searches
in huge amounts of semi-structured data in natural language. In such circumstances, new effective
methods for dealing with large amounts of information that can convey the exact content of a
document in a concise form are of particular importance.</p>
        <p>One of these methods is automatic summarisation as a type of analytical and synthetic
document processing [Sujit et al., 2013], [Shari, 2018], [Radev et al., 2002], which allows the
required information support [Molchanov, 2015], [Jansen, 2010]. The purpose of the study is to
evaluate the effectiveness of automatic text summarisation using index methods, including the
use of frequency-morphological analysis.</p>
        <p>In order to achieve this purpose, the following tasks were set: To analyse approaches to text
auto summarisation based on index methods. Selecting Frequency indexing Methods to generate
automatic summarisation of text materials in Russian. To modify the selected index methods
using frequency-morphological analysis as the main one. To evaluate and compare the results
of auto summarisation, frequency and frequency morphological analysis in terms of accuracy,
completeness and amount reduction from the source text and the standard.
3</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Automatic summarisation algorithm</title>
      <p>When using index automatic summarisation methods, any Dj text in natural language can
be represented as a set of words: W = w1, w2,... wn. Where, each word wn has an index
F obtained by calculating the frequency indexing method. When using the frequency analysis
based indexing method, text is represented at the elementary level, thus the main elementary
unit of frequency analysis is the word.</p>
      <p>We propose to complicate frequency indexing methods by adding morphological analysis,
thanks to which we will be able to obtain another set of V index.</p>
      <p>The morphological index V is determined on the basis of the importance of the part of
the NL speech on which the text is written. As a result of document indexing by means of
frequency-morphological analysis, index P will be obtained, for which the following statements
are correct (formula 1):</p>
      <p>P = F</p>
      <p>V
(1)</p>
      <sec id="sec-2-1">
        <title>The document indexing process can be represented as a number of steps:</title>
      </sec>
      <sec id="sec-2-2">
        <title>1. Indexing of the document.</title>
      </sec>
      <sec id="sec-2-3">
        <title>2. Carrying out morphological analysis of text.</title>
      </sec>
      <sec id="sec-2-4">
        <title>3. Obtaining a combined frequency-morphological index.</title>
        <p>3.1</p>
        <sec id="sec-2-4-1">
          <title>The index of the document</title>
          <p>The first step in automatic text summarisation is to apply a frequency indexing method
that will allow the calculation of index F for each word in the text. The calculation of index
F depends on the algorithm or method of indexing, in the following the obtained index will be
used for calculations with the index of morphological analysis V. In this paper, we selected the
following as the main algorithms.</p>
          <p>TF-IDF. Luh in 1957 developed a method of analysing text information, which allows to
identify the most significant, relevant words that were supposed to be used to classify documents
in natural language [Luhn, 1958]. At the heart of the TF-IDF method is frequency analysis,
and the hypothesis that the most important words in the test are, in more often than the
rest of the words in the text. Thanks to this approach, the TF-IDF method can be used not
only to classify documents, but also to expand its application, using it to reduce information
redundancy. Sentences that do not contain the most significant words are removed from the text.
The remaining test is then subject to linguistic analysis to agree on the remaining sentences in
the text.</p>
          <p>TF-ISF. TD-IDF modification [Luhn, 1958]: aimed at testing the hypothesis that the most
important words are used more than once in a single sentence, but are rarely found throughout
the document.</p>
          <p>Collocations. Technology of identification of significant sentences in text, which is based on
analysis of weight of phrases. Indexing of significant sentences is calculated as a sentence with
a common word, of the total number of sentence. Position analysis of offers. This technology of
indexing the most important sentences is limited to the hypothesis that all the main sentences are
used at the beginning and end of the text to be indexed, thus, the largest index dials sentences
at the beginning and end of the text, which are then auto summarised.</p>
          <p>The signal method. Theory-based technology that key and most important sentences use
specific words: meaningful, complex, heavy, tasks, goals, etc. The words are used from a special
dictionary developed by H.P. Edmans.</p>
          <p>Neural networks. Deep machine learning - which appeared relatively long ago, actively
developing direction, which has found application to a wide range of problems: robotics, training
and recognition of graphic information and intelligent search. One of the most important works in
the field of in automatic summarisation in recent decades was the results of Collport’s research
[Shari, 2018], [Molchanov, 2015], which developed a unified procedure for machine analysis of
text. Many modern automatic summarisation software use Collobert method [Collobert, 2008].
Using the Collobert approach allows to index by semantic and linguistic importance parts of
the text: sentences, paragraphs, etc. automatic summarisation of text using depth learning in
a neural network, differs from a conventional neural network by the number of layers, which
contributes to more complex calculations.
3.2</p>
        </sec>
        <sec id="sec-2-4-2">
          <title>Morphological analysis</title>
          <p>The use of morphological analysis in the indexing of documents allows to apply more
complex methods of calculation of indices of documents taking into account the special specificity of
NL.</p>
          <p>After investigation of many different texts on classification and automatic summarisation,
it was found that the most used words in the sentence, and the most rarely found in the text,
are much less important than the words located in a certain area of frequency of use (see Figure
1).
text , Qs - frequency of use of combinations Sn part of speech witch other part speech, h – is
the count of parts of speech in the NL on which the analysed text is written. Figure 2 shows the
results of indexing parts of speech, based on 200 artistic texts in Russian.</p>
          <p>Figure 2 Shows morphological indices calculated on the basis of formula 2. The results of
morphological indexing of parts of speech in Russian language showed that the most commonly
used part of speech are verbs, and nouns, as well as combinations thereof. The third most
important part of speech was adjectives, which were often used in combination with nouns,
further increasing the index of the given part of speech. The smallest index is received by service
parts of speech, which are not included in the main parts of sentences.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiment</title>
      <p>Based on the analysis of [Yogesh, 2014], [Tarasov, 2010], [Gambhir, 2016] methods for auto
summarisation, the “Rouge” method was chosen [Yogesh, 2014], [Gambhir, 2016] because the
method is easy to modify, has many varieties and is less prone to the element of chance. Accuracy
calculations are carried out using the freely distributed “Rouge” application on [GitHub “Rouge”].</p>
      <p>The “Rouge” method is based on the use of bigrams. A unique feature of this method is the
detail of units of measure, Rouge - 1 for example, considers a word as a minimum unit, in Rouge
- 2 a minimum unit is a bigram, this evaluation method is called “Rouge – N”. There are other
types of evaluation methods, “Rouge - S” – takes measurements based on the bigrams taking
into account changes in the text sequence, “Rouge - L” takes measurements based on the longest
chain of bigrams of the matching sequence between the template and the text abstract, etc.</p>
      <p>According to the “Rouge” metric, each summary is compared by two indicators called “F1
score” [Getahun, 2017] Precision, Completeness Recall, and based on these two indicators,
another overall indicator is calculated - the measure of accuracy of the test measures. You can
read more about methods of evaluating the results of summarisation on the Internet portal for
natural language processing “RxNLP” [Internet portal “Portal”, “NLP text-mining”].
Frequencymorphological analysis is performed by special software. To date, there is quite a large number of
morphological analysis libraries. In the study we will apply the morphological analysis
methodology developed by the authors [Fomin et al., 2019].</p>
      <p>The difference of calculating indices using frequency-morphological analysis is the limitation
of text, i.e. bringing all words into the initial form, identifying chats of speech of each sentence
member, breaking according to dictionaries of morphological modules, identifying parts of speech,
determining grammatical basis of sentences and chains of parts of speech.
4.1</p>
      <sec id="sec-3-1">
        <title>Text corpora</title>
        <p>In order to conduct an experiment comparing frequency and frequency-morphological
analysis in indexes, two special corpora of text were collected. For implementation of comparison,
assessment of accuracy, completeness and efficiency of summarising, methods need a reference
text, as a rule, the reference text means the paper, the composition, the essay written by the
person.</p>
        <p>An important element in auto summarisation, reduction of the text redundancy, is the
preservation of the basis of the text, which allows to preserve the main theme of the text.
Within the framework of automatic summarisation it is common to define two types of texts
[Yogesh, 2014], [Tarasov, 2010], [Gambhir, 2016]:
context-identifiable - these texts are expected to describe specific issues, problems, topics;
context-indelible - in these texts, there is no clearly marked theme, it can be hidden,
including from the reader, in the general context1.</p>
        <p>Context-identifiable corpora The corpora “Dissertation” refers to context-identifiable issue
automatic summarisation and is represented by graduate works for obtaining the PhD degree,
collected from various sites of universities of the Russian Federation in different directions and
specialties, to each thesis an autoabstract is attached.</p>
        <p>All dissertations and autoabstracts presented in the corpora of texts are published during
the 2007 to 2019 period. The average size of one dissertation: 142 pages or 76,964 words, the
average length of the autoabstract is 22 pages or 8,464 words. The works presented by one subject
area have different topics and directions, for example, for Jurisprudence, the works discuss the
problems of the civil code of the Russian Federation, the judicial document of production, the
Customs Code of the Customs Union, etc. The reference text of the abstract in this corpora
of texts, is considered the autoabstract to the thesis written by the author. As a comparison,
1As a rule, context-indelible group of texts includes artistic works
abstracts created by joint application of indexing technologies with morphological analysis are
used.</p>
        <p>Context-indelible corpora The “Art literature” corps refers to Context-indelible corps and
is represented by various artistic works in Russian, different time eras. As the reference text,
works and essays taken from the Internet with a retelling of the content of the artistic ration are
used. The body “Art literature” is presented in Table 2</p>
        <p>The second comparative corpora of texts of automatic summarisation is made similar to
comparative abstracts in the corpora “Dissertation” where abstracts were created using different
indexing technologies together with morphological analysis.</p>
        <p>Evaluation of automatic summarisation of context-identifiable corpora
We will evaluate the automatic summarisation generated by various indexing methods,
which are based exclusively on frequency analysis. Results of evaluation of automatic
summari</p>
        <sec id="sec-3-1-1">
          <title>Evaluation Method 1 Precision 0,2110718</title>
          <p>TF-IDF Recall 0,1750908</p>
          <p>M-measures 0,191405</p>
          <p>Precision 0,2036052
TF-ISF Recall 0,168897</p>
          <p>M-measures 0,1846341</p>
          <p>Precision
Collocations Recall</p>
          <p>M-measures
Position Precision 0,1940082
analysis of Recall 0,1378786
offers M-measures 0,161197
Tmhetehsoigdnal PRMre-ecmcaielslaiosnures 000,,,111835302756597417443</p>
          <p>Precision 0,1778164
Neural networks Recall 0,1778225</p>
          <p>M-measures 0,1778195
sation by “Rouge-N” metric the results are presented in Table 3 below.</p>
          <p>The absence of Rouge-1 indexing in the “Collocation” indexing method is a consequence of
the inability to use unigrams in text indexing.</p>
          <p>In the evaluation of automatic summarisation by the Rouge-1 method, the highest accuracy
is achieved in autoabstract generated by the method of Neural Network indexing, but the best
overall correspondence (m-measures) was achieved by TF-IDF This situation can be explained
by the fact that the “Recall” of autoreferences obtained by TF-IDF indexing is less than the
“Recall” of autoabstract, but the number of words that often coincided with the the reference
increased, while neural networks used service parts of speech more often (43.34%) than in the
TF-IDF method.</p>
          <p>When using bigrams (Rouge-2), the best “Precision” and “Recall” was shown by the
“Collocations” method, which suggests the presence of coherence in words in the referenced texts.
Increasing the Rouge-3 sequence, after the bigrams, decreases the accuracy of almost all methods
except the phrase. Increase the sequence of n-grams resulted in reduced “Precision” and “Recall”
with the maiming of the chain of n-grams. The best result when using 6 words long n-grams,
was shown by the TF-IDF method.</p>
          <p>Table 4 presents the results of automatic summarisation of dissertation, in which indexing
was carried out on the basis of frequency-morphological analysis. As with frequency analysis,
after indexing and shortening the text, morphological libraries were used to reconcile sentences.</p>
          <p>The best indicator of “Recall” in the evaluation of unigrams (Rouge-1) was shown by the
method of positional analysis of sentences method.</p>
          <p>When using the position analysis of offers method the “Recall” of the main sections
“Introduction”, problem, “Conclusions”, almost did not decrease, because positional they are located
at the beginning and end of the texts, in case the volume of conclusions was sufficiently small,
the abstract included relevance, problem and methodological part of the dissertation.</p>
          <p>With high “Recall”, the number of words matching the reference text was extremely small,
which made the overall “Precision” of the summarisation of the method low.</p>
          <p>The lowest estimate was found in auto summarisation, where the indexing of texts was
carried out with method TF-IDF. The average accuracy of TF-ISF is higher than that of
TFIDF, this result indicates that, with increasing text volume, words that are often used within a
single sentence are often used in autoreferences written by humans.</p>
          <p>Better “Precision” is achieved with any method when using unigrams. When evaluating
autoreferences with unigrams, the similarity with the reference falls, except for the collocations
method. Neural networks reached the best M-measures in auto summarisation, with neural
networks generating the largest volume abstracts.
4.3</p>
          <p>Evaluation of automatic summarisation of context-indelible corpora
We will evaluate automatic summarisation generated by various indexing techniques based
on frequency analysis. Results of evaluation of automatic summarisation by “Rouge-N” metric
the results are presented in Table 5.</p>
          <p>As a result of estimation by the method of Rounge-1 of automatic summarisation generated
on the basis of frequency indexing methods, the best method was neural networks, where the
value of the function reached 43.36%. The best correspondence when evaluation the match
with the standard of bigrams, showed automatic summarisation generation by the method of
“Collocations” where the text of the reference was similar to the writing in 30.5% of cases. In
the evaluation of n-grams of Rouge-3-6, automatic summarisation obtained by the “Collocations”
indexing method also have the best m-measure value with the reference. Now we will carry out
comparative analysis by Rouge-N procedure, using frequency-morphological analysis in indexing,
results of which are presented in Table 6.</p>
          <p>Evaluation of automatic summarisation generated on the basis of frequency-morphological
analysis using the Rouge-1 technique showed that the highest compliance with the reference, was
achieved in methods where the method of neural networks was used. The total compliance with
the standard of 50.6% was achieved, which is 17.94% better than when using frequency analysis
method. The “Precision” indicator reached 97%, which allows saying that almost all words used
in automatic summarisation are also found in the reference text.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>The experiments results show that moving from simple unigram frequency analysis to more
complex frequency-morphological analysis has a great impact on automatic text summarisation
quality. The “Rouge-N” method, used for evaluating the efficiency of automatic text
summarisation, showed that autoabstracts made on base of frequency-morphological analysis were 16%
closer to the original than autoabstracts based on frequency analysis only, moreover automatic
text summarisation of fiction was 31,28% more accurate using frequency-morphological analysis
than using frequency analysis.</p>
      <p>While using frequency-morphological analysis in indexing documents was found that all
methods except TF-IDF increased the similarity with the original text. Comparing to the original
texts, which were written by people, it was found that text on average is accurate by 48% and
the text reduction reached 93,05% while dissertation referencing.</p>
      <p>The experiments results let us suppose the high potential of using index methods on base
of neural networks using frequency-morphological analysis in smart search or in
informationeducational fields. We are expecting to expand the scope of application fields of
frequencymorphological analysis in fiction auto reference and in indexing of parts of speech usage frequency,
in order to classify data on NL.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>The research was supported by the Russian Science Foundation (RSF), Project
“Digitalisation of the high school professional training in the context of education foresight 2035” No
19-18-00108.
[GitHub “Rouge”] Application ¾Rouge¿, ¾GtiHub¿ Avaible at: https://github.com/kylehg/
summariser/blob/master/rouge/ROUGE-1.5.5.pl
[Internet portal “Bookzip”] Avaible at:
boeviki-ostrosjuzhetnaja-literatura/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Brandow et al.,
          <year>1995</year>
          ]
          <string-name>
            <surname>Brandow R. Mitze</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>and Lisa F. R.</surname>
          </string-name>
          (
          <year>1995</year>
          )
          <article-title>Automatic condensation of electronic publications by sentence selection // Inf</article-title>
          . Process.
          <source>Manag</source>
          . Vol.
          <volume>31</volume>
          . Pp.
          <volume>1</volume>
          -
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Baxendale et al.,
          <year>1958</year>
          ] Baxendale P. B. and
          <string-name>
            <surname>etc</surname>
          </string-name>
          (
          <year>1958</year>
          )
          <article-title>Machine-made index for technical literature: An experiment // IBM</article-title>
          <string-name>
            <surname>J. Res. Dev.</surname>
          </string-name>
          , Vol.
          <volume>2</volume>
          , Pp.
          <fpage>354</fpage>
          -
          <lpage>363</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[Lei</source>
          , 2017] Lei L. and
          <string-name>
            <surname>etc.</surname>
          </string-name>
          (
          <year>2017</year>
          )
          <article-title>Redundancy checking algorithms based on parallel novel extension rule</article-title>
          .
          <source>//Journal of Experimental Theoretical Artificial Intelligence</source>
          Vol.
          <volume>29</volume>
          ,
          <fpage>2017</fpage>
          - Issue 3.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Said et al.,
          <year>2017</year>
          ]
          <article-title>Said A.S. and etc (2017) Using Text Mining Techniques for Extracting Information</article-title>
          from Research //
          <source>Intelligent Natural Language Processing: Trends and Applications</source>
          Vol.
          <volume>1</volume>
          . Pp.
          <volume>373</volume>
          -
          <fpage>397</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Salloum et al.,
          <year>2017</year>
          ]
          <string-name>
            <surname>Salloum</surname>
            <given-names>S.A.</given-names>
          </string-name>
          (
          <year>2017</year>
          )
          <article-title>A Survey of text mining in socialmedia: facebook</article-title>
          and twitter perspectives //Advances in Science,
          <source>Technology and Engineering Systems</source>
          Journal Vol.
          <volume>2</volume>
          . Pp.
          <volume>127</volume>
          -
          <fpage>133</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Clayton et al.,
          <year>2011</year>
          ]
          <article-title>Clayton S.and etc (2011) Experiments in Automatic Text Summarisation Using Deep Neural Networks // Machine Learning</article-title>
          , Fall Vol.
          <volume>1</volume>
          <fpage>2011</fpage>
          . Avaible at: https://www.semanticscholar.org/paper/ 545-Machine-Learning-%
          <string-name>
            <surname>2C-Fall-</surname>
          </string-name>
          2011
          <string-name>
            <surname>-Final-</surname>
          </string-name>
          Project-in-Ben-Rahul/
          <year>8f4f64e15553baf9fd0c2933c631b78c97c8f0bc</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Radev et al.,
          <year>2002</year>
          ]
          <string-name>
            <given-names>Radev D. R.</given-names>
            ,
            <surname>Hovy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            , and
            <surname>McKeown</surname>
          </string-name>
          <string-name>
            <surname>K.</surname>
          </string-name>
          (
          <year>2002</year>
          )
          <article-title>Introduction to the special issue on summarisation</article-title>
          . //Comput. Linguist, Vol. No 28.
          <string-name>
            <surname>Pp</surname>
          </string-name>
          .
          <volume>399</volume>
          -
          <fpage>408</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Dragomir</source>
          , 2012]
          <string-name>
            <surname>Dragomir</surname>
            <given-names>R.R.</given-names>
          </string-name>
          (
          <year>2012</year>
          )
          <article-title>Single-document and multi-document summary evaluation via relative</article-title>
          utility University of Michigan, Ann Arbor MI 48109 2012 Avaible at: https://www.eecs.umich.edu/techreports/cse/2007/CSE-TR-
          <volume>538</volume>
          -07.pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Derczynski</source>
          , 2016] Derczynski
          <string-name>
            <surname>L.</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <article-title>Complementarity, F-score</article-title>
          ,
          <source>and NLP Evaluation // Proceedings of the International Conference on Language Resources and Evaluation</source>
          , Vol.
          <volume>1</volume>
          . Pp.
          <volume>1</volume>
          -
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Sujit et al.,
          <year>2013</year>
          ]
          <string-name>
            <surname>Sujit R. Sujit</surname>
            <given-names>V.</given-names>
          </string-name>
          <article-title>and etc (2013) Classification of News and Research Articles Using Text Pattern Mining IOSR Journal of Computer Engineering (IOSR-JCE</article-title>
          ) Vol.
          <volume>14</volume>
          , Issue 5 . Pp.
          <volume>120</volume>
          -
          <fpage>126</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[Luhn</source>
          , 1958]
          <string-name>
            <surname>Luhn P.</surname>
          </string-name>
          (
          <year>1958</year>
          )
          <article-title>The automatic creation of literature abstracts IETE //</article-title>
          <source>Journal of research J. Res. Dev.</source>
          , Vol.
          <volume>2</volume>
          , no.
          <issue>2</issue>
          . Pp.
          <volume>159</volume>
          -
          <fpage>165</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Fomin et al.,
          <year>2019</year>
          ]
          <string-name>
            <surname>Fomin</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osochkin</surname>
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zhuk</surname>
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>Frequency and morphological patterns of recognition and thematic classification of essay and full text scientific publications</article-title>
          //NESinMIS-2019,
          <fpage>12</fpage>
          -Jul-2019, CEUR-WS
          <year>2019</year>
          . Vol.
          <volume>2401</volume>
          ,
          <fpage>69</fpage>
          -
          <lpage>84</lpage>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[Evdokimenko</source>
          , 2013]
          <string-name>
            <surname>Evdokimenko E. Y.</surname>
          </string-name>
          (
          <year>2013</year>
          )
          <article-title>The Concept of Information Noise in the Social and Human Sciences // Molodoy ucheniy</article-title>
          . Vol.
          <volume>10</volume>
          . Pp.
          <volume>564</volume>
          -
          <fpage>566</fpage>
          ,
          <year>2013</year>
          . Avaible at: https: //moluch.ru/archive/57/7765/
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [
          <string-name>
            <surname>Al-Emran</surname>
          </string-name>
          ,
          <year>2017</year>
          ]
          <string-name>
            <surname>Al-Emran</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaalan</surname>
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2017</year>
          )
          <article-title>Academics'awareness towards mobile learning in</article-title>
          <source>Oman // Int.J. Com. Dig. Sys</source>
          . Vol.
          <volume>6</volume>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <source>[Shari</source>
          , 2018]
          <string-name>
            <surname>Shari</surname>
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2018</year>
          )
          <article-title>Optimize Optimize the A Commentary</article-title>
          .
          <source>Journal Search Voice</source>
          , Vol.
          <volume>1</volume>
          , Pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>[Molchanov</source>
          , 2015]
          <article-title>Molchanov A.N and etc</article-title>
          .(
          <year>2015</year>
          )
          <article-title>A mathematical model of natural language text that takes into account the coherence property // Internet-journal “Science of science”</article-title>
          . Vol.
          <volume>7</volume>
          , No 1, 2015 Avaible at:https://naukovedenie.ru/PDF/70TVN115.pdf
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>[Jansen</source>
          , 2010] Jansen,
          <string-name>
            <given-names>B. J.</given-names>
            and
            <surname>Rieh</surname>
          </string-name>
          ,
          <string-name>
            <surname>S</surname>
          </string-name>
          (
          <year>2010</year>
          )
          <article-title>The Seventeen Theoretical Constructs of Information Searching</article-title>
          and Information Retrieval//
          <source>Journal of the American Society for Information Sciences and Technology</source>
          . Vol
          <volume>61</volume>
          . Pp.
          <volume>1517</volume>
          -
          <fpage>1534</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>[Yogesh</source>
          , 2014]
          <string-name>
            <surname>Yogesh</surname>
            <given-names>M .</given-names>
          </string-name>
          et al. (
          <year>2014</year>
          )
          <article-title>Analysis of Sentence Scoring Methods for Extractive Automatic Text Summarisation //</article-title>
          <source>Proceedings of the International Conference on Information and Communication Technology for Competitive Strategies</source>
          . - ACM: NY, USA, vol.
          <volume>1</volume>
          . Pp.
          <volume>89</volume>
          -
          <fpage>97</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>[Gambhir</source>
          , 2016]
          <string-name>
            <surname>Gambhir</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Gupta</surname>
          </string-name>
          . V. (
          <year>2016</year>
          )
          <article-title>Recent automatic text summarisation techniques: a survey // Artificial Intelligence Review</article-title>
          . vol.
          <volume>1</volume>
          Pp.
          <fpage>1</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[Tarasov</source>
          , 2010] Tarasov S.D.
          <article-title>Modern methods of automatic referencing // Scientific and technical statements of SPBPU</article-title>
          . //Journal “Computer science,
          <source>telecommunications and management”</source>
          ,
          <source>vol. No</source>
          <volume>6</volume>
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          2010.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Internet portal “Portal”, “
          <string-name>
            <surname>NLP</surname>
          </string-name>
          text-mining”] Avaible at: http://rxnlp.
          <source>com(update:30.09</source>
          .
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>[Getahun</source>
          , 2017] Getahun T. and etc. (
          <year>2017</year>
          )
          <article-title>Automatic Amharic Text Summarisation using NLP Parser international</article-title>
          . //Journal of Engineering Trends and
          <string-name>
            <surname>Technology (IJETT</surname>
          </string-name>
          ) - Vol.
          <volume>53</volume>
          , Pp.
          <fpage>52</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>[Collobert</source>
          , 2008] Collobert,
          <string-name>
            <given-names>R.</given-names>
            and
            <surname>Weston</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2008</year>
          )
          <article-title>A unified architecture for natural language processing: Deep neural networks with multitask learning</article-title>
          . // Conference: Machine Learning,
          <source>Proceedings of the Twenty-Fifth International Conference (ICML</source>
          <year>2008</year>
          ), Helsinki, Finland, June vol.
          <volume>1</volume>
          , Pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>