<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Identification of Cognitive Actions Speech Genres of Scientific Theoretical Text Constituting</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dmitry Devyatkin</string-name>
          <email>devyatkin@isa.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ludmila Vladimir Salimovsky</string-name>
          <email>salimovsky@rambler.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kadzhaya</string-name>
          <email>kadzhaya@psu.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natalia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chudova</string-name>
          <email>nchudova@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valery</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Federal Research Center “Computer Science and Control” RAS</institution>
          ,
          <addr-line>9 60-let Oktyabrya ave, Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Perm State University</institution>
          ,
          <addr-line>15 Bukireva str., Perm</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Translation Studies, Shandong University</institution>
          ,
          <addr-line>27 Shanda Nanlu, Jinan</addr-line>
          ,
          <country country="CN">P.R.China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents speech genres as forms of social and cultural activity at the stage of its objectification via a system of speech actions in a text as a communication unit. We implement the speech genre typology of a scientific text due to projecting information on the structure of the research and cognitive process onto the text that implements this structure. Moreover, we consider each of the studied speech genres as a cognitive-communicative action system characterized by a specific linguistic marker set. Those markers have been used as a linguistic base for the proof-of-concept implementation of a cognitive action parser for theoretical scientific text. Namely, we applied that linguistic knowledge to build compact high-level features, allowing the parser to be reliably trained on a small manually annotated corpus. The experiments on the corpus show the parser can accurately identify the actions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Speech genre</kwd>
        <kwd>genre form</kwd>
        <kwd>cognitive-speech action</kwd>
        <kwd>scientific text</kwd>
        <kwd>relational-situational model</kwd>
        <kwd>sequence labeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        theoretical activity involves transforming and developing conceptual apparatus [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We regard the
stages of empirical and theoretical cognition as the extralinguistic basis of appropriate speech genres.
      </p>
      <p>The goal of this paper is to study the methods and develop proof-of-concept algorithms to identify
cognitive-communicative actions in scientific texts. The objectives are:</p>
      <p>1. Сharacterize the proposed linguistic and psychological bases for automatic identification of
cognitive actions in scientific texts.</p>
      <p>2. Develop and adapt the linguistic and software tools for detecting descriptions of cognitive
actions in theoretical scientific publications.</p>
      <p>3. Study the developed method and algorithms for detecting descriptions of cognitive actions in a
theoretical scientific text.</p>
      <p>The study material includes 160 scientific theoretical texts — papers and monograph chapters —
in physics, biology, psychology, and linguistics (40 for each science).</p>
      <p>
        The linguistic method is implemented in two stages. The first stage implies projecting the structure
of the cognitive process [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] onto a text array and then correlating the primary goal of a particular text
with one or another section of this structure. Therefore, the obtained speech genres typology of
scientific empirical and theoretical texts looks as follows (Table 1 and 2):
      </p>
      <p>
        The table shows the results of our previous work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The latter proves that the main stages of
empirical scientific research, such as the formation of basic empirical knowledge, distribution of
experimental data into groups, and working out the empirical laws, correspond to the specific speech
genres: description of a new scientific phenomenon (mineral, plant, animal), classification genre, and
report on the empirical law of a cause-and-effect type. A similar pattern is found in the study of
theoretical texts (Table 2).
      </p>
      <p>As can be seen from the diagram, the formation of theoretical ontology corresponds to the
problem-stating theoretical text. Meanwhile, the construction of a theory is carried out in explicating
its central concept by a system of less general ones. Explanation of facts by theory serves as its
confirmation, verification.</p>
      <p>We study each speech genre as a set of interrelated cognitive-communicative actions subordinated
to a common goal. For example, the speech genre « Problem-stating theoretical text» is constituted by
the actions:
1.1. presentation of theories that form available knowledge,
1.2. author's assessment of available knowledge,
The speech genre «Explication of its main concept » involves:
2.1. definition of the concept,
2.2. emphasizing an important thought,
2.3. explanation and clarification of the author's idea.</p>
      <p>The speech genre «Verification text (experimental theory testing)» is comprised of the actions:
3.1. definition of a hypothesis to test,
3.2. description of the experimental methodology,
3.3. analysis and explanation of experimental data,
3.4. conclusion on the confirmation or refutation of the checked hypotheses.</p>
      <p>
        Since the results of automatic analysis of empirical texts were presented earlier [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], we consider
only the texts that embody the main stages of theoretical knowledge.
      </p>
      <p>The second stage of linguistic analysis describes the multi-level language markers of utterances
that realize the studied cognitive actions. At the same time, we become aware that it is crucial to
create software tools that model human cognitive functions to solve problems in artificial intelligence.
In our case, the point is that the perception of the utterance language form necessarily specifies the
author's intention that ensures the adequacy of utterance understanding.</p>
      <p>
        It should be pointed out that the indicator of a particular cognitive action is not the presence of
particular linguistic means in the text itself, but the special nature of the speech system [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ],
generated by peculiarities of choice, recurrence, placement, combination, and modification of
multilevel language units.
      </p>
      <p>
        We use the approach which differs from the numerous pieces of research, based on the genre-study
concept of John M. Swales [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ] and applied in reviewing scientific texts [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ]. While in
those works, the object of analysis is generally a scientific paper Introduction (the authors have
referred to the «Discussion» section recently), we study the whole text, not the set of rhetorical moves
(cognitive and communicative actions) found in a paper’s typical compositional part, but realization a
substantial part of the cognitive process structure in a text as an integral communicative unit. Thus,
our research is cognitive-oriented.
      </p>
      <p>The work, perforce, is of a research outline: only the primary units of the scientific and cognitive
process are considered, and within each unit – the most regularly implemented cognitive actions.</p>
      <p>
        From the technical side, considering the speech system means that the cognitive action parser
should consider multi-level linguistic features from each analyzed text fragment's wide context.
Besides, as far as we analyze natural language, the lexis related to the particular actions can vary
intensively. The typical approach to consider all those context-dependent features is to train deep
neural models with recurrent and attention layers or to utilize a language model, such as BERT [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] or
GPT-2 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>However, cognitive actions tagging is a new one, which means a lack of annotated corpora.
Together with imbalanced classes, all those deep-learning models tend to over-fit. We propose a
multi-step tagging approach to tackle this problem. Firstly, we extract various linguistic features for
every clause; then, we apply the templates to generate a compact feature-set and fit a sequence
labeling model with this reduced feature-set. The final step is the disambiguation of the clauses with
multiple labels.</p>
      <p>
        Such the approach also has reasoning from cognitive science. Namely, perception implies
selectivity concerning the properties of the environment. It begins with a particular aim to receive
such information on the objects' properties that contributes to the most appropriate behavior [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
According to the ideas of development psychology about the existence of cultural standards of
perception and thought [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], as well as the concept of objectification of mental reflection [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ],
information received by the analyzer system inputs, is interpreted considering objects' application
practice in society (for animals – following the practice prescribed by instinctive behavior programs).
As a result, the sensory-perceptual system processes not the isolated signals but the information about
the whole objects. The objective nature of perception, defined by the organizing action criterion,
allows recognizing the essential objects and events for an individual, their actions, and the others,
intentions of subjects' interaction. Instead of tracking the whole set of parameters measured by
neurophysiological detectors, a psychic subject operates with non-random structure discrimination
rules – structures reflecting something meaningful in their life. After that, the application of the
objectification principle in the suggested method for recognizing cognitive-communicative actions
enables us to move from statistical comparisons for a wide range of parameters to handling linguistic
features relevant for identifying a subject's intentions.
      </p>
      <p>
        The rest of the paper is organized as follows. Section 2 briefly provides the results of the recent
studies in the closest research topic, which is RST (Rhetorical Structure Theory) parsing [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Section
3 contains a high-level description of the proposed approach to identifying cognitive actions and
detailed descriptions of each approach's essential step. Eventually, Section 4 has the results of the
experimental evaluation on a manually annotated corpus of scientific texts and some speculations
regarding the reasons for the drawbacks identified.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        The closest NLP-field to cognitive action identification is RST parsing. RST parsers are widely
applicable for text summarization and information extraction. For example, paper [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] utilizes the
hierarchical discourse-level structure of fake and real news articles to distinguish them. Such an
approach has hardly been applied to the fake news detection problem because of the following issues.
First, there is a lack of labeled corpora to train the methods for capturing the discourse-level structure
for fake news. The second issue is how to extract useful information from those identified structures.
To tackle these problems, they propose a hierarchical discourse-level structure for the fake news
detection approach. This structure learns and builds a discourse-level structure for fake or real news
pieces in an automated and data-driven manner. The researchers also revealed structure-related
properties that can describe the revealed structures and increase detection accuracy. Another
application of the RST is presented in paper [22], which proposes an approach to analyzing scientific
discourse structures and the extraction of “evidence fragments” from a corpus of biomedical
experimental research papers. The researchers trained and validated a scientific discourse tagger on
two scientific discourse tagging corpora and checked if it can be transferred to a new dataset. They
also show the benefit of utilizing scientific discourse tags for claim-extraction and evidence fragment
detection. The experiment results show the applicability of evidence fragments derived from image
spans for improving the quality of scientific claims by cataloging, indexing, and reusing evidence
fragments as independent texts. In paper [23], researchers propose a new three-step approach to
automatic text summarization. First, vector space modeling is used to compute coverage and fidelity
scores. Then they apply fuzzy logic to evaluate an aggregated fidelity-coverage score. The last step is
applying a discourse analysis on top of sentences, which have the highest fidelity-coverage scores to
achieve coherence. The experiments on a labeled dataset show that the approach outperforms the state
of the art extractive summarization models.
      </p>
      <p>
        Although RST is a well-known theory, there is still a lack of human-annotated corpora to train
parsers with machine-learning techniques. Therefore, the main research efforts are focused on
unsupervised and semi-supervised approaches, such as methods to build dense contextualized vector
representations (embeddings) of texts. These representations allow using a simpler machine-learning
model to tackle the problem. For example, paper [24] presents an unsupervised automatic text
summarization approach that combines rhetorical structure theory, deep neural model, and domain
knowledge. This approach contains three crucial parts: domain knowledge base construction with
representation learning, attentional autoencoder network for rhetorical parsing, and a
subroutinebased network for text summarization. They use domain knowledge to increase the quality of
unsupervised rhetorical parsing and utilize the concept of translation to tackle the lack of data to train
the rhetorical parsing module. The summarization model hardly depends on the revealed discourse
structure and can generate content-balanced results. They also present an unsupervised metric to
evaluate the obtained results. The experiments show that the proposed approach has the same
accuracy as other modern approaches. In [25], researchers propose a method that evaluates the
applicability of language models for rhetorical analysis. Namely, they test their abilities to encode a
set of linguistic features obtained from RST. The experiments show that BERT-based models [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
outperform others because they identify richer discourse features in their intermediate layer
representations. It has also been shown in other studies [26] that the same BERT layers are
responsible for holding syntax dependencies. This may be a clue to automatic revealing the
dependence between syntax and discourse. However, GPT2 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and XLNet [27] encode less
rhetorical knowledge. They suggest this is because BERT considers context from both the left and
right sides, which seems to be crucial for the task. It also does not permute context elements;
therefore, it does not distort the original meaning of analyzed texts. The performance of the RST
parser is also of research interest. Lin proposes an approach for sentence-level discourse analysis [28].
The process is two-step. First of all, they use a discourse segmenter to detect the elementary discourse
units in a text and then apply a discourse parser that builds a discourse tree in a top-down manner.
Both the components use Pointer Networks [29] and work in linear time.
      </p>
      <p>Summarizing, the main focus of the current research in genre and discourse research lies in
obtaining compact high-level feature sets, allowing one to train the parsers, and on the approaches to
extend annotated corpora or transfer the training results. In this study, we also tackle this problem;
however, to do so, we tried to consider the linguistic and cognitive background of speech perception,
which is not assumed in pure statistical methods.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Cognitive action parser</title>
    </sec>
    <sec id="sec-4">
      <title>3.1. Overall parser schema</title>
      <p>Fig. 1 shows the overall scheme of the parser. Because of the small size of the training corpora, it
is impossible to use end-to-end machine learning models to analyze clauses characterized by raw
high-dimensional feature sets. Therefore, we separate the extraction of cognitive actions into several
following steps: obtaining raw features with the full linguistic analysis of the clauses, matching the
clauses with the context-free templates to generate a high-level feature-set with a smaller dimension,
training of a model for sequence labeling with this compact feature-set, then disambiguation of the
clauses, labeled with several cognitive actions.</p>
      <p>Sequence labeling
High-level features</p>
      <p>Set of actions
Text</p>
      <p>Template matching</p>
      <p>Clause
disambiguation</p>
      <p>Classified text fragments</p>
      <p>As we noted earlier, another reason for such an approach lies in psychology and cognitive science.
Namely, an individual perceives information considering its own experience. As a result, the
sensoryperceptual system processes not the isolated signals but the information about the whole objects. This
schema allows recognizing the essential items and events for an individual, their actions, intentions,
and interactions.</p>
      <p>We gathered a corpus of 688 text fragments of research papers from various fields (Biology, Math,
and Geology) to train and evaluate the parser. Each fragment has from 1 to 10 clauses and from 10 to
1K tokens [30].
3.2.</p>
    </sec>
    <sec id="sec-5">
      <title>Markers of the cognitive-communicative action</title>
      <p>Considering the very limited scope of this paper, we describe only the most significant speech
markers of the Russian texts that realize the given action. We use those markers as the linguistic
background to build the set of context-free templates. As an example of utterance markers’
description, let us examine the means of implementing one of the main cognitive and communicative
actions of a scientific theoretical text — definition of a concept (as shown in [31], in a particular text,
definitive constructions can be considered as constituents of its mental space.). As a main syntactic
model for defining a scientific concept, a construction with two Nominative cases and a quasi-copula
это is used, and also the linking verbs есть, суть (is) (in constructions with plural of both N1),
including a zero one. E.g.: N1 — это …N1; Фонем-а — это функциональная фонетическая
единиц-а…; Фонем-а есть функциональная фонетическая единица…; Фонемы
суть функциональные фонетические единиц-ы…; Фонем-а [ø] – функциональная
фонетическая единиц-а…(dash is an important formal marker). If the machine processes the
corpora of only scientific texts, it is sufficient to determine the case of the nominal components of this
structural model to automatically identify the given genre (utterances like Убийца и есть дворецкий
in scientific texts are possible only as illustrations). While analyzing text arrays that are not
discursively defined, the number of markers includes semantic characteristics of both N1 specified in
pre-formed lists, and also in the morphemic structure of the construction’s nominal components.</p>
      <p>In the subject position (defined notion), it is expected either a substantive term from the object
class (животное, кристалл) or feature names (значение, коммуникация), or — more often —
phrases with a coordinated or uncoordinated attribute (usually in the form of N2) with a reference
name belonging to the class of generalized nominations (функциональн-ый стиль-ø —
это [разновидность-ø язык-а…]. Nomination terms in subject position (Nтерм) are identified by the
word-formation markers, mainly suffixes: -ит- (лигнит, хромит), ант- (инвариант, адресант),
ид, (хлорид, пестицид), -оид- (коллоид, суффиксоид), -ин- (экзотоксин, папаверин),
-ема (синтаксема, фразема), -а-ци-я-, а/ ени-е, -изм, -ость, -ств-о, etc. Occurrence of commonly
used nouns with a similar morphemic structure (or sound composition) is very rare (names like
апельсин, маргарин in such syntactic models are unlikely).</p>
      <p>The predicative position in this model is replaced by highly generalized categories of objects,
attributes, processes (unit, type, class, type, variety, variant; category, concept; process, action,
construction, construct, education; attribute, property, nomination, word, material, substance;
component, part, section, fragment [+ N2], etc.), which subordinate the differentiating components
required in this model — coordinated or uncoordinated attributes. For example,
Кванторные местоимени-я — это языков-ые эквивалент-ы логическ-их оператор-ов – [кванторов
существования и общности].</p>
      <p>Variations of the studied model are the structures with the anaphoric pronoun это (this, that) in
the subject position (its antecedent is contained in the immediately preceding sentence; E.g., Это —
класс позиционно чередующ-ихся звук-ов) and the structures with relative clauses, subordinate to
N1 at the predicative position which is the antecedent of the relative pronoun который (N1 —
это … [так-ой] N1 [Prep] котор...; E.g.: Дискурсивн-ая деятельность-ø —
это такая разновидность-ø речев-ой деятельност-и, [Prep] котор-…).
3.3.</p>
    </sec>
    <sec id="sec-6">
      <title>Context-free template matching</title>
      <p>In the first step, we use MyStem [32] and UDPipe [33] to detect lexemes, morphological features,
and syntactic dependencies in the analyzed text clauses. Then we obtain predicate-argument structures
and semantic roles with the parser developed at the FRC CS&amp;C RAS [34]. Finally, we combine all
those results to form relational-situational models of each clause. G. Osipov defines a
relationalsituational model as a heterogeneous semantic network (HSN) with the following structure [35]:
 =&lt;  ,  ,  ,  ,  &gt;,
(1)
Where
 = { 1,  2, … ,   } is a set of feature sets.</p>
      <p>S – a set of tuples like &lt;   , ∆ &gt; and   – a value of a syntaxeme from a name set  , ∆ ⊆   =
  ×   × … .×   – features of the syntaxeme for each  = 1,2, … , | | and  ≤  .</p>
      <p>R is a set of relationships on N2.</p>
      <p>F is a set of functions  →   , j=1,..,m.</p>
      <p>In other words, the relational-situational model is an HSN, in which vertices are syntaxemes and
edges define semantic relationships between the vertices. These syntaxemes are minimal indivisible
semantic-syntactic structures of language. In our case, set D contains several morphological features
(POS tags, grammatical cases, and moods, etc.) and embeddings of the lexemes. We use pre-trained
FastText (ruscorpora_none_fasttextskipgram_300_2_2019) [36, 37] to build the character-level
embeddings to deal with lexical richness and potential misspellings.</p>
      <p>Let every clause of the analyzed text can be represented with this model, therefore all the text is
the following sequence of HSNs:  =&lt;  1,  2, . . ,   &gt;.</p>
      <p>Define the context-free templates as a tuple  =&lt;  1,  2, . . ,   &gt; of HSNs. We have built more
than 100 such templates based on the cognitive-communicative action markers presented in the
previous section. Since the HSNs from T are templates, they hold only the feature descriptions
essential for the identification. Here we presume that all the HSNs from T and H have the same
feature sets D.</p>
      <p>Eventually we define a reflection:  :   ×   → {0,1} in the following manner. For every text
model &lt;  ,   ,   ,   ,   &gt;∈  and every template &lt;  ,   ,   ,   ,   &gt;∈  we set   = 1 if  ∩   =
  , and  ∩   =   , otherwise we set  = 0. Fig. 2 represents an example of such matching, where
the whole network represents a clause, and the red fragment is the part that matches a template.</p>
      <p>Although the vital question remains is how to intersect syntaxeme values. In our implementation,
the analyzer matches two syntaxemes from different HSNs if they have the same morphology features
and syntax dependencies (defined in the template’s HSN), and the cosine distance between their lexis
embeddings is less than the empirically-defined threshold. It is worth noting that this
templatematching process is context-free. That means this process can be implemented with the algorithms,
which have good performance and applicable for analyzing large texts.</p>
      <p>Eventually, each clause is represented with a binary vector, which encodes if the clause matches
the templates. Since these vectors are sparse, it is reasonable not to use them for further training
directly but to build some dense embeddings before. We applied the implementation of SVD
transformation from Scikit-Learn [38] to generate them.</p>
      <p>In the second step, we apply machine learning to train the models to classify the clauses. Since one
clause can represent several actions, we tackle this classification problem as a multi-label one. We
have tested the following machine-learning methods with the sliding window approach:
1. Decision tree ensembles, such as a Random forest [39] and Gradient boosting on decision
trees (XGBoost implementation) [40].</p>
      <p>2. Linear SVM based classifier with L2 regularization.</p>
      <p>
        The training corpus is relatively small; therefore, in this study, we did not use recurrent networks
or CRF for sequence labeling because they tend to over-fit even with our small artificial feature set, as
we revealed earlier for the empirical text parser [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>Since both SVM and tree ensembles are single-label classifiers, we have trained an independent
binary classifier for each cognitive action. We selected all the training hyperparameters and the
window size for the classifiers with a grid search on three-fold cross-validation.
3.5.</p>
    </sec>
    <sec id="sec-7">
      <title>Clause disambiguation</title>
      <p>In the last step, we have to clarify the type of clauses, labeled with several cognitive actions. In
more detail, the task is to build an mXn mapping between clause’s tokens and labels. Since the
annotated corpus is small, we have automatically labeled an additional dataset with the first two
components of our approach and define this task as the multi-instance learning (MIL) problem [41].
According to the MIL, each clause (bag) is marked with a class label if this clause contains at least
one token corresponding to that class. However, the problem definition, in our case, is non-classical.
Instead of label ranking, we train a model to score a clause’s tokens for each given label. Namely, we
train Pointer Network-based regressor [29] to distinguish token labels inside the clause. The templates
are again used to build features in this step, but now we apply them in the level of distinct tokens (Fig.
2).</p>
    </sec>
    <sec id="sec-8">
      <title>4. Experiments and Results</title>
      <p>The evaluation of classification scores was carried out using the statistical procedure of
crossvalidation [42]. Table 3 shows the normalized confusion matrix (Here, we hold the same numeration
as in the introduction). Each cell of the table (i, j) contains the ratio of text fragments with action i,
incorrectly classified as an action j. The table shows that classification errors are mainly associated
with the recognition of actions “1.1”, “1.2” (“Presentation of theories that form available
knowledge”, “Author's assessment of available knowledge”) and “2.1” (“Definition of a concept”). In
contrast, the rest of the cognitive actions are detected with an insignificant level of errors. The reason
for that may lie in the simplicity of models we use, which catch quite a narrow context of a clause;
therefore, this issue can be fixed if we label more data and use more complex models.</p>
      <p>We do not provide the evaluation results for the clause disambiguation step because there is no
representable manually labeled corpus for that, and this is out of bonds of our proof-of-concept study;
therefore, that is a topic of further research.</p>
    </sec>
    <sec id="sec-9">
      <title>5. Conclusion</title>
      <p>The experiments show that the proposed cognitive actions can be accurately identified in scientific
texts. We suppose that considering the patterns of speech perception, which is not assumed in pure
statistical methods of the broadest possible set of linguistic features, can be the step towards
developing intelligent systems focused on human cognitive functions. At the same time, it is essential
that the object of our study, in contrast to others, is not only the introductory part of a scientific text or
its largely standardized sections but the entire text.</p>
      <p>The implemented parser can also be useful for solving applied problems related to science
development. These problems include generating article’s abstracts, detecting promising research
areas, subject and methodological gaps, or denoting interdisciplinary interests, etc.</p>
      <p>Although we got quite accurate results, it is clear that the further extension of the training corpus is
necessary if we are building a full cognitive action parser. It seems that active learning would be an
appropriate framework for that. It should be noted that the most complex and time-consuming part of
the study is obtaining the linguistic markers. However, we believe there is the possibility to tackle this
issue. It has been shown that linguistic models, such as BERT [25], can identify syntactic and
discourse features in their intermediate layer representations. Therefore, the required linguistic
markers could be extracted in a (semi-) automated manner if one had a reliable approach to the
extraction of those features from the models.</p>
    </sec>
    <sec id="sec-10">
      <title>6. Acknowledgements</title>
      <p>This study is supported by Russian Foundation for Basic Research, grant No 17-29-07049 ofi_m.</p>
    </sec>
    <sec id="sec-11">
      <title>7. References</title>
      <p>[22] X. Li, G. Burns, N. Peng Discourse tagging for scientific evidence extraction //arXiv preprint
arXiv:1909.04758, (2019).
[23] A. B. Ayed, I. Biskri, J.G. Meunier Automatic Text Summarization: A New Hybrid Model Based
on Vector Space Modelling, Fuzzy Logic and Rhetorical Structure Analysis. in: International
Conference on Computational Collective Intelligence, Springer, Cham, 2019, pp. 26-34.
[24] S. Hou, R Lu. “Knowledge-guided unsupervised rhetorical parsing for text summarization”</p>
      <p>Information Systems, 94 (2020): 101615.
[25] Z. Zhu et al. Examining the rhetorical capacities of neural language models //arXiv preprint
arXiv:2010.00153, (2020).
[26] C. D. Manning et al. “Emergent linguistic structure in artificial neural networks trained by
selfsupervision” Proceedings of the National Academy of Sciences (2020).
[27] Z. Yang et al. “Xlnet: Generalized autoregressive pretraining for language understanding”</p>
      <p>Advances in neural information processing systems (2019): 5753-5763.
[28] X. Lin et al. A unified linear-time framework for sentence-level discourse parsing //arXiv
preprint arXiv:1905.05682, (2019).
[29] O. Vinyals, M. Fortunato, N. Jaitly “Pointer networks” Advances in neural information
processing systems (2015): 2692-2700.
[30] Mental actions dataset. http://nlp.isa.ru/mental_actions, last accessed 26/11/2020/.
[31] D.A. Devyatkin, Y. M. Kuznetsova ”Mentalnye deistviya i predmety v prostranstve nauchnogo
discursa [Mental Actions and Mental Objects in the Space of Science Discourse]” Iskusstvennyi
intellect I privyatie reshenyi, 1 (2020): 50-69. (In Russian).
[32] Mystem analyzer, https://tech.yandex.ru/mystem/doc/index-docpage, last accessed 26/11/2020
[33] M. Straka, J Straková. Tokenizing, pos tagging, lemmatizing and parsing ud 2.0 with udpipe, in:
Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal
Dependencies, 2017, pp. 88-99.
[34] A. Shelmanov, D. Devyatkin Semantic role labeling with neural networks for texts in Russian,
in: Computational Linguistics and Intellectual Technologies. Papers from the Annual
International Conference" Dialogue, 2017, 1, pp. 245-256.
[35] G. S. Osipov, I. V. Smirnov, I. A. Tikhomirov “Relational-situational method for text search and
analysis and its applications “Scientific and Technical Information Processing, 37 6 (2010):
432437.
[36] A. Kutuzov, E. Kuzmenko Building web-interfaces for vector semantic models with the
webvectors toolkit, in: Proceedings of the Software Demonstrations of the 15th Conference of
the European Chapter of the Association for Computational Linguistics, 2017, pp. 99-103.
[37] T. Mikolov et al. Advances in pre-training distributed word representations, arXiv preprint
arXiv:1712.09405, (2017).
[38] F. Pedregosa et al. “Scikit-learn: Machine learning in Python” The Journal of machine Learning
research. 12 (2011): 2825-2830.
[39] L. Breiman “Random forests” Machine learning 45 1 (2001): 5-32.
[40] J. H. Friedman “Stochastic gradient boosting” Computational statistics &amp; data analysis 38 4
(2002): 367-378.
[41] T. G. Dietterich, R. H. Lathrop, T. Lozano-Pérez “Solving the multiple instance problem with
axis-parallel rectangles”, Artificial intelligence, 89 1-2 (1997): 31-71.
[42] P. Flach Machine learning: the art and science of algorithms that make sense of data, Cambridge
University Press, 2012.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>M.M. Bakhtin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Holquist</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGee</surname>
            <given-names>V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Emerson</surname>
            <given-names>C.</given-names>
          </string-name>
          <article-title>The Problem with Speech Genres. Speech Genres and Other Late Essays</article-title>
          . Austin: University of Texas Press,
          <year>1986</year>
          ,
          <volume>43</volume>
          pages.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>M.M. Bahtin</surname>
          </string-name>
          ,
          <article-title>Pod maskoj</article-title>
          .
          <article-title>Maska vtoraya. [Behind the mask</article-title>
          . Second mask] (In Russian),
          <string-name>
            <given-names>P.N.</given-names>
            <surname>Medvedev</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.M. Bakhtin</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          <string-name>
            <surname>Wehrle</surname>
          </string-name>
          , “
          <article-title>The formal method in literary scholarship: A critical introduction to sociological poetics”</article-title>
          , Harvard University Press,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.A.</given-names>
            <surname>Salimovsky</surname>
          </string-name>
          <article-title>Zhanry rechi v funktsional'no-stylisticheskom osveshchenii (nauchnij akademicheskij text) [Speech genres in functional stylistic perspective (scientific text)]</article-title>
          . Perm,
          <string-name>
            <surname>PSU</surname>
          </string-name>
          ,
          <year>2002</year>
          . 236 p.
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>A. N.</surname>
          </string-name>
          <article-title>Leont'ev. Activity, consciousness and personality</article-title>
          . Englewood Cliffs, NJ: Prentice Hall,
          <year>1978</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>V.S.</surname>
          </string-name>
          <article-title>Shvyrev Teoreticheskoe i empiricheskoe v nauchnom poznanii [Theoretical and empirical in scientific knowledge]</article-title>
          .
          <source>Nauka</source>
          ,
          <year>1978</year>
          . p.
          <volume>382</volume>
          (in Russian).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>A.S. Maidanov</surname>
          </string-name>
          <article-title>Metodologiya nauchnogo tvorchestva [Methodology of scientific creativity]</article-title>
          . Moscow, LKI Publ.,
          <year>2008</year>
          . 512 p.
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Devyatkin</surname>
          </string-name>
          <article-title>Extraction of Cognitive Operations from Scientific Texts</article-title>
          .
          <source>In: Proceedings of Russian Conference on Artificial Intelligence</source>
          , Springer: Cham,
          <year>2019</year>
          , pp.
          <fpage>189</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.N.</given-names>
            <surname>Kozhina</surname>
          </string-name>
          <article-title>Rechevedenie: teoriya funkcional'noj stilistiki: izbrannye trudy [Speech studies: theory of functional stylistics: selected works]</article-title>
          . ed.
          <source>Flinta: Nauka</source>
          ,
          <year>2014</year>
          . 624 p.
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>B. N.</surname>
          </string-name>
          <article-title>Golovin Osnovy kultury rechi [The basics of speech culture]</article-title>
          . Moscow, Vyssh. sk. Publ.,
          <year>1988</year>
          . 320 p.
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>J.M. Swales</surname>
          </string-name>
          <article-title>Genre analysis: English in academic and research settings</article-title>
          . Cambridge: Cambridge University Press,
          <year>1990</year>
          . 261 p.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[11] J. Swales Research Genres: Explorations and Applications</source>
          . Cambridge: Cambridge University Press,
          <year>2004</year>
          . 314 p.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Teufel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carletta</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Viens</surname>
          </string-name>
          <article-title>An annotationscheme for discourse-level argumentation in researcharticles</article-title>
          .
          <source>in: Proceedings of EACL'99: Ninth Conference of the European Chapter of the Association for ComputationalLinguistics</source>
          ,
          <fpage>8</fpage>
          -
          <lpage>12</lpage>
          June 1999. University of Bergen, Norway,
          <year>1999</year>
          , pp.
          <fpage>110</fpage>
          -
          <lpage>117</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Liakata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Teufel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siddharthan</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Batchelor Corpora for conceptualisation and zoning of scientificpapers</article-title>
          .
          <source>in: Proceedings of the 7th International Conference onLanguage Resources and Evaluation</source>
          . Paris, France: ELDA.LREC,
          <year>2010</year>
          , pp.
          <fpage>2054</fpage>
          -
          <lpage>2061</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.I.</given-names>
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.M.</surname>
          </string-name>
          <article-title>Swales “Strengthening move analysis methodology towards bridging the function-form gap” Journal of English for Academic Purposes</article-title>
          .
          <volume>50</volume>
          (
          <year>2017</year>
          ):
          <fpage>40</fpage>
          -
          <lpage>63</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          ,
          <source>in: Proc. 2019 Conf. North Am. Chapter Assoc. Comput. Linguist. Hum. Lang. Technol. NAACL-LT</source>
          <year>2019</year>
          ,
          <article-title>Minneapolis</article-title>
          , MN, USA, June 2-7,
          <year>2019</year>
          , Vol.
          <volume>1</volume>
          (
          <string-name>
            <given-names>Long</given-names>
            <surname>Short Pap</surname>
          </string-name>
          .,
          <year>2019</year>
          : pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .https://aclweb.org/anthology/papers/N/N19/N 19-1423/.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radrof</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <source>Language Models are Unsupervised Multitask Learners</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bruner</surname>
          </string-name>
          <article-title>Psihologiya poznaniya</article-title>
          .
          <article-title>Za predelami neposredstvennoj informacii [Psychology of cognition. Beyond the immediate information]</article-title>
          .
          <source>Progress</source>
          ,
          <year>1977</year>
          , 413 p.
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.V.</given-names>
            <surname>Zaporozhec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.A.</given-names>
            <surname>Venger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.P.</given-names>
            <surname>Zinchenko</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.G.</surname>
          </string-name>
          <article-title>Ruzskaya Vospriyatie i dejstvie [Perception and action]</article-title>
          .
          <source>Prosveshchenie</source>
          ,
          <year>1967</year>
          , 323 p.
          <article-title>(in Russian)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>A.N.</surname>
          </string-name>
          <article-title>Leontyev O putyah issledovaniya vospriyatiya (vstupitel'naya stat'ya) [On ways to study perception (introductory article)]. Vospriyatie i deyatel'nost'</article-title>
          . / ed. by
          <string-name>
            <surname>Leontyev</surname>
            <given-names>A.N.. MSU</given-names>
          </string-name>
          ,
          <year>1976</year>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>27</lpage>
          . (in Russian).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>W.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thompson</surname>
          </string-name>
          , Sandra A.
          <article-title>"Rhetorical structure theory: toward a functional theory of text organization"</article-title>
          <source>Text: Interdisciplinary Journal for the Study of Discourse</source>
          .
          <volume>8</volume>
          (
          <issue>3</issue>
          ) (
          <year>1988</year>
          ):
          <fpage>243</fpage>
          -
          <lpage>281</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>Karimi</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>Tang Learning hierarchical discourse-level structure for fake news</article-title>
          detection //arXiv preprint arXiv:
          <year>1903</year>
          .
          <volume>07389</volume>
          , (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>