<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Topic Modelling of the Russian Corpus of Pikabu Posts: Author-Topic Distribution and Topic Labelling</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga Mitrofanova</string-name>
          <email>o.mitrofanova@spbu.ru</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Veronika Sampetova</string-name>
          <email>nikasampetova@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ivan Mamaev</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna Moskvina</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kirill Sukharev</string-name>
          <email>sukharevkirill@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Saint-Petersburg Electrotechnical University</institution>
          ,
          <addr-line>Russia, 197376, Saint-Petersburg, ul. Professora Popova, 5</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Speech Technology Center</institution>
          ,
          <addr-line>Russia, 194044, Saint-Petersburg, Vyborgskaya emb, 45E</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>St. Petersburg State University</institution>
          ,
          <addr-line>Russia, 199034, Saint-Petersburg, Universitetskaya emb. 11</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>101</fpage>
      <lpage>116</lpage>
      <abstract>
        <p>The paper discusses development of a corpus of Russian posts with hash tags based on Pikabu social network. We developed a balanced and representative corpus as regards the impact of certain authors, the amount and size of their posts. Our study is aimed at the development of probabilistic topic models revealing the authors' interests and preferences, as well as correlation of topics within the corpus. We performed a series of experiments including standard LDA topic modelling and Author-Topic modelling. In course of topic modelling we used algorithms from Python libraries. Experiments allowed to extract groups of authors with similar and related interests. We used topic label assignment based on manually introduced hash tags and labels automatically extracted from the lexical database RuWordNet. That facilitates linguistic interpretation of results.</p>
      </abstract>
      <kwd-group>
        <kwd>Social Networks</kwd>
        <kwd>Pikabu</kwd>
        <kwd>Russian</kwd>
        <kwd>Topic Modelling</kwd>
        <kwd>LDA</kwd>
        <kwd>ATM</kwd>
        <kwd>Topic Label Assignment</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Information environment gradually penetrated into our daily life, and the growth of
network devices gave rise to a peculiar virtual world with its own rules of digital
discourse. Communication within the virtual world is governed by technical equipment
of «speakers» and «listeners», the texts created by the digital discourse exhibit the
features of various types and forms of speech, the roles of «speakers» and «listeners»
turn out to be diversified, communication in itself becomes spectacular, it requires
reinforcement by visual content. Therefore, the study of digital discourse should
combine methods of cognitive linguistics, content analysis, computational linguistics,
sociology and adjacent fields of knowledge.</p>
      <p>At present the attention of computational linguists and sociologists is focused on
multilevel analysis of social media texts, the core tasks to be solved in empirical
stud</p>
      <p>Copyright ©2020 for this paper by its authors.</p>
      <p>
        Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
ies being author profiling [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ] and topical analysis of online communities [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5,
6</xref>
        ]. These tasks require corpora collection from web-sources and software elaboration.
Proper linguistic processing of social media corpora opens wide opportunities for
studies of social opinion and Russian web discourse, cf. recent publications of
E. Koltsova and colleagues, LINIS HSE [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and SCILA [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]; T. Litvinova and
colleagues, RusProfiling Lab [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]; S. Bodrunova, I. Blekanov and colleagues,
WebMetrics Research group, SPbSU [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], etc.
      </p>
      <p>
        Our present study is devoted to the creation and processing of the social media
corpus containing posts of various authors from the social network Pikabu [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which
has not been properly investigated. Pikabu is a Russian language community founded
in 2009 which is considered as an elaborated analogue to Reddit [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The rise of
attention to Pikabu was caused by the famous meme Zhdun which flooded social media in
2017. The attractiveness of Pikabu as a source of linguistic data is explained by the
medium size of posts (they are not so brief as in Twitter) and by the abundance of the
users’ hash tags indicating the subject matter of the posts.
      </p>
      <p>
        Most corpora developed for Russian social networks use Twitter, LiveJournal,
Facebook, VKontakte as sources of textual data, e.g., Taiga social network segment
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], GICR [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Twitter sentiment corpus [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and the like. Such social media
corpora are used for elaboration and evaluation for NLP algorithms, models and tools, cf.
Dialogue Evaluation Competition [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Social media provide a huge platform for
studying topics, opinions, discourse structure of web-communication, that requires
corpus-based linguistic resources: lexical databases, sentiment lexicons, formal
ontologies, etc. Nowadays predictive distributed word representations are in great demand
for text classification, collocation and construction analysis, that’s why research
community welcomes access to word2vec embedding models pretrained on various
text corpora, and social media corpora as well (cf. RusVectōrēs [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]). In the previous
work we described and evaluated word2vec models for Pikabu [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] which prove to be
useful for further studies of the given source.
      </p>
      <p>
        In course of experiments we process a newly developed dataset of the Russian
Pikabu corpus by means of state-of-the-art algorithms and NLP tools. That gives us the
opportunity to form a baseline for further elaboration of our methodology. For the
first time we carry out experiments on author-topic modelling of the corpus and
obtain data on thematic coherence of posts and on authors’ covert clusters. The novelty
of our study consists in thorough linguistic interpretation of topic models strengthened
by topic label assignment which takes into account hash tags introduced by the
authors as well as labels automatically extracted from RuWordNet [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] lexical database.
Thereby, our research fills in the gaps existing in contemporary Russian corpus
linguistics and social media analysis.
1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Topic Modelling</title>
      <p>
        In recent years we witness the rise of interest in the development and application of
topic modelling as a research procedure for data mining and content analysis [
        <xref ref-type="bibr" rid="ref17 ref18 ref19">17, 18,
19</xref>
        ]. In fact, topic modelling is a variety of fuzzy clustering performed for words
and documents, latent semantic relations within a corpus in this case being described
in terms of a family of probability distributions over a set of topics [
        <xref ref-type="bibr" rid="ref20 ref21 ref22">20, 21, 22</xref>
        ].
      </p>
      <p>Early versions of topic models were based on algebraic transformations, e.g.
classical Vector Space Model (VSM) and Latent Semantic Analysis (LSA), which take
into account term-document distribution, term co-occurrence frequency, and may be
expanded with dimensionality reduction techniques, e.g. Singular Value
Decomposition (SVM). Gradually these models gave way to probabilistic topic models, the most
notable of them being Probabilistiс Latent Semantic Analysis (PLSA), Latent
Dirichlet Allocation (LDA), Expectation-Maximization (EM) Algorithm, etc.</p>
      <p>Probabilistic topic models are based on the assumptions that ordering of words
within documents and documents within a corpus may be ignored; frequent and rare
words do not affect the quality of the topic model; a topic t should be considered as a
discrete distribution over a set of words w, and a separate document d as a discrete
distribution over a set of topics t; the occurrence of words in a document d is
determined by a particular distribution p(w|t).</p>
      <p>
        In our study we use a topic model which is based on Latent Dirichlet Allocation,
our choice is explained by the advantages of LDA compared with previously
developed methods, as well as its availability in a set of libraries, including gensim [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]
and scikit-learn [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] for Python.
      </p>
      <p>Topic modelling allows to build multimodal models which include metadata
alongside with intrinsic features of a corpus, e.g., polylingual models for information
retrieval, temporal models taking into account the time of document creation,
authortopic models which include authorship parameter. The latter type of topic models
satisfies the conditions of our experiments.</p>
      <p>
        The Author-Topic model (ATM) [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] may be considered as a refined extension of
LDA, combines a topic model reproducing relations between words, documents and
topics, and an author model describing relations between documents and authors. In
recent years ATM is often used in linguistic and sociological studies, especially in the
tasks of user profiling (age and gender detection [
        <xref ref-type="bibr" rid="ref26 ref27">26, 27</xref>
        ]) and authorship attribution
[
        <xref ref-type="bibr" rid="ref28 ref29 ref30">28, 29, 30</xref>
        ].
2
2.1
      </p>
      <sec id="sec-2-1">
        <title>Corpus collection</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Development of a corpus of Pikabu Russian posts</title>
      <p>
        The corpus of Pikabu Russian posts includes texts downloaded from Pikabu social
network. Collection of posts was carried out with the help of Pikabu parser adapted
from [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. The parser was developed for Python 3.7 [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] and maintained with lxml
[
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] and requests [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] libraries.
      </p>
      <p>
        We improved the original parser by adding the option of arranging posts as regards
their authorship. We also added an option of post filtering: deletion of non-Russian
texts, images, media-content, punctuation marks, etc. Html-pages parsing was
performed by means of BeautifulSoup library [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. We parsed no less than 100 posts for
each author, so preliminary selection of the most productive authors was necessary.
We took into account productivity ratings [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] which were published in 2017−2018
but still preserved their actuality in 2019. The parsed posts were ordered from the
latest (2019 – end of 2018) to the earliest (middle 2018 and further). Each post was
saved in a *.txt file, its name containing the author’s ID. Joint data about the authors
being saved in a separate document. After preprocessing the corpus size turned out to
be 2 161 681 tokens, the total number of texts being 3 059, maximum number of texts
for a single author – 100, minimal number of texts for a single author – 16. Some of
the authors fell out of the final list of Pikabu users as they posted texts in the image
format, e.g. Oblomoff (cooking recipes in JPEG) и IriskaVRF (comics/pictures).
2.2
      </p>
      <sec id="sec-3-1">
        <title>Corpus processing</title>
        <p>
          Corpus processing included tokenization, lemmatization and stop-word removal was
performed by means of spaCy [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ] and pymorphy2 [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ] libraries. We modified
standard stop-word list by adding Wiktionary data [
          <xref ref-type="bibr" rid="ref39">39</xref>
          ] which was parsed by means of
BeautifulSoup library [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] in order to extract interjections, pronouns, particles,
prepositions, parenthetic expressions, numerals. We added proper names and pejoratives
into a stop-word list. All in all, the size of a stop-word list is 1400 items. After
stopword removal the corpus size reduced to 1 144 812 tokens which constitutes about
53% of the initial corpus size.
        </p>
        <p>
          Before topic modelling we detected bigrams and trigrams in the corpus. In most
cases topic models are designed as unigram models which don’t take into account
regular syntagmatic relations in contexts. At the same time such models may fail to
reflect lexical constructions (collocations and idioms) which are broken into separate
lemmata, the content of the whole phrase being lost [
          <xref ref-type="bibr" rid="ref40 ref41">40, 41</xref>
          ]. That’s why we came to a
conclusion that in the process of topic model development bigrams and trigrams with
frequency more than 20 should be added into the dictionary of the model. In our case
retrieved n-grams turned out to be frequent functional set expressions, e.g.
любой_случай, всякий_случай ‘any_case’, etc., that’s why only a few of them occurred
among top 10 topic words in the output.
        </p>
        <p>We also compressed the dictionary by omitting high- and low-frequency items, so
that the final size of the corpus turned out to be 8320 tokens in 3059 documents of 39
authors.
2.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Hash tag analysis</title>
        <p>Alongside with posts we parsed users’ hash tags which could facilitate linguistic
interpretation of the topics generated by our models. We selected three most frequent
hash tags for each author, e.g. видео ‘video’, мое ‘my’, длинностекст ‘long-text’,
длиннопост ‘long-post’, etc.; hash tags duplicating users’ names: varlamov, mtd,
goodmix, etc.</p>
        <p>
          While processing frequent hash tags we joined semantically correlated hash tags
which could possibly introduce a single topic, e.g.: авторский рассказ ‘author story’
→ рассказ ‘story’; строительная история ‘building story’ → строительство
‘building’; кулинария, рецепт ‘cooking, recipe’→ кулинария ‘cooking’, etc. In cases
of low interpretability of hash tags we borrowed topic labels from the community
titles: e.g. the user Region89 [
          <xref ref-type="bibr" rid="ref42">42</xref>
          ] uses the hash tag bash im as the most frequent one,
instead of it we labeled his posts by group names Истории из жизни ‘Life stories’,
Лига диетологов ‘League of nutritionists’, etc. In some cases hash tags admitted
generalization: e.g. country → continent: США, Канада ‘USA, Canada’ → Северная
Америка ‘North America’; Бразилия, Латинская Америка ‘Brazil, Latin America’
→ Латинская Америка ‘Latin America’; Уганда, Руанда ‘Uganda, Rwanda’ →
Африка ‘Africa’, etc.: hyponym → hypernym: андроид ‘android’, ios → телефон
‘telephone’; супергерой, комиксы ‘superhero, comics’ → комиксы ‘comics’, etc.
The given hash tag transformations were necessary for evasion of false diversity of
topics which could complicate data analysis. Resulting correspondencies are
illustrated in Table 1.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Topic Modelling of a corpus of Pikabu Russian posts</title>
      <p>
        LDA topic models were developed for the whole corpus and for its subcorpora. We
used gensim library [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] to train the models and pyLDAvis library [
        <xref ref-type="bibr" rid="ref43">43</xref>
        ] for Python to
visualize topic distributions. The procedure includes 3 stages:
1) LDA parameter choice and model training;
2) evaluation of topic coherence with UMass-measure [
        <xref ref-type="bibr" rid="ref44">44</xref>
        ];
3) topic visualization.
      </p>
      <p>For LDA development we split the corpus to segments including representing posts
of particular authors. Parameter choice was performed taking into account the variety
of author’s hash tags which envisages a set of expected topics, and UMass measure as
it reflects the level of topic coherence which is treated as a level of human
interpretability of the model based on relatedness of words and documents within a topic:
where
while</p>
      <p>
        corresponds to the number of documents containing words
shows the number of documents contains [
        <xref ref-type="bibr" rid="ref44">44</xref>
        ].
and
UMass measure was calculated for each author’s subcorpus, thus, the most
appropriate number of topics in LDA models for different authors was established ad hoc
(ranging from 3 to 30). In all cases the topic size was settled as 10 lemmata per topic.
Visualization allows to view the most significant lemmata characterizing author’s
subcorpora and word-topic distributions. Below we present topic distributions and
their visualizations for posts of random authors BadVadim, CometovArt,
FrancoDictator and Griffel (cf. Fig. 1 – 4).
Username ID: CometovArt; hash tags: Игры ‘Plays’
Topic 1: игра, ролик, фракция, парад, террана, делать, карта, являться, нормальный,
хороший… ‘play, roller, fraction, parade, terran, make, card, be, normal, good…’
Topic 2: проект, скорость, работа, Бог, проблема, древние, Зот, идея, эфир, хороший…
‘project, speed, work, God, problem, ancient, Thot, idea, airing, good…’
Topic 3: сделать, игра, сезон, серия, момент, играть, трейлер, хороший, увидеть,
делать… ‘make, play, season, episode, moment, play, trailer, good, see, make…’
Topic 4: день, проект, колода, получить, эльф, система, игра, работа, праздник,
подборка… ‘day, project, block, get, elf, system, play, work, holiday, collection…’
Topic 5: орда, альянс, сила, фракция, победить, герой, игра, убить, объединить,
зачистить… ‘horde, alliance, force, fraction, win, hero, play, kill, join, clean…’
      </p>
      <p>
        Subcorpora preparation was strengthened by stylometric analysis performed with
JGAAP toolkit [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ]. Data processing gives evidence in favour of stylometric
parameter diversity between subcorpora and their unity within sets of documents written by
particular authors.
3.2
      </p>
      <sec id="sec-4-1">
        <title>ATM Topic Model</title>
        <p>
          ATM was built by means of gensim library [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. The procedure includes 4 stages:
1) ATM parameter choice and model training;
2) evaluation of topic coherence with UMass-measure;
3) authors’ posts similarity estimation by means of Hellinger distance [
          <xref ref-type="bibr" rid="ref46">46</xref>
          ];
4) modelling of authors’ clusters as regards similarity of their posts.
        </p>
        <p>
          A series of experiments was carried out to choose the appropriate number of topics
for the whole corpus, this number was changed from 10 up to 40 with step 5, the
constant size of topics being 10 lemmata. UMass measure was used to define the best
experimental settings. The highest UMass values corresponded to topic modelling
with 25, 30 and 40 topics. We took into account lowest scores of lemmata repetitions
between topics, that were characteristic of the model with 30 topics, this value was
selected as the best for the tasks of our study. For each author we selected the most
relevant topics generated by ATM and matched them with hash tags. After ATM
constructions we evaluated similarity of authors’ posts by Hellinger distance [
          <xref ref-type="bibr" rid="ref46">46</xref>
          ]
which estimates the distance between probability distributions describing topic variety
for the authors, thus, author clusters were formed within our model. Examples of such
clusters for users BadVadim, CometovArt, FrancoDictator and Griffel are given in
Table 2.
        </p>
        <p>
          Author clusters proved to be consistently interpretable with respect to the content
of their posts. The user Griffel [
          <xref ref-type="bibr" rid="ref47">47</xref>
          ] is a participant of the society «Pikabu Users of
North America», his associates within a cluster being immigrants and/or travelers, e.g.
Varlamov.ru, a well-known blogger writing on urbanistics and adventures [
          <xref ref-type="bibr" rid="ref48">48</xref>
          ]; the
user goodmix is a businessman and sauna proprietor, and his associates turn out to be
builders and repairmen (AnnrR, alekseev77, Scrypto), etc.
        </p>
        <p>CometovArt
Игры
‘Plays’
Griffel
FrancoDictador
Северная
Америка
‘North
America’
Политика,
Испания
‘Politics,
Spain’</p>
        <p>L4rever
AlexGyver
Scrypto
Varlamov.ru
Griffel
Varlamov.ru
ShamovD
FrancoDictad
or
L4rever
Esmys
Varlamov.ru
ShamovD
Griffel
L4rever
CometovArt
0,75
0,70
0,69
0,68
0,68
0,99
0,98
0,97
0,79
0,78
0,99
0,98
0,97
0,79
0,68
Германия,
Бонусы‘Germany, bonuses’
Своими руками ‘With my
hands’
Ремонт техники
‘Equipment repair’
Городская среда,
архитектура ‘Urban
environment, architecture’
Северная Америка ‘North
America’
Городская среда,
архитектура ‘Urban
environment, architecture’
Япония ‘Japan’
Политика, Испания
‘Politics, Spain’
Германия,
Бонусы‘Germany, bonuses’
Латинская Америка,
кулинария ‘Latin America,
cooking’
Городская среда,
архитектура ‘Urban
environment, architecture’
Япония ‘Japan’
Северная Америка ‘North
America’
Германия,
Бонусы‘Germany, bonuses’
Игры ‘Plays’</p>
        <p>ATM allows to distribute users over certain groups in accordance with the major
topics discussed in the posts. We will mention only a few examples: user group of
travelling (Griffel, ShamovD, FrancoDictador, L4rever, Esmys, etc. describe different
cities and countries), user group of narrators (MadTillDead, smile2, CreepyStory,
DoktorLobanov, denisslavin, ozymandia, femme.kira, svoemnenie, 889900, Region89,
etc. write short stories), user group of builders and repairmen (alekseev77, AnnrR,
BadVadim, Scrypto, AlexGyver, etc.).
3.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Topic Label Assignment</title>
        <p>There are certain approaches to the improvement of topic models. Traditional topic
modelling algorithms, LDA being among them, do not include label assignment as an
internal procedure.</p>
        <p>
          At the same time, topic labelling allows to improve informativeness of the models.
Labels are considered as single terms or phrases generalizing the topic content. From
the semantic point of view, such labels are expected to be either strict hypernyms,
holonyms, or at least more abstract lexical items covering the meaning of separate
topic words. By default one can choose the first word of a topic as a label, but such
labels often turn out to be unconvincing. Although topic words are ranked within a
topic and may be related to each other by various syntagmatic and paradigmatic
relations [
          <xref ref-type="bibr" rid="ref49">49</xref>
          ], such ordering is not obligatorily hierarchical. Labels can be assigned
manually in course of human expertise, but in this case they may reflect subjective
treatment of topics. Thus, it is necessary to find a tradeoff solution to the problem.
Consequently, NLP researches proposed several ways of automatic label assignment [
          <xref ref-type="bibr" rid="ref50 ref51 ref52">50,
51, 52</xref>
          ] which differ in the source of labels (intrinsic data extracted from corpora or
extrinsic information from external resources – lexical databases (WordNet,
Wikipedia, etc.) or search engines (Google, Yandex, etc.). In our previous studies
experiments on automatic label assignment were performed using such external resources as
Russian Wikipedia distributional model accessible via ESA (Explicit Semantic
Analysis) and Yandex search engine output with morphosyntactic parsing and statistical
ranking [
          <xref ref-type="bibr" rid="ref53 ref54">53, 54</xref>
          ]. Taking into account stylistic peculiarities of social media (Pikabu
corpus representing one of them), we expect that labels extracted from encyclopaedic
texts (Wikipedia) and news headings (Yandex) may fail to match with the content of
Pikabu topics. Therefore, in this study we chose a lexical database RuWordNet [
          <xref ref-type="bibr" rid="ref16 ref55">16,
55</xref>
          ] suitable for experiments with social media data of mixed stylistic character.
        </p>
        <p>
          The procedure of label assignment implemented in our project implies a hybrid
approach combining human expertise and automatic data processing, involving internal
and external sources of candidate labels. On the one hand, we use manually assigned
hash tags extracted from Pikabu as topic labels. As hash tags are consciously
introduced by the authors, they may be considered more reliable than the first topical
words. On the other hand, we extract hypernyms for topic words as candidate labels
from RuWordNet. The idea to use lexical hierarchy as a source of topic labels keeps
close to the task of automatic extraction of «IS-A» relations and corpus-based
taxonomy enrichment. The procedure used in our study implied selection of hypernyms for
each topic word which were united in a list and ranked. Both hash tags and
hypernyms are ranked in accordance with ipm frequencies from A Frequency Dictionary of
Contemporary Russian (based on the Russian National Corpus) by O.N.Lashevskaya
and S.A.Sharoff [
          <xref ref-type="bibr" rid="ref56">56</xref>
          ].
        </p>
        <p>Samples from our dataset are described in Table 3. As expected, in all four cases
the users’ hash tags provide the best fit. It should be noted that throughout the corpus
hash tags are repeated among top 10 topical words, but not necessarily at the head
part. This gives us the reasons to consider them as a solid baseline dataset in further
experiments.</p>
        <p>As regards the first topical words, in our example they are suitable as topic labels
for the topics extracted from the posts of the users Brahmanden and yulianovsemen,
but it is not the case as regards the users upitko and Malfar: although the words
готовить ‘prepare’ и дело ‘case’ have abstract meanings, they turn out to be too
ambiguous and vague for being topic labels. As for the corpus in general, the set of top one /
three topical words is rather heterogeneous in meaning, so we may consider them as a
tentative – less reliable as hash tags – dataset for evaluation procedure.
Author</p>
        <p>Topic example
Labels extracted from RuWordNet and assigned to the main topics of the users
Brahmanden and upitko correspond to the content of the topics and partially intersect with
hash tags on the lexical level (repetitions are marked in bold: кулинария, фильм,
рецензия ‘cooking, film, review’). That doesn’t hold true for the users yulianovsemen
and Malfar: RuWordNet labels turn out to be rather more general than users’ hash
tags and topical words (преступление, право, издание ‘crime, law, edition’). All in
all, according to our observations, RuWordNet topic labels, being semantically
correlated with the topics, seem to be rather generalized in comparison with label
candidates selected by other methods. In order to improve the results we upgraded the
procedure of hypernym selection by using word2vec embeddings extracted from the
pretrained corpus model: we assumed that possible label vectors may be similar to
averaged topic vectors, but our expectations were partially fulfilled as candidate
labels enhanced by word2vec data remained general by meaning. That inspires further
experiments with combination of topic labeling with distributed vector
representations.</p>
        <p>However, three types of labels constitute a scale of acceptability which is limited
by the first topical words as formal labels and RuWordNet labels as the most general
ones, the golden mean being the users’ hash tags. The combination of expert-based
and knowledge-based approaches to topic label assignment requires further
quantitative analysis and evaluation, but even in the qualitative aspects it seems to be fruitful
as it provides data for topic expansion and may be useful in text rubrication.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Summary</title>
      <p>In the given study we managed to create an author-topic model for the corpus of
Pikabu Russian posts.</p>
      <p>We worked out a procedure for corpora development suitable for processing mixed
data from social media: texts and metadata (author’s usernames and hash tags). A
multilevel analysis of our corpus was performed.</p>
      <p>We developed standard LDA models with visualization for subcorpora containing
posts of separate authors, that provides data on their interests. A complex
AuthorTopic model was constructed for the whole corpus, which allows to detect clusters of
authors writing on similar topics.</p>
      <p>Finally, we carried out experiments on topic label assignment, topic labels being
obtained from two sources: manually assigned users’ hash tags and hypernyms for
topical words automatically extracted from RuWordNet lexical database.</p>
      <p>Results achieved by now allow to expand our studies and put forward the next set
of tasks: development and processing of various social media corpora, elaboration of
the procedure for latent community detection, enhancement of the procedure of
hypernym extraction and ranking for automatic label assignment.</p>
      <p>Acknowledgements. The authors express their sincere gratitude to Dr. Prof. Natalia
Loukachevich (Moscow State University) and colleagues for the access to RuWordNet thesaurus.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Panicheva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Litvinova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Litvinova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Author Clustering with and Without Topical Features</article-title>
          . In: Speech and Computer 21st International Conference,
          <string-name>
            <surname>SPECOM</surname>
          </string-name>
          <year>2019</year>
          , Istanbul, Turkey,
          <source>August 20-25</source>
          ,
          <year>2019</year>
          , Proceedings. LNAI, vol.
          <volume>11658</volume>
          , pp.
          <fpage>348</fpage>
          -
          <lpage>358</lpage>
          . Springer (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Litvinova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sboev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Panicheva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Profiling the Age of Russian Bloggers</article-title>
          .
          <source>In: Artificial Intelligence and Natural Language</source>
          , 7th International Conference, AINL 2018,
          <article-title>St</article-title>
          . Petersburg, Russia,
          <source>October 17-19</source>
          ,
          <year>2018</year>
          , Proceedings, issue
          <volume>930</volume>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>177</lpage>
          . Switzerland: Springer (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Panicheva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mirzagitova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ledovaya</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Semantic Feature Aggregation for Gender Identification in Russian Facebook</article-title>
          .
          <source>In: Artificial Intelligence and Natural Language</source>
          , 6th Conference,
          <string-name>
            <surname>AINL</surname>
          </string-name>
          <year>2017</year>
          ,
          <article-title>St</article-title>
          . Petersburg, Russia,
          <source>September 20-23</source>
          ,
          <year>2017</year>
          , Revised Selected Papers, issue
          <volume>789</volume>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>15</lpage>
          . Switzerland: Springer (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>LINIS</surname>
            <given-names>HSE</given-names>
          </string-name>
          , https://linis.hse.ru/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. SCILA, https://scila.hse.ru/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. RusProfiling Lab, https://rusprofilinglab.ru/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. WebMetrics Research group, SPbSU, http://www.apmath.spbu.ru/ru/structure/depts/tp/ webometrics.html,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Pikabu, https://pikabu.ru/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Reddit, https://www.reddit.com,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Taiga</surname>
          </string-name>
          , https://tatianashavrina.github.io/taiga_site/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. GICR, http://www.webcorpora.ru/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. Twitter sentiment corpus, https://study.mokoron.com/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Dialogue Evaluation Competition, http://www.dialog-
          <volume>21</volume>
          .ru/evaluation/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. RusVectōrēs, https://rusvectores.org/ru/models/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Antipenko</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrofanova</surname>
            ,
            <given-names>O.A.</given-names>
          </string-name>
          :
          <article-title>Comparative Study of Word Associations in Social Networks Corpora by means of Distributional Semantics Models for Russian</article-title>
          .
          <source>International Journal of Open Information Technologies</source>
          ,
          <volume>8</volume>
          (
          <issue>1</issue>
          ) (
          <year>2020</year>
          ), http://www.injoit.org/ index.php/j1/article/view/871/834, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. RuWordNet, https://ruwordnet.ru/ru, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Latent Dirichlet allocation</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>3</volume>
          ,
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          . (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Blei</surname>
          </string-name>
          , J.:
          <article-title>Probabilistic topic models</article-title>
          .
          <source>In: Communications of the ACM</source>
          , vol.
          <volume>55</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>77</fpage>
          -
          <lpage>84</lpage>
          . (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Vorontsov</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potapenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Tutorial on Probabilistic Topic Modeling: Additive Regularization for Stochastic Matrix Factorization</article-title>
          . In: D.I. Ignatov,
          <string-name>
            <given-names>M.Y.</given-names>
            <surname>Khachay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Konstantinova</surname>
          </string-name>
          , R. Yavorskiy (Eds.).
          <source>Analysis of Images, Social Networks and Texts</source>
          . Third International Conference, AIST 2014, Yekaterinburg, Russia,
          <source>April 10-12</source>
          ,
          <year>2014</year>
          ,
          <string-name>
            <given-names>Revised</given-names>
            <surname>Selected</surname>
          </string-name>
          <string-name>
            <surname>Papers</surname>
          </string-name>
          ,
          <string-name>
            <surname>CCIS</surname>
          </string-name>
          , vol.
          <volume>436</volume>
          , pp.
          <fpage>29</fpage>
          -
          <lpage>46</lpage>
          . Springer, Cham (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Bodrunova</surname>
            ,
            <given-names>S</given-names>
          </string-name>
          , Blekanov,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Kukarkin</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Topic modeling for Twitter discussions: Model selection and quality assessment</article-title>
          .
          <source>In: Proceedings of the 6th SGEM International Multidisciplinary Scientific Conferences on Social Sciences and Arts SGEM2018, Science and Humanities</source>
          . Sofia, Bulgaria: STEF92
          <string-name>
            <given-names>Technology</given-names>
            <surname>Ltd</surname>
          </string-name>
          ., pp.
          <fpage>207</fpage>
          -
          <lpage>214</lpage>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Apishev</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koltcov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koltsova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikolenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Vorontsov</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Additive Regularization for Topic Modeling in Sociological Studies of User-Generated Texts</article-title>
          . In: G. Sidorov &amp; O. Herrera-Alcántara (eds.),
          <source>Advances in Computational Intelligence</source>
          , pp.
          <fpage>169</fpage>
          -
          <lpage>184</lpage>
          . Springer International Publishing (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Nikolenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koltcov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Koltsova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Topic modelling for qualitative studies</article-title>
          .
          <source>Journal of Information Science</source>
          ,
          <volume>43</volume>
          (
          <issue>1</issue>
          ),
          <fpage>88</fpage>
          -
          <lpage>102</lpage>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. gensim, https://radimrehurek.com/gensim/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24. scikit-learn, https://scikit-learn.org/stable/index.html,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Rosen-Zvi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            <given-names>Griffiths</given-names>
          </string-name>
          , Th.,
          <string-name>
            <surname>Steyvers</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The Author-Topic Model for Authors and Documents</article-title>
          , http://arXiv:
          <fpage>1207</fpage>
          .4169, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Use of language and author profiling: identification of gender and age</article-title>
          .
          <source>In: NLPCS 2013 10th International Workshop on Natural Language Processing and Cognitive Science CIRM</source>
          , Marseille, France, pp.
          <fpage>177</fpage>
          -
          <lpage>185</lpage>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Panicheva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mirzagitova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ledovaya</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Semantic feature aggregation for gender identification in Russian Facebook</article-title>
          . In: A.
          <string-name>
            <surname>Filchenkov</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Pivovarova</surname>
          </string-name>
          , J.
          <source>Žižka (Eds.) Artificial Intelligence and Natural Language</source>
          , 6th Conference,
          <string-name>
            <surname>AINL</surname>
          </string-name>
          <year>2017</year>
          ,
          <article-title>St</article-title>
          . Petersburg, Russia,
          <source>September 20-23</source>
          ,
          <year>2017</year>
          , Revised Selected Papers, pp.
          <fpage>3</fpage>
          -
          <lpage>15</lpage>
          . Springer (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Argamon</surname>
          </string-name>
          , Sh.,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>A.J.:</given-names>
          </string-name>
          <article-title>Automatically profiling the author of an anonymous textю Communications of the ACM - Inspiring Women in Computing</article-title>
          ,
          <volume>52</volume>
          (
          <issue>2</issue>
          ),
          <fpage>119</fpage>
          -
          <lpage>123</lpage>
          . (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Seroussi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bohnert</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zukerman</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Authorship Attribution with Author-aware Topic Models</article-title>
          .
          <source>In: Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics</source>
          , vol.
          <volume>2</volume>
          ,
          <string-name>
            <surname>Short</surname>
            <given-names>Papers</given-names>
          </string-name>
          , pp.
          <fpage>264</fpage>
          -
          <lpage>269</lpage>
          . Jeju Island,
          <string-name>
            <surname>Korea</surname>
          </string-name>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , H.,
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Authorship Attribution for Short Texts with Author-Document Topic Model</article-title>
          . In: W. Liu,
          <string-name>
            <given-names>F.</given-names>
            <surname>Giunchiglia</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          Yang (Eds.) Knowledge Science,
          <source>Engineering and Management KSEM</source>
          <year>2018</year>
          ,
          <article-title>LNCS</article-title>
          , vol.
          <volume>11061</volume>
          , pp.
          <fpage>29</fpage>
          -
          <lpage>41</lpage>
          . Springer, Cham (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31. Pikabu parser, https://github.com/silver4one/pikabu_parser,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32. Python 3.7, https://www.python.org/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33. lxml, https://lxml.de/, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34. requests, https://2.python-requests.org/en/master/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35. BeautifulSoup, https://www.crummy.com/software/BeautifulSoup/bs4/doc/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36. Pikabu productivity ratings, https://clck.ru/GK8gP, https://clck.ru/GK8jn, https://clck.ru/GK8kL, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37. spaCy, https://spacy.io/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38. pymorphy2, https://pymorphy2.readthedocs.io/en/latest/,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>Wiktionary</surname>
          </string-name>
          data, https://ru.wiktionary.org/wiki/Заглавная_страница,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <string-name>
            <surname>Loukachevitch</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nokel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ivanov</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Combining Thesaurus Knowledge and Probabilistic Topic Models</article-title>
          .
          <source>In: International Conference on Analysis of Images, Social Networks and Texts</source>
          . Springer, Cham,
          <year>2017</year>
          , pp.
          <fpage>59</fpage>
          -
          <lpage>71</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          41.
          <string-name>
            <surname>Sedova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrofanova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Topic Modelling of Russian Texts based on Lemmata and Lexical Constructions</article-title>
          .
          <source>In: Computational Linguistics and Digital Ontologies</source>
          , vol.
          <volume>1</volume>
          , pp.
          <fpage>132</fpage>
          -
          <lpage>144</lpage>
          . Saint-Petersburg,
          <string-name>
            <surname>ITMO</surname>
          </string-name>
          (
          <year>2017</year>
          ) [in Russian].
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          42. Region89, https://pikabu.ru/@Region89, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          43. pyLDAvis, https://github.com/bmabey/pyLDAvis, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          44.
          <string-name>
            <surname>Stevens</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kegelmeyer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andrzejewski</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buttler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Exploring topic coherence over many models and many topics</article-title>
          .
          <source>In: Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning</source>
          ,
          <source>Jeju Island, Korea, July 12-14</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>952</fpage>
          -
          <lpage>961</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          45. JGAAP, https://github.com/evllabs/JGAAP, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          46.
          <string-name>
            <surname>Nikulin</surname>
            ,
            <given-names>M.S.:</given-names>
          </string-name>
          <article-title>Hellinger distance</article-title>
          .
          <source>In: Encyclopedia of Mathematics</source>
          , Kluwer Academic Publishers (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          47.
          <string-name>
            <surname>Griffel</surname>
          </string-name>
          , https://pikabu.ru/profile/griffel, last accessed
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          48.
          <string-name>
            <surname>Varlamov</surname>
          </string-name>
          .ru, https://varlamov.ru,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          49.
          <string-name>
            <surname>Koltsov</surname>
            ,
            <given-names>S.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koltsova</surname>
            ,
            <given-names>O.Ju.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrofanova</surname>
            ,
            <given-names>O.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shimorina</surname>
            ,
            <given-names>A.S.:</given-names>
          </string-name>
          <article-title>Interpretation of Semantic Relations in the texts of the Russian LiveJournal Segment based on LDA Topic Model</article-title>
          .
          <source>In: Proceedings of the XVIIth All-Russia Joint Conference «Internet and Modern Society» IMS-2014</source>
          , Saint-Petersburg,
          <string-name>
            <surname>ITMO</surname>
          </string-name>
          , November
          <volume>19</volume>
          -
          <issue>20</issue>
          ,
          <year>2014</year>
          , pp.
          <fpage>135</fpage>
          -
          <lpage>142</lpage>
          . Saint-Petersburg,
          <string-name>
            <surname>ITMO</surname>
          </string-name>
          (
          <year>2014</year>
          ) [in Russian].
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          50.
          <string-name>
            <surname>Aletras</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stevenson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Court</surname>
          </string-name>
          , R.:
          <article-title>Labelling Topics using Unsupervised Graph-based Methods</article-title>
          .
          <source>In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics</source>
          , vol.
          <volume>2</volume>
          ,
          <string-name>
            <surname>Short</surname>
            <given-names>Papers</given-names>
          </string-name>
          , pp.
          <fpage>631</fpage>
          -
          <lpage>636</lpage>
          . Baltimore, Maryland,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          51.
          <string-name>
            <surname>Allahyari</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pouriyeh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochut</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arabnia</surname>
            ,
            <given-names>H.R.:</given-names>
          </string-name>
          <article-title>A Knowledge-based Topic Modeling Approach for Automatic Topic Labeling</article-title>
          .
          <source>International Journal of Advanced Computer Science and Applications</source>
          ,
          <volume>8</volume>
          (
          <issue>9</issue>
          ),
          <fpage>335</fpage>
          -
          <lpage>349</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          52.
          <string-name>
            <surname>Bhatia</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lau</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Automatic Labelling of Topics with Neural Embeddings</article-title>
          .
          <source>In: 26th COLING International Conference on Computational Linguistics</source>
          , pp.
          <fpage>953</fpage>
          -
          <lpage>963</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          53.
          <string-name>
            <surname>Erofeeva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrofanova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Automatic assignment of topic labels in topic models for Russian text corpora</article-title>
          .
          <source>Structural and Applied Linguistics</source>
          ,
          <volume>12</volume>
          ,
          <fpage>122</fpage>
          -
          <lpage>147</lpage>
          . St. Petersburg University (
          <year>2019</year>
          ) [in Russian].
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          54.
          <string-name>
            <surname>Kriukova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erofeeva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrofanova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sukharev</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Explicit Semantic Analysis as a Means for Topic Labelling</article-title>
          . In: D.
          <string-name>
            <surname>Ustalov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Filchenkov</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Pivovarova</surname>
          </string-name>
          , J. Žižka (eds.).
          <source>Artificial Intelligence and Natural Language Processing: 7th International Conference, AINL</source>
          <year>2018</year>
          ,
          <article-title>St</article-title>
          . Petersburg, Russia,
          <source>October 17-19</source>
          ,
          <year>2018</year>
          , Proceedings, pp.
          <fpage>167</fpage>
          -
          <lpage>177</lpage>
          . Springer, Cham (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          55.
          <string-name>
            <surname>Loukachevitch</surname>
            ,
            <given-names>N.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lashevich</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerasimova</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ivanov</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dobrov</surname>
            ,
            <given-names>B.V.</given-names>
          </string-name>
          :
          <article-title>Creating Russian WordNet by Conversion</article-title>
          .
          <source>In: Computational Linguistics and Intellectual Technologies: Proceedings of the Annual International Conference «Dialogue»</source>
          , vol.
          <volume>15</volume>
          , pp.
          <fpage>405</fpage>
          -
          <lpage>415</lpage>
          . Moscow: RSUH (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          56.
          <string-name>
            <surname>Lashevskaya</surname>
            ,
            <given-names>O.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharoff</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          :
          <article-title>A Frequency Dictionary of Contemporary Russian (based on the Russian National Corpus)</article-title>
          , http://dict.ruslang.ru/freq.php,
          <source>last accessed</source>
          <year>2020</year>
          /05/08.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>