<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>What's in the News? Identification of Trending Topics in Alternative and Mainstream Lithuanian Media</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Baltic Institute of Advanced Technology</institution>
          ,
          <addr-line>Pilies str. 16, Vilnius 01124</addr-line>
          ,
          <country country="LT">Lithuania</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Justina Mandravickaitė</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vytautas Magnus University</institution>
          ,
          <addr-line>K. Donelaičio str. 58, Kaunas 44248</addr-line>
          ,
          <country country="LT">Lithuania</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>2014</fpage>
      <lpage>2016</lpage>
      <abstract>
        <p>It is no longer surprising that internet media is a significant appliance in reflecting and shaping public opinion. Tracking topics dynamics and focus in different media channels is an important tool for opinion-forming mechanisms and process analysis. Information collect, text analytics and Artificial Intelligence tools allows identification of trending topics in different media sources, while exploratory visual analytics tools provide means to identify prevalence of topics in different sources, and their dynamics. In this paper we discuss an ongoing research and demonstrate applicability of such approach to main Lithuanian news portal (delfi.lt) and alternative unconventional media channels - sarmatas.lt and netiesa.lt.</p>
      </abstract>
      <kwd-group>
        <kwd>Topic modelling</kwd>
        <kwd>Framing</kwd>
        <kwd>Media Monitoring</kwd>
        <kwd>NLP</kwd>
        <kwd>Lithuanian language</kwd>
        <kwd>Artificial Intelligence</kwd>
        <kwd>LDA</kwd>
        <kwd>stm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Internet media is an important tool in reflecting and shaping public opinion. Modern
tools and technologies allow automatic tracking and comparing dynamics of different
topics in different media channels, and analysis of the results using visual tools. We
apply a set of such tools for the two types of Lithuanian news portals: main WWW
news channel - delfi.lt1 and two alternative unconventional media channels -
sarmatas.lt2 and netiesa.lt3. We apply topic modelling methods for (trending) topics
identification, and visual results for the further analysis.</p>
      <p>Topic modelling is a text mining technique to discover common topics in a collection
of documents. In practice researchers attempt to fit appropriate model parameters to the
data corpus using one of several heuristics for maximum likelihood fit.</p>
      <p>Copyright © 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>
        Media Framing Dynamics of the ‘European Refugee Crisis’ is analyzed in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This
study investigates the national media discourses in Hungary, Germany, Sweden, the
United Kingdom and Spain for this time period. LDA was applied 130,042 articles in
5 languages from 24 news outlets. It shows that country-specific media tracks the
overall course of the refugee debate, uncovers dynamics and shifts in discourses.
      </p>
      <p>
        Turkish news analysis is presented in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The dataset consists of 4200 Turkish news
titles belonging to 7 classes. NMF was the most successful method for three classes,
while for five and seven classes LSA was the most successful method. Comparative
study is presented in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] as well.
      </p>
      <p>
        There is an interesting study of topic modelling of news articles for two consecutive
elections in South Africa [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Articles are classified using pairwise cosine similarity to
identify similar topics in different periods of elections.
      </p>
      <p>
        Critical evaluation of the utility of the thematic grouping of texts into ‘topics’
emerging from a large collection of online patient comments about the National Health
Service (NHS) in England is presented in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Results show that topic modelling allowed
to group texts into topics that were truly thematically coherent with a mixed degree of
success, while the more traditional approaches to discourse analysis consistently
provided a more nuanced perspective on the data which was ultimately closer to the reality
of the texts it contains.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] paper, authors describe their work in developing a model for topic modelling
and detection of hot topics being discussed in the local Malay news publisher. This
model explored different features for article clustering and topic modelling, and then
applied the TextRank algorithm to identify hot topics in the news.
      </p>
      <p>The tremendous growth of social media content on the Internet has inspired the
development of the text analytics to understand and solve real-life problems. Leveraging
statistical topic modelling helps researchers in better comprehension of textual content
as well as provides useful information for further analysis.</p>
      <p>
        Authors [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] have tested Dengue epidemics tracking using Twitter content
classification and topic modelling. Classifier achieves a prediction accuracy of about 80 % based
on a small training set of about 1,000 instances, but the need for manual annotation
makes it hard to track seasonal changes in the nature of the epidemics, such as the
emergence of new types of virus in certain geographical locations. In contrast,
LDAbased topic modelling scales well, generating cohesive and well-separated clusters from
larger samples.
      </p>
      <p>
        Another experiment with Twitter data set on topic modelling was for identification
of vaccine reactions. The study [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] compared Gensim LDA, MALLET, and
jLDADMM DMM models to determine the most effective model for detecting vaccine
safety signals, assisted by an evaluation process that used an adjusted F-Scoring
technique over a labelled subset of the documents.
      </p>
      <p>
        Paper [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] uses 18,552 tweets dated from 2015 up to 2018 to analyze the dynamics of
the LGBT conversation among Indonesian peoples. In this research, they explore the
main topic of the LGBT conversation using LDA. The result shows that there are seven
main categories that people normally talked about regarding LGBT.
      </p>
      <p>
        Study [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] summarizes the message content of four data sets of Twitter messages
relating to challenging social events in Kenya. They use LDA topic modelling to
analyze the content. This study uses two evaluation measures: Normalized Mutual
Information (NMI) and topic coherence analysis, to select the best LDA models. The
obtained LDA results show that the tool can be effectively used to extract discussion
topics and summarize them for further manual analysis.
      </p>
      <p>
        Investigations can be done with short texts as well. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] conduct a topic modelling
of 6854 Instagram posts made by Ramzan Kadyrov (the head of the autonomous
Chechen Republic in the Russian Federation). Researchers analyze the verbal framing of
24 dominant topics. The study concludes that the main rhetorical device that Kadyrov
employs is a merging of personal and political themes throughout his posts.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data and Methods</title>
      <sec id="sec-2-1">
        <title>Corpora</title>
        <p>
          Corpus consists of 5000 delfi.lt articles (a random sample from News category of
delfi.lt corpus [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]), 1145 sarmatas.lt articles and 2411 netiesa.lt articles, both
published in a period of 2014 – 2016 years. Delfi.lt is the mainstream news portal, the most
readable and visited channel in Lithuania, while sarmatas.lt and netiesa.lt are alternative
source of media in selected geographical indication. Sarmatas.lt is one of the most
important sources in terms of dissemination of information (project Research Meadow4,
2014) and netiesa.lt is unconventional but quite popular news portal among Lithuanian
portal readers.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Methods</title>
        <p>
          Topic analysis is a Natural Language Processing (NLP) technique that allows
automatically extract meaning from texts by identifying recurrent themes or topics. The
goal of the structural topic model is to discover topics and estimate their relationship to
document metadata. LDA is a particularly popular method for fitting a topic model [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
It treats each document as a mixture of topics and handles each topic as a mixture of
words [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. This allows documents to "overlap" with content, rather than grouping them
in a way that reflects the normal use of natural language.
        </p>
        <p>
          The structural topic model allows researchers to flexibly estimate a topic model that
includes document-level metadata [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Estimation is accomplished through a fast
variation approximation. In this research the stm package [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] was used, it provides many
useful features, including rich ways to explore topics, estimate uncertainty, and
visualize quantities of interest. Structural topic modeling operating principle:
1. The generative model begins at the top, with document-topic and topic-word
distributions generating documents that have metadata associated with them;
• a topic is defined as a mixture over words where each word has a probability
of belonging to a topic.
4 http://mokslopieva.lt/, last accessed 2020/03/15
• a document is a mixture over topics, meaning that a single document can be
composed of multiple topics. As such, the sum of the topic proportions across
all topics for a document is one, and the sum of the word probabilities for a
given topic is one.
2. Topical prevalence refers to how much of a document is associated with a topic
(described on the left hand side) and topical content refers to the words used
within a topic (described on the right hand side). Hence metadata that explain
topical prevalence are referred to as topical prevalence covariates, and variables
that explain topical content are referred to as topical content covariates
In this work, the R [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] package stm [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] for structural topic modeling was used.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Overall Process</title>
        <p>
          We used the following process for the analysis:
1. corpora were collected from the corresponding portals (not part of this research);
2. corpora were created from the random sample from delfi.lt and selected sarmatas.lt,
netiesa.lt articles;
3. all texts were lemmatized and lowercased using SpaCy5 Core Lithuania models;
4. stopwords6, numbers, symbols and punctuation marks were removed;
5. documents were represented as bag-of-words (a text is represented as the bag
(multiset) of words, disregarding grammar and even word order but keeping
frequencies.);
6. low frequency words (5% of the least frequent words in the whole corpora) and 5%
of words that occurred in all the texts were removed;
7. Latent Dirichlet Allocation (LDA) [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and stm R function [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] were applied for
structural topic modelling;
8. results were visualized.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>Topic modeling is part of a class of text analysis methods that analyze “bags” or groups
of words together—instead of counting them individually–in order to capture how the
meaning of words is dependent upon the broader context in which they are used in
natural language. So foremost investigation was for finding the expected proportions in
the data (see Fig. 1).
5 https://spacy.io/, last accessed 2020/03/15
6 https://github.com/tokenmill/ltlangpack, last accessed 2020/03/15</p>
      <p>After this approach we have to select and take into account the words with the
highest (raw) probabilities and the highest FREX (Frequency and Exclusivity, i.e., words
that are most frequent and exclusive to the topic) (see Fig. 2).</p>
      <p>We apply LDA for delfi.lt and sarmatas.lt dataset. Analysis shows rather different
combination of topics in the portals, see Fig. 3. In delfi.lt (left) orange topics
(democracy, traffic accidents, referendum on preventing foreigners from owning land in
Lithuania, ceasefire negotiations in Ukraine, activities of the state security department of
Lithuania, etc.) prevail, while in sarmatas.lt (right) blue-purple topics (Islam and
terrorism, industry, Maidan, taxes, migrants and refugees, etc.) are significant part of
content. Summaries of identified topics (highly probable words) were assigned by experts
after qualitative analysis.</p>
      <p>Interpretability of topics built by topic modeling is an important issue for researchers
applying this technique. Our investigation showed that higher semantic coherence
indicates topics that have more consistent words (more interpretable) while exclusivity
measures how exclusive the words are to the topic relative to other topics (e.g. low
values mean topics that are vague and share a lot of words with other topics while high
values indicate words that are very unique/exclusive to the topic) (see Fig. 4).</p>
      <p>Fig. 4. Topic interpretability: the exclusivity and semantic coherence (X axis represents
semantic coherence, Y axis – exclusivity).</p>
      <p>After examination of the whole set, we focused on the distribution of topics across
different media channels. The stm is a general framework for topic modeling with
document-level covariate information. The covariates can improve inference and
qualitative interpretability and are allowed to affect topical prevalence, topical content or both.
The software package implements the estimation algorithms for the model and also
includes tools for every stage of a standard workflow from reading in and processing
raw text through making publication quality figures. Topical prevalence refers to how
much of a document is associated with a topic and topical content refers to the words
used within a topic. Expected difference in topic probability be media type (with 95 %
confidence intervals) is shown below (see Fig. 5 and Fig. 6).</p>
      <p>Following examining the distribution of all topics, we focused our research on key
sensitive topics. We find that the model captures important events and differences
between different media channel’ depictions of these events (see Annex 1).</p>
      <p>Topic correlation network creation results are depicted in Fig. 7. The way these
algorithms work is by assuming that each document is composed of a mixture of topics,
and then trying to find out how strong a presence each topic has in a given document.
This is done by grouping together the documents based on the words they contain, and
noticing correlations between them. A topic model captures this intuition in a
mathematical framework, which allows examining a set of documents and discovering, based
on the statistics of the words in each, what the topics might be and what each
document's balance of topics is.</p>
      <p>Topic models have become a standard tool within quantitative text analysis for many
different reasons. Topic models can be much more useful than simple word frequency
or dictionary based approaches depending upon the use case. Topic models tend to
produce the best results when applied to texts that are not too short (e.g. tweets), and those
that have a consistent structure.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Plans</title>
      <p>We discussed an ongoing research of: (1) text analytics and Artificial Intelligence tools
to identify trending topics in different media sources; (2) exploratory visual analytics
tools to identify prevalence of topics in different sources &amp; their dynamics. We
demonstrated the applicability of such approach to mainstream Lithuanian news portal
(delfi.lt) and two alternative/unconventional media channels – sarmatas.lt and
netiesa.lt. Early stage analysis shows considerable difference of prevalent topics in
different media channels, which allows identifying targets of the channel.</p>
      <p>We plan to extend research to wider set of media sources, change of topics in time
(more detailed) and relations between topics and media channels (more detailed).</p>
      <p>Annex 1
Explanation:
• Blue – mainstream media portal;
• Red – unconventional media portal;
• Line -- expected probabilities;
• Dash line – sample median.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Heidenreich</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lind</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eberl</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boomgaarden</surname>
            ,
            <given-names>H. G.</given-names>
          </string-name>
          :
          <article-title>Media Framing Dynamics of the 'European Refugee Crisis'. A Comparative Topic Modelling Approach</article-title>
          ,
          <source>Journal of Refugee Studies</source>
          ,
          <volume>32</volume>
          (
          <issue>1</issue>
          ),
          <fpage>i172</fpage>
          -
          <lpage>i182</lpage>
          (
          <year>2019</year>
          ), https://doi.org/10.1093/jrs/fez025, last accessed
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Güven</surname>
            ,
            <given-names>Z. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diri</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Çakaloğlu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Comparison of Topic Modeling Methods for Type Detection of Turkish News</article-title>
          .
          <source>In: 4th International Conference on Computer Science and Engineering (UBMK)</source>
          , pp.
          <fpage>150</fpage>
          -
          <lpage>154</lpage>
          , Samsun, Turkey (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kherwa</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Bansa,l P.
          <article-title>: Topic Modeling: A Comprehensive Review</article-title>
          ,
          <string-name>
            <surname>SIS</surname>
          </string-name>
          , EAI (
          <year>2019</year>
          ), doi: 10.4108/eai.13-
          <fpage>7</fpage>
          -
          <year>2018</year>
          .159623, last accessed
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Moodley</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marivate</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Topic Modelling of News Articles for Two Consecutive Elections in South Africa</article-title>
          .
          <source>In: 6th International Conference on Soft Computing &amp; Machine Intelligence (ISCMI)</source>
          , pp.
          <fpage>131</fpage>
          -
          <lpage>136</lpage>
          , Johannesburg, South Africa (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Brookes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McEnery</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The utility of topic modelling for discourse studies: A critical evaluation'</article-title>
          .
          <source>Discourse Studies</source>
          <volume>21</volume>
          (
          <issue>1</issue>
          ),
          <fpage>3</fpage>
          -
          <lpage>21</lpage>
          (
          <year>2019</year>
          ), doi: 10.1177/1461445618814032, last accessed
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Weiying</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pham</surname>
            ,
            <given-names>D.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hai</surname>
            ,
            <given-names>N.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ong</surname>
            ,
            <given-names>H. H.</given-names>
          </string-name>
          :
          <article-title>Topic Modelling for Malay News Aggregator</article-title>
          .
          <source>In: Fourth International Conference on Advances in Computing, Communication &amp; Automation (ICACCA)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ,
          <string-name>
            <given-names>Subang</given-names>
            <surname>Jaya</surname>
          </string-name>
          ,
          <string-name>
            <surname>Malaysia</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Missier</surname>
            <given-names>P.</given-names>
          </string-name>
          et al.:
          <article-title>Tracking Dengue Epidemics Using Twitter Content Classification and Topic Modelling</article-title>
          . In: Casteleyn S.,
          <string-name>
            <surname>Dolog</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pautasso</surname>
            <given-names>C</given-names>
          </string-name>
          . (eds) Current Trends in Web Engineering,
          <source>ICWE 2016, Lecture Notes in Computer Science</source>
          , vol
          <volume>9881</volume>
          , Springer, Cham (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Habibabadi</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haghighi</surname>
          </string-name>
          , P. D.:
          <article-title>Topic Modelling for Identification of Vaccine Reactions in Twitter</article-title>
          .
          <source>In: Proceedings of the Australasian Computer Science Week Multiconference (ACSW</source>
          <year>2019</year>
          ),
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA, Article
          <volume>31</volume>
          ,
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          (
          <year>2019</year>
          ), https://doi.org/10.1145/3290688.3290735, last accessed
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Arslina</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liebenlito</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Sequential Topic Modelling: A Case Study on Indonesian LGBT Conversation on Twitter</article-title>
          .
          <source>In: Prime: Indonesian Journal of Pure and Applied Mathematics</source>
          ,
          <volume>1</volume>
          .
          <fpage>10</fpage>
          .15408/inprime.v1i1.
          <volume>12726</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Sokolova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramisch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sazonova</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Black</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orwa</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ochieng</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sambuli</surname>
          </string-name>
          , N.:
          <article-title>Topic Modelling and Event Identification from Twitter Textual Data (</article-title>
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Rodina</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dligach</surname>
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Dictator's Instagram: personal and political narratives in a Chechen leader's social network</article-title>
          .
          <source>Caucasus Survey</source>
          ,
          <volume>7</volume>
          (
          <issue>2</issue>
          ),
          <fpage>95</fpage>
          -
          <lpage>109</lpage>
          (
          <year>2019</year>
          ), doi: 10.1080/23761199.
          <year>2019</year>
          .
          <volume>1567145</volume>
          , last accessed
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Bielinskienė</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boizou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bumbulienė</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalevskaitė</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krilavičius</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandravickaitė</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rimkutė</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vilkaitė-Lozdienė</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          : DELFI.lt corpus, Vilnius, Lithuania (
          <year>2019</year>
          ), https://www.clarin.vdu.lt/xmlui/handle/20.500.11821/30, last accessed
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Silge</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Text Mining with R- A Tidy Approach. O'Reilly Media</surname>
          </string-name>
          , Sebastopol, California, USA (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Roberts</surname>
            <given-names>M. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stewart</surname>
            <given-names>B. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tingley</surname>
            <given-names>D.</given-names>
          </string-name>
          : Stm:
          <article-title>An R Package for Structural Topic Models</article-title>
          .
          <source>Journal of Statistical Software</source>
          <volume>91</volume>
          (
          <issue>2</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          (
          <year>2019</year>
          ),
          <source>doi: 10.18637/jss.v091.i0, last accessed</source>
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>R</given-names>
            <surname>Core Team: R:</surname>
          </string-name>
          <article-title>A language and environment for statistical computing</article-title>
          .
          <source>R Foundation for Statistical Computing</source>
          , Vienna, Austria (
          <year>2014</year>
          ), http://www.R-project.org/,
          <source>last accessed</source>
          <year>2020</year>
          /03/15.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lafferty</surname>
            ,
            <given-names>J. D.</given-names>
          </string-name>
          :
          <article-title>Topic models</article-title>
          . In Text mining, pp.
          <fpage>101</fpage>
          -
          <lpage>124</lpage>
          , Chapman and Hall/CRC, Boca
          <string-name>
            <surname>Raton</surname>
          </string-name>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>