<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Topic-Based Approach to Multiple Corpus Comparison</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jinghui Lu</string-name>
          <email>Jinghui.Lu@ucdconnect.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maeve Henchion</string-name>
          <email>Maeve.Henchion@teagasc.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brian Mac Namee</string-name>
          <email>Brian.MacNamee@ucd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight Centre for Data Analytics, University College Dublin, Ireland Tegasc the Agriculture and Food Development Authority</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Corpus comparison techniques are often used to compare different types of online media, for example social media posts and news articles. Most corpus comparison algorithms operate at a word-level and results are shown as lists of individual discriminating words which makes identifying larger underlying differences between corpora challenging. Most corpus comparison techniques also work on pairs of corpora and do need easily extend to multiple corpora. To counter these issues, we introduce Multi-corpus Topic-based Corpus Comparison (MTCC) a corpus comparison approach that works at a topic level and that can compare multiple corpora at once. Experiments on multiple real-world datasets are carried demonstrate the effectiveness of MTCC and compare the usefulness of different statistical discrimination metrics - the 2 and Jensen-Shannon Divergence metrics are shown to work well. Finally we demonstrate the usefulness of reporting corpus comparison results via topics rather than individual words. Overall we show that the topic-level MTCC approach can capture the difference between multiple corpora, and show the results in a more meaningful and interpretable way than approaches that operate at a word-level.</p>
      </abstract>
      <kwd-group>
        <kwd>Corpus Comparison</kwd>
        <kwd>Topic Modelling</kwd>
        <kwd>Jensen-shannon Divergence</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Many corpus comparison techniques are proposed in the literature to reveal the
divergence between corpora [
        <xref ref-type="bibr" rid="ref16 ref8">8, 16</xref>
        ], especially corpora of web-based content such as online
news, social media posts, and blog posts [
        <xref ref-type="bibr" rid="ref10 ref5">5, 10</xref>
        ]. Although these approaches have been
shown to be effective, almost all of them are limited by comparing corpus at a word
level. Consequently, the results are communicated to a user as a list of unrelated words
that are divergent across two corpora, which makes identifying larger underlying
differences between the corpora challenging. Additionally, these studies focus on comparing
pairs of corpora instead of multiple corpora.
      </p>
      <p>
        In recent years, topic modelling [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] has become a widely used method for
revealing thematic information, which is described by a series of high related words called
topic descriptors, in a collection of documents. We assume the combination of topic
modelling techniques and corpus comparison methods has the potential to eliminate the
problem described above that arise with word-based corpus comparison approaches. In
other words, a approach uses topic modelling as a basis for corpus comparison to
offer users the divergence that can be explained at a topic-level rather than an individual
word-level.
      </p>
      <p>This paper describes the Multi-corpus Topic-based Corpus Comparison (MTCC)
approach to corpus comparison, that leverages topic modelling and statistical
discrimination metrics to conduct a topic-based comparison. We describe the approach and
demonstrate it on 8 real-world datasets. In a series of experiments we demonstrate that
the topics extracted by our models contain divergence information, as well as comparing
the effectiveness of different statistical discrimination metrics applied in the algorithm.
We also compare the output of MTCC with a word-based corpus comparison method
to show that the results output by MTCC are more meaningful and more interpretable
than those produced by the word-based method. The contributions of this paper area:
– Multi-corpus topic-based Corpus Comparison, a new corpus comparison technique
that leverages topic modelling and statistical discrimination metrics.
– An experiment to show that the topics extracted by the statistical discrimination
metrics are capturing divergence.
– An experiment that investigates the effectiveness of different statistical
discrimination metrics applied in MTCC.
– A demonstration of the usefulness of topic-based divergence explanations over
multiple real-world datasets.</p>
      <p>The rest of this paper is organized as follows: Section 2 presents related work;
Section 3 describes the MTCC approach; Section 4 describes experiments to measure the
divergent information contained in topics found and the effectiveness of different
statistical divergence metrics; Section 5 compares the output of the MTCC approach with
a word-based method; and, finally, Section 6 summarizes the work and suggests future
directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Corpus comparison approaches extract the distinct content from two corpora [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
Typically, researchers attempted to compute the contribution to the divergence of individual
words by applying many statistical discriminating metrics over word frequencies
calculated from different corpora, and those words which highly contribute to the difference
are selected to be presented as the divergence [
        <xref ref-type="bibr" rid="ref16 ref5 ref8">5, 8, 16</xref>
        ]. Therefore, it is very intuitive to
show the comparison results using words.
      </p>
      <p>
        There is a variety of statistical discrimination metrics in the literature. For instance,
Leech and Fallon [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] use 2 to measure the discriminative power of a word;
loglikelihood [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and relevance-frequency (RF) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] were utilized to extract divergent
information over corpora of different domains. Besides, Information Gain (IG), Gain
Ratio (GR) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], Kullback-Leibler divergence [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and Jensen-Shannon divergence (JSD)
[
        <xref ref-type="bibr" rid="ref13 ref5">5, 13</xref>
        ] have been widely employed in corpus comparison.
      </p>
      <p>
        Latent Dirichlet Allocation[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], as a widely used strategy for exploiting topic
information from texts, can automatically infer the distribution of membership to set of
topics in a large collection of documents. It has been shown to have a great ability to
find latent topics and cluster documents [
        <xref ref-type="bibr" rid="ref1 ref6">1, 6</xref>
        ].
      </p>
      <p>
        As far as we know, Zhao et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] were the first to use topic modelling, specifically
LDA, in conjunction with statistical discrimination to find topics specific to a corpus in
a pair of corpora being compared, and to use these to explain the differences between
the corpora. Zhao et al. first performed independent topic modelling on the corpora
being compared, and then applied Jensen-Shannon divergence (JSD) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] over the topic
descriptors to measure the pair-wise similarities between the sets of topics found in
the two corpora. If the similarity of the nearest match to a topic in one corpus to a
topic in the other corpus was below a specific threshold then that topic was said to be
discriminatory. The set of discriminatory topics was then used to explain the differences
between the two corpora. Zhao et al. demonstrated this approach by comparing corpora
from Twitter and the New York Times. Similarly, Murdock et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and Sievert et al.
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] used JSD applied to the distributions of word frequencies in topics to measure the
distance between topics, but this was not done in a corpus comparison scenario.
      </p>
      <p>Zhao et al’s method trains independent LDA model for each corpus which increases
the instability of the whole system due to the non-deterministic nature of LDA. Besides,
the output highly relies on the setting of thresholds, namely, the small thresholds tend
to result in too many similar topics as divergence, but the large thresholds will lose
some discriminating topics. The MTCC approach proposed in this paper differs from
the work of Zhao et al in two key ways. First, in MTCC a global topic model is trained
across a combined corpus that contains all documents from all corpora being compared.
Second, because a single topic model is built, intuitively, JSD can be applied directly
to topic membership vectors rather than applying JSD to word distributions between
topics. As compared to matching similar topics, MTCC skips the empirical setting of
thresholds to further automate the comparison process and, since one global topic model
is trained, the topic proportion of each corpus can be inferred on which multiple corpus
comparison can be based.</p>
      <p>corpus</p>
      <p>P1
corpus</p>
      <p>P2
corpus</p>
      <p>Pn
combined
corpus
...
train</p>
    </sec>
    <sec id="sec-3">
      <title>The Multi-corpus Topic-based Corpus Comparison (MTCC)</title>
    </sec>
    <sec id="sec-4">
      <title>Approach</title>
      <p>In this section, we will first provide a brief description of the use of LDA for topic
modelling, then describe the Multi-corpus Topic-based Corpus Comparison (MTCC)
approach in detail.
3.1</p>
      <sec id="sec-4-1">
        <title>Latent Dirichlet Allocation</title>
        <p>
          Latent Dirichlet Allocation (LDA) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is proposed to infer topic distribution in a
collection of documents. The model generates a document-topic matrix , and a
topicterm matrix (see Figure 1). Specifically, each row of the document-topic matrix is a
topic-based representation of a document where the ith entry determines the degree of
association between the ith topic and the document. Each row of the topic-term matrix
represents the word distribution of the corresponding topic. Usually, several most
common words in one topic will be chosen to be presented as the topic descriptors. We use
the LDA implemented in Python Gensim package 1.
        </p>
        <p>
          Properly setting the number of latent topics to be found plays a vital role in the
performance of LDA as well as other topic modelling algorithms. There are many
approaches for seeking the appropriate choice of the number of latent topics, k, in the
literature. O’Callaghan et al [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] proposed a topic coherence metric via word
embeddings reflecting the semantic relatedness of topic descriptors, which has been previously
used in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] to determine the number of topics to be found. We also adopt this method, in
the experiments, we use word embeddings built across our own test corpora using the
FastText algorithm [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2 Infer Topic Proportion for Corpus</title>
        <sec id="sec-4-2-1">
          <title>1 https://radimrehurek.com/gensim/models/ldamodel.html 2 https://radimrehurek.com/gensim/models/fasttext.html 3 https://www.nltk.org/api/nltk.html</title>
          <p>where [i; t] denotes the the proportion of tth topic in the ith documents in the combined
corpus and p is the number of documents from corpus Pi.
3.3</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Strategies for Comparing Multiple Corpora</title>
        <p>Since most of the corpus comparison approaches listed in Section 2 is focused on
comparing pairs of corpora, we adopt a one-versus-all strategy in order to conduct a multiple
corpus comparison.</p>
        <p>
          The algorithm is quite simple, a specific corpus Pi and a mixture corpus which
is a concatenation of all corpora except corpus Pi are considered two corpora being
compared. Hence, the proportion of any topic for corpus Pi and the mixture corpus can
be given by Equation 1, on which the statistical metrics are applied. Then we can derive
the divergence score of tth topic in terms of corpus Pi following [
          <xref ref-type="bibr" rid="ref11 ref17 ref5 ref9">5, 9, 11, 17</xref>
          ], which is
denoted by dt(Pi) in this paper.
        </p>
        <p>However, in order to rank the discrimination power of each topic, t, across the full
set of corpora to be compared, a single divergence score per topic is required
necessitating an aggregation strategy. We define three aggregation strategy which is given as
follows:
– sum: divt(P1 jj P2 jj ::: jj Pn) = Pn</p>
        <p>i=1 dt(Pi)
– weighted sum: divt(P1 jj P2 jj ::: jj Pn) = Pn
i=1 (Pi)dt(Pi)
– maximum: divt(P1 jj P2 jj ::: jj Pn) = maxin=1dt(Pi)
where n is the number of corpora for comparison and (Pi) is the proportion of corpus
Pi in the combined corpus, divt(P1 jj P2 jj ::: jj Pn) denotes the global divergence score
of topic t.</p>
        <p>
          There is a modification to the Jensen-Shannon divergence metric, extended JSD [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]
that can be used directly across multiple corpora at a word-level without the need for
the one-versus-all approach. In extended JSD divergence score for the tth word over
multiple corpora P1; P2; :::; Pn is defined as:
divJS;t(P1 jj P2 jj ::: jj Pn) =
n
mt log mt + 1 X pit log pit
n
i=1
(2)
where pit is the probability of seeing word t in corpus Pi, and mt is the probability
of seeing word t in M . Here, M is a mixed distribution of n corpora where M =
1 Pn
n i=1 Pi. In extended JSD, the global divergence score of the tth topic could be
derived from Equation 2 by replacing pit with memt(Pi) calculated by Equation 1.
        </p>
        <p>Subsequently, the topics can be ranked by their global divergence score in
descending order and the top n topics along with their topic descriptors can be selected to
represent the difference between multiple corpora.</p>
        <p>A github repository containing the code to implement the MTCC approach and all
experiments described in the following section is publicly available.4
4 https://github.com/GeorgeLuImmortal/topic-based corpus comparison</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Comparing Topic-based Discrimination Metrics</title>
      <p>
        In this set of experiments we compare the ability of different statistical discrimination
metrics to identify discriminative topics across corpora. Corpus comparison is an
unsupervised procedure which makes this evaluation somewhat challenging. To overcome
these challenges, following [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], we reframe the corpus comparison evaluation as a
document classification task. This section describes the experimental approach and the
datasets used, and discusses the results of these experiments.
4.1
      </p>
      <sec id="sec-5-1">
        <title>Datasets</title>
        <p>
          Our experiments are carried out on 8 real-world datasets described in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]: 6 news article
datasets bbc, bbc-sport, guardian-2013, irishtimes-2013, nytimes-1999, nytimes-2003
and 2 Wikipedia datasets wikipedia-high and wikipedia-low. Each dataset has different
sections, for example, bbc includes sections: business, politics, entertainment, sports,
tech. Table 1 shows the total number of documents, the size of vocabulary, the total
number of terms, and the number of sections and the value of k in each dataset. We
divide each dataset into corpora following these sections.
Following the approach used in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] we base our evaluation on the assumption that
discriminative topics should be useful features for classifying documents as belonging to
different corpora. Thus, we reframe the evaluation of corpus comparison to a document
classification task.
        </p>
        <p>The procedure for estimating discriminativeness of a set of topics found by MTCC
is as follows:
1. extract the top n most informative topics (MTCC is run and the n most
discriminative topics are extracted)
2. construct vector representations of documents based on the selected n topics.
3. build a classification model using the vector representation and measure its
performance.</p>
        <p>Dataset
bbc
bbc-sport
guardian-2013
irishtimes-2013
nytimes-1999
nytimes-2003
wikipedia-high
wikipedia-low</p>
        <p>The performance of the second 10 cross-validation is used to assess the
discriminativeness of the n topics.</p>
        <p>
          We repeat this procedure for values of n 2 [ 1k0 ; k], where k is the number of topics
used within MTCC and n increases by 1k0 every time. In other words, given a statistical
discrimination metric and a dataset, we will run the above procedure 10 times. Also, we
should note here, we rank the topics according to their divergence score from MTCC
model in descending order which means the first n topics are the most distinctive topics
whereas the remaining topics are not so informative. We also use a baseline, in which
n topics are selected randomly.
We compare the performance of MTCC models utilizing the different statistical
discrimination metrics in Section 2, i.e. IG, GR, 2, RF, JSD and extended JSD
(extJSD) across 8 real-world datasets. For each dataset, we treat a section as an
individual corpus—for instance, the bbc dataset has 5 corpora. Therefore, this is a multiple
corpora comparison task. After extensive preliminary experiments, we choose to use
Linear-SVM [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] classifier when measuring performance.5 Also, the maximum
aggregation strategy (see Section 3.3) has been adopted as it was shown to perform better
than the other strategies in preliminary experiments. For RF, we set the threshold to
0.01. The baseline method classifiers are trained with the document representations in
terms of n topics chosen arbitrarily. To reduce the effect of randomness, for a given n,
we run baseline methods 10 times with different random seeds and report the averaged
micro-averaged f1 score.
        </p>
        <p>We set the random seed for the LDA model to 1984; = \auto" which
indicates the automatically tuning the hyperparameter ;6 and the random seed for twice
cross-validation to 2018 and 0 respectively to make sure the results are consistent and
independent of random initial states. The best choice for k found for each dataset using
the approach described in Section 3.1 is also reported in Table 1.
We present the document classification results of MTCC models with respect to
different statistical discrimination metrics in different datasets in Figure 2. The x-axis denotes
n, the percentage of total topics used for training the classifier. The y-axis represents
the performance which is the micro-averaged f1 score in this case.</p>
        <p>We can observe that, almost in all situations, JSD, extended JSD or 2 achieves
high micro-averaged f1 scores that outperform the baseline (random selection) by a
5 http://scikit-learn.org/stable/modules/generated/sklearn.svm.LinearSVC.html
6 https://rare-technologies.com/python-lda-in-gensim-christmas-edition/
large margin. This demonstrates that the discrimination metrics can identify the topics
that contain divergence information. We can also see from Figure 2 that the best results
are achieved by JSD, extended JSD or 2 across 8 datasets with very low values for
n which denotes the percentage of the total topics used for training (usually &lt;= 30).
Moreover at these low numbers of topics the performance from the other metrics and
from the random baseline is very low. For example, in the guardian-2013 (Figure 2 (c))
and wikipedia-high (Figure 2 (g)) datasets, classifiers nearly reach the best performance
(above 0.8) using only the top 10% of topics, meanwhile the performance of baseline is
only a little higher than 0.4. This implies that the top 10% topics selected by JSD,
extended JSD and 2 carry almost all of the divergence information in these two datasets.
Similarly, as we can see in bbcsport (Figure 2 (b)), the classifiers based on topics
selected using JSD, extended JSD and 2 achieve the best results at n = 20 and these
results even surpass the result where all topics are used (n = 100). This indicates that
the top 20% of topics are very good at distinguishing between corpora while the
remaining topics not only do not contain divergence information but even produce noise
for classification.</p>
        <p>It is also interesting to note that JSD, extended JSD and 2 result in almost the same
performance across all datasets. These three metrics are reasonably similar to each other
so this is not too surprising. It is interesting to note, however, that although all three will
usually select the same topics at the very highest ranks, they do select different topics
further down the ordering.</p>
        <p>To conclude, JSD, extended JSD and 2 can effectively select the discriminative
topics and outperform the random baseline and other metrics by a large margin.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Demonstrating Topic-based Corpus Comparison</title>
      <p>Since we are interested in whether the results of corpus comparison at topic-level are
more meaningful and interpretable, in this section, we compare the output of a corpus
comparison based on the MTCC topic-level approach to a word-level approach over 4
real-world datasets. We describe the datasets used and an analysis of the results
produced by each approach.
5.1</p>
      <sec id="sec-6-1">
        <title>Setup</title>
        <p>
          In this demonstration we perform a corpus comparison that compares the bbc dataset,
guardian-2013 dataset, irishtimes-2013 and nytimes-2003 dataset described in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Each
dataset is treated as one corpus which is different from the settings in Section 4 and
extended JSD is selected as the discrimination metric since its effectiveness shown in
Section 4. We therefore conduct a comparison over 4 corpora at one time. These four
datasets are all news articles dataset. Also there is both overlap and difference in the
sections present in those four datasets (summarised in Table 2). This makes these four
corpora an interesting test case for corpus comparison approaches as there are differences
that we can expect a corpus comparison approach to discover—for example sections in
one corpus that are not in the others (e.g. music section in guardian-2013).
        </p>
        <p>
          Extended JSD, which is commonly used in extracting keywords from different
corpora and has proven its effectiveness in [
          <xref ref-type="bibr" rid="ref13 ref5">5, 13</xref>
          ], is adopted as the word-level corpus
comparison method. The contributions to divergence of individual words across all
corpora are computed by Equation 2.
        </p>
        <p>We set the random seed for the LDA model to 1984; = \auto" following the
previous experiments; the best choice k are tested in preliminary experiments varies
from 100 to 300 in steps of 10. After pre-experiments, 300 is the optimal number for k
according to the topic coherence measure described in Section 3.1.</p>
        <p>education
273 health
movies</p>
        <p>942</p>
        <p>The topics extracted using MTCC can be compared to the most discriminative words
for each dataset found by the word-based JSD approach which are shown in Table 4.</p>
        <p>The first thing that is apparent from comparing Tables 3 and 4 is that the
topiclevel approach has the advantage of presenting the user with coherent sets of terms in
the topic descriptors rather than a long list of disconnected words. This makes it much
easier for the user to understand the differences between different corpora.</p>
        <p>Looking more deeply at Table 3 we can see that the MTCC approach has
successfully identified the divergence information of each dataset, which is evidenced by the
presence of distinct topics. For example, the topic descriptors of topic 109 suggest that
this is a topic relating to articles describing court proceedings. We can see from this that
MTCC has identified the fact that the irishtimes-2013 corpus has a crime-law section
not present in other corpora. Similarly, the presence of other discriminative topics (e.g.
topic 289, topic 31, topic 42 etc.) has demonstrated the effectiveness of the algorithm
in detecting the difference between multiple corpora.</p>
        <p>When we look at Table 4, we can find that though word-level corpus comparison
properly define the divergence information in some datasets such as guardian-2013 and
nytimes-2013. This approach fails to depict the distinctive content in bbc corpus where
a set of unrelated words are presented.
6</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>In this paper, we introduce the Multi-Corpus Topic-based Corpus Comparison (TMCC)
approach to discover distinctive topics across multiple corpora for corpus comparison
tasks. We compared the performance of different discrimination metrics on 8 real-world
datasets. The results showing that using JSD, extended JSD, or</p>
      <sec id="sec-7-1">
        <title>2 can extract topics</title>
        <p>that contain the most divergence information. We also demonstrated TMCC and
compared its output to a word-level approach. Overall we believe that the example presented
demonstrates the advantages of using topics for corpus comparison rather than using
word-level approaches.</p>
        <p>
          However, because we apply topic modelling on multiple corpora, we are likely to
take the risk of losing or blurring topics compared to applying topic modelling over a
single corpus (as done by Zhao et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]). Hence, a further exploration that investigates
whether or not we are losing or blurring topics together in the multi-corpus topic-based
corpus comparison is scheduled in future work.
        </p>
        <p>Acknowledgement. This research was kindly supported by a Teagasc Walsh
Fellowship award (2016053) and Science Foundation Ireland (12/RC/2289 P2).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Belford</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Mac</given-names>
            <surname>Namee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Greene</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Stability of topic modeling via matrix factorization</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>91</volume>
          ,
          <fpage>159</fpage>
          -
          <lpage>169</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of machine Learning research 3(Jan)</source>
          ,
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>arXiv preprint arXiv:1607.04606</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Degaetano-Ortlieb</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kermes</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khamis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teich</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>An information-theoretic approach to modeling diachronic change in scientific english. Selected papers from VariengFrom data to evidence (d2e) (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gallagher</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reagan</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Danforth</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dodds</surname>
            ,
            <given-names>P.S.:</given-names>
          </string-name>
          <article-title>Divergent discourse between protests and counter-protests:# blacklivesmatter and# alllivesmatter</article-title>
          .
          <source>PloS one 13(4)</source>
          ,
          <year>e0195644</year>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Greene</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Callaghan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>How many topics? stability analysis for topic models</article-title>
          .
          <source>In: Joint European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          . pp.
          <fpage>498</fpage>
          -
          <lpage>513</lpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kelleher</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Mac</given-names>
            <surname>Namee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>D'Arcy</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Fundamentals of machine learning for predictive data analytics: algorithms, worked examples, and case studies</article-title>
          . MIT Press (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Comparing word frequencies across corpora: Why chi-square doesn't work, and an improved lob-brown comparison</article-title>
          . In:
          <string-name>
            <surname>ALLC-ACH Conference</surname>
          </string-name>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Low</surname>
          </string-name>
          , H.B.:
          <article-title>Proposing a new term weighting scheme for text categorization</article-title>
          .
          <source>In: AAAI</source>
          . vol.
          <volume>6</volume>
          , pp.
          <fpage>763</fpage>
          -
          <lpage>768</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lawrence</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sides</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farrell</surname>
          </string-name>
          , H.:
          <article-title>Self-segregation or deliberation? blog readership, participation, and polarization in american politics</article-title>
          .
          <source>Perspectives on Politics</source>
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <fpage>141</fpage>
          -
          <lpage>157</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Leech</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fallon</surname>
          </string-name>
          , R.:
          <article-title>Computer corpora-what do they tell us about culture</article-title>
          .
          <source>ICAME journal 16</source>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Divergence measures based on the shannon entropy</article-title>
          .
          <source>IEEE Transactions on Information theory 37</source>
          (
          <issue>1</issue>
          ),
          <fpage>145</fpage>
          -
          <lpage>151</lpage>
          (
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henchion</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>MacNamee</surname>
          </string-name>
          , B.:
          <article-title>Extending jensen shannon divergence to compare multiple corpora</article-title>
          .
          <source>In: 25th Irish Conference on Artificial Intelligence and Cognitive Science</source>
          , Dublin, Ireland, 7
          <article-title>-8 December 2017</article-title>
          .
          <article-title>CEUR-WS. org (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Murdock</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Visualization techniques for topic model checking</article-title>
          .
          <source>In: Twenty-Ninth AAAI Conference on Artificial Intelligence</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>O</given-names>
            <surname>' Callaghan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Greene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Carthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>An analysis of the coherence of descriptors in topic modeling</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>42</volume>
          (
          <issue>13</issue>
          ),
          <fpage>5645</fpage>
          -
          <lpage>5657</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rayson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garside</surname>
          </string-name>
          , R.:
          <article-title>Comparing corpora using frequency profiling</article-title>
          .
          <source>In: Proceedings of the workshop on Comparing corpora-Volume</source>
          <volume>9</volume>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . Association for Computational Linguistics (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sajgalik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barla</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bielikova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Searching for discriminative words in multidimensional continuous feature space</article-title>
          .
          <source>Computer Speech &amp; Language</source>
          <volume>53</volume>
          ,
          <fpage>276</fpage>
          -
          <lpage>301</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Sievert</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shirley</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Ldavis: A method for visualizing and interpreting topics</article-title>
          .
          <source>In: Proceedings of the workshop on interactive language learning, visualization, and interfaces</source>
          . pp.
          <fpage>63</fpage>
          -
          <lpage>70</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>W.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>E.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Comparing twitter and traditional media using topic models</article-title>
          .
          <source>In: European conference on information retrieval</source>
          . pp.
          <fpage>338</fpage>
          -
          <lpage>349</lpage>
          . Springer (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>