<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Interactive Visualization for Topic Model Curation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Guoray Cai</string-name>
          <email>cai@ist.psu.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Feng Sun</string-name>
          <email>fzs122@psu.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yongzhong Sha</string-name>
          <email>shayzh@lzu.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lanzhou University</institution>
          ,
          <addr-line>Lanzhou, Gansu</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Penn State University</institution>
          ,
          <addr-line>University Park, PA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Understanding the content of a large text corpus can be assisted by topic modeling methods, but the discovered topics often do not make clear sense to human analysts. Interactive topic modeling addresses such problems by allowing a human to steer the topic model curation process (generate, interpret, diagnose, and refine). However, human have limited ability to work with the artifacts of computational topic models since they are difficult to interpret and harvest. This paper explores the nature of such challenges and provides a visual analytic solution in the context of supporting political scientists to understand the thematic content of online petition data. We use interactive topic modeling of the White House online petition data as a lens to bring up key points of discussions and to highlight the unsolved problems as well as potentials utilities of visual analytics methods.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Author Keywords
Topic models; Information visualization; visual analytics
INTRODUCTION
Topic modeling has been advanced as a solution to the
challenge of making sense of large corpora of textual data. With
the help of machines, valuable themes buried in a large
document collection can emerge and provide a better representation
of the documents. The most popular topic modeling
techniques, LDA (Latent Dirichlet Allocation) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and its variants,
such as supervised LDA [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and supervised anchor LDA [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
have been proven useful in many applications [
        <xref ref-type="bibr" rid="ref25 ref29">29, 25</xref>
        ],
including online petition analysis [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Topic modeling assists
qualitative and quantitative research over user-generated texts
coming from the blogs or social media. By studying the set
of topics learned from social media conversations over some
period of time, it may become possible to find out what users
are talking about, identify underlying topical trends, and
follow them through time. Topic similarities among documents
©2018. Copyright for the individual papers remains with the authors.
Copying permitted for private and academic purposes.
      </p>
      <p>ESIDA ’18, March 11, 2018, Tokyo, Japan
also help to identify the most relevant documents for a specific
topic. Ideally, an analyst may be able to draw conclusions
from word distributions for topics and use such insight to
conduct a more in-depth study on documents with high affinities
for specific topics.</p>
      <p>
        Despite such advances, topic models have not been widely
adopted by data analysts for practical use of understanding
large corpora [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Topics discovered by LDA and other
algorithms often have both “good” and “bad” topics judged by
users. Topics could be bad because (1) they often confuse
two or more themes into one topic; (2) they often pick up two
different topics that are (nearly) duplicates for human; and
(3) nonsense topics [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], (4) topics with too many generic
words (e.g., “people, like, mr”) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], (5) topics with disparate
or poorly connected words [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], (6) topics misaligned with
human interpretation [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], (7) irrelevant topics [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], (8)
missing associations between topics and documents [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and (9)
multiple similar topics [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The presence of poor-quality topics
has been cited as the primary obstacle to the acceptance of
statistical topic models outside of the machine learning
community [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The root of these problems lies in the fact that the
objective function that topic models optimize does not always
correlate well with human judgments of topic quality [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Due
to these problems, the use of topic models to analyze
domainspecific texts often requires manual validation of the latent
topics to ensure that they are meaningful [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Addressing the above issues to make topic models usable by
analysts who are not machine learning experts, a variety of
human-in-the-loop methods have been proposed to allow
analysts to manipulate and incrementally refine a topic model
of a target text corpus [
        <xref ref-type="bibr" rid="ref17 ref18 ref19 ref2">17, 18, 19, 2</xref>
        ]. These methods
typically involve the use of interactive visualization and direct
manipulation of topic models to diagnose poor topics and fix
them through operations such as adding or removing words in
topics, adjusting the weights of words within topics, splitting
generic topics, and merging similar topics [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. For example,
ITM [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] allows users to add, emphasize, and ignore words
within topics, while UTOPIAN [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] allows users to adjust the
weights of words within topics, merge and split topics, and
create new topics. Additionally, iVisClustering [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] lets users
manually create or remove topics, merge or split topics, and
reassign documents to another topic, with the help of visually
exploring topic-document associations in a scatter plot.
While these operations can be supported by direct
manipulation and algorithmic extensions, it is more challenging to
diagnose the quality concerns of machine-discovered topics,
and in assessing if a refinement strategy results in topic
improvement. This is where interactive visualization methods are
most helpful. Topic Browser [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] uses a tabular visualization
technique to assist assessing term orders within each topic,
and Termite [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] focuses on supporting effective evaluation of
term distributions associated with LDA topics through
visualizations. TopicNets [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] used a web-based interactive visual
interface to enables users to discover topics of increasing
granularity through an informed selection of relevant subsets of
documents.
      </p>
      <p>While these visualization tools help users to assess and refine
static topic models, they run short in supporting the whole
topic curation process. Topic model curation goes beyond
human validation of machine-generated topics to include the
whole human-directed process of discovering topics that are
useful specific to a domain of applications. For example,
public opinion researchers may be interested in discovering what
is the range of policy preferences expressed in blog-spheres.
Crisis managers may be interested in conversations in social
media that are especially informative to their decisions on how
to allocate resources and dispatch rescue teams. For such
applications, the use of topic models is not a one-shot process but
is a broader process of seeking, assessing, relating, and
structuring topics with the help of supervised and unsupervised
topic models. A typical topic curation process starts with a
vanilla topic model (purely unsupervised probabilistic model
such as LDA), and let users conduct a full diagnostics to
recognize good and bad topics. Good topics will be collected and
kept in a “bag”, while bad topics improved or removed. For
the set of bad topics, users may explore multiple ways to
adjust topic models (merging/splitting topics, adding/removing
words from a topic, modifying orders or weights of words in
a topic). Depending on the consequence of imposed
correlations and constraints, a new round of modeling and refinement
can be initiated to explore the topic space of the document
collection either in breadth or depth.</p>
      <p>
        Towards supporting topic curation, this paper focuses on
understanding the specific challenges of topic curation in the
context of analyzing online petition data. We gained insight
by actually practicing interactive topic modeling on the
petition data we collected from the White House online petition
website “We the People”. This data set is considered a unique
source for understanding citizens’ policy concerns and
preferences [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The insight gained from this practice is used
to inform the design of a visual analytic system that supports
topic model diagnostics, refinement, and evaluation. We
reflect the use of visual analytic methods to enable users to
interactively curate topic models.
      </p>
      <p>
        INTERACTIVE TOPIC MODELING OF PETITION DATA
Electronic petitioning (e-petitioning) is becoming a prevalent
form of political action for enabling direct democratic
engagement [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The data used for this study comes from the online
petitioning platform “We the People”, hosted by the White
House. It contains 5,177 petitions accumulated over the course
of six years (2011-2016). We further selected 4,095 petitions
that are in English. Each petition has four fields: (1) a petition
ID, (2) a title, (3) a description, and (4) category tags.
As topic models treat documents as “bag-of-words”, the first
step of preparation before model training is tokenization,
which splits each petition into a set of words. As words may
have various forms, lemmatization is then applied to transform
them into a common base form. Compared with the
stemming technique that shares a similar goal, lemmatization takes
advantage of vocabulary analysis and thus can produce the
dictionary form of words that users can interpret. Bigrams
are also used here for performance purpose [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. Finally,
stopwords are removed from the texts, as well as the overly
common terms that appear frequently (top 50), to avoid
possible discrimination. The resulting corpus contains 11,189
unique terms.
      </p>
      <p>
        System Design
Figure 1 shows the user interface of interacting with petition
documents and topic words. This system has two functional
areas. The lower part is a topic-word visualization that supports
direct manipulation of words-to-topics correlation.
The upper part is designed for exploring the topic quality from
the perspective of how the petitions (documents) are
clustered according to the space defined by the topics. The points
cloud map provides a visual overview of the petition space
where topically similar petitions are positioned adjacently. It
is generated using t-SNE (t-distributed stochastic neighbor
embedding) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to reduce the high-dimensional petition data
to a 2-D vector space that human can perceive easily. Due to
its nature of being nondeterministic, t-SNE usually transforms
a high-dimensional data point to a different 2-D vector.
However, the relationships between the data points will remain
almost the same. An example of visualized petitions is shown
in the Figure 1.
      </p>
      <p>Each petition is assigned to one cluster based on its most
salient topic and is color-coded correspondingly. Users can
apply filters and highlighters on topics to manipulate the
petition overview map. Highlighting enables users to review
petitions in context while filtering allows users to focus on the
petitions of interest. When hovering over a document point,
a pop-up window displays the title, body, and topics of the
document. In the meantime, the topic distribution (in terms
of weights) of the selected document is visualized as a bar
chart. By clicking a topic label, its topic-words distribution is
visualized as color-coded bars.</p>
      <p>
        At the back-end of the system, we choose Correlation
Explanation (CorEx) [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] as the topic modeling algorithm to perform
interactive topic curation. Built on the theory of Correlation
Explanation [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] in information science, CorEx strives to
represent the substrate information in a document collection that
maximizes the informativeness of the data. Due to its fast
training time and capability of supporting anchoring, CorEx can
be easily tailored to incorporate human imposed correlations
or constraints for semi-supervised topic modeling, making it
an ideal choice for supporting interactive topic modeling [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
Using CorEx, users can anchor multiple words to one topic,
anchor one word to multiple topics, or any other creative
combination of anchors in order to discover topics that do not
naturally emerge. By leveraging CorEx’s capability of topic
seeding through anchor words in our system, human analysts
can incorporate their knowledge and insights into the process
of refining topic models.
      </p>
      <p>TOPIC CURATION
Using our system for topic curation involves three phased of
activities, with a number of iterations.</p>
      <p>
        Topic Discovery
The first step is to use topic modeling algorithm with random
seeds to run an unsupervised discovery of topics. The user
must specify how many topics is to be produced, with the
understanding that different numbers of topics can be chosen
to analyze the petitions data on different levels of granularity
and it is likely to generate a different set of topics [
        <xref ref-type="bibr" rid="ref14 ref24">14, 24</xref>
        ].
After initial unsupervised topic modeling with CorEx, users
assess the topic model and conduct diagnostic analysis on
topics. In particular, users will inspect topics, both individually
and as a group, to evaluate their qualities by examining topic
words. Those topics that users recognize as good ones should
be kept. For those bad ones, users can file complaints and
come up with one or more strategies to address them.
Topic Refinement
Topic refinement is achieved through manipulating
topicsword representations at the bottom part of Figure 1. We
included an anchoring mechanism to be coupled with CorEx
models. It allows users to anchor one or more words to one
topic, anchor one word to multiple topics, and anchor one or
more words to some topics while not others. With this
anchoring mechanism, topic revision interactions are supported
by operations such as splitting a topic, merging by joining,
and merging by absorbing (following [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]). More complicated
operations can be achieved through a combination of above
basic operations. For example, investigating more fine-grained
topics can be accomplished by splitting topics iteratively.
Split a topic
If a topic is considered be “bad” based on the observation that
it confuses two or more meaningful topics into one topic, a
solution could be to split the topic into two or more topics. To
do so, the user can check the topic he/she intends to split and
then click the “split” button. Before applying the operation,
the user is provided with the option to configure the number
of resulting topics. Once confirming, the underlying model
training will re-run under the new constraint that only the
selected topic is decomposed while the others remain the same
in terms of word allocation. Updated results will be generated
and visualized.
      </p>
      <p>
        In the backend, splitting a topic into n topics involves training
a word2vec model to produce word embeddings [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The
resulting model is used to calculate the semantical similarity
between words. After that, a similarity matrix of the words
within this topic is produced, and spectral clustering is applied
to the matrix to categorize the words into n clusters. The
n clusters of words are encoded into the previous model as
anchor words and will produce n new topics to replace the
original one.
      </p>
      <p>Merge topics by joining
If several topics are judged to have something common in their
semantic meaning, they can be merged into one topic. This is
accomplished by selecting these topics and then clicking
“apply” button. The system automatically apply the constraint that
words assigned to the topics to be merged have to appear in
the resulting topic. underlying model will be updated.
Accordingly, the visualization will be re-rendered. In the backend,
the words that appeared in the two topics are now anchored
under the same one.</p>
      <p>Merge topics by absorption
If one or more words in a topic are considered intruders and
fit better to a different topic, the user can re-allocate topic
words through drag-and-drop operations. Specifically, a user
can select a word that is considered allocated incorrectly and
move it to a more related topic. After reallocation of words is
done, the petition view will update to reflect the modification.
In the back end, Merging topics by absorbing is basically
a reallocation process where selected words in one topic is
anchored to the other one and a new model is trained. The
rest of the topic-word assignments remain the same through
anchoring as well.</p>
      <p>Evaluating Topics Interactively
Evaluating the quality of the topics in the current model is
necessary for both the diagnoses of good/bad topics as well
as assessing the impact of topic revisions. Evaluating topic
quality is done by assessing two aspects: (1) are the words in
a topic coherent and contributing to some collective meaning?
(2) are the topics aligned with the information needs of the
intended application? As such, we designed the interface
in Figure 1 that visualizes topically represented petitions to
support the following functions for evaluating the quality of
topics:
Inspecting quality of every single topic. Users can evaluate
topics by looking at the coherence of the component words
and their relative weights (see the bars next to words) on a
topic. Topics are also color-coded in the visualization window.
Clicking on the legend of a topic results in all the petitions with
sufficient weights on that topic being highlighted (while other
petitions are dimmed). These functions allow users to explore
the patterns of how petitions of the same topic clustered. A
good topic tends to create a cluster of petitions that are less
mixed with petitions.</p>
      <p>
        Comparing topics. Users can evaluate one or more topics
together by observing semantic relations to spatially close or
remote topics, and by looking at the spatial relationships
(overlapping clusters, adjacent clusters, non-intersecting clusters)
between petitions of the two topics. Applying filters to leave
fewer topics on the figure helps reduce visual clutters.
TOPIC MODEL CURATION SCENARIO
We practiced topic curation process on the online petition
dataset to experience how well our system supports topic
diagnostics and refinement. Firstly, we run the CorEx topic
modeling and generated 20 topics. A fixed random seed was
used to make sure the same results can be reproduced. Table 1
shows 5 samples out of 20 produced from a topic model. The
initial result from the CorEx topic modeling reveals interesting
topic clusters from the data set. In the provided samples, topic
0 mainly talks about “disease”, topic 4 generally discusses
“economy”, topic 5 describes “election”, and topic 16
represents “law enforcement”. The bottom part of the table shows
the results after applying certain topic revision operations.
id
0
4
5
6
16
0’
4’
6.1
6.2
Moving Intruder Words
By examining the above table, we find that topic 4 contains a
word “health” that is clearly different from other words (see
Figure 2). We also find that some petitions related to health but
has nothing to do with “economy” are assigned to this topic
during the petition exploration phase. One example petition
is “place mental health as a required course in junior high and
middle schools”. In order to correct this topic assignment,
we performed topic refinement by moving the intruder word
“health” from topic 4 to topic 0. The re-generated topic words
are shown in Table 1 as topic 0’ and topic 4’.
In order to assess if such a strategy of refining topics has
led to a better outcome, we rendered the petition clusters in
relation to the new topic definition and the result is shown
in Figure 3. From this figure, we can clearly see how topic
groups are isolated and cut. Compared with Figure 1, outliers
are nicely scattered apart and small clusters of outliers
disappear. Such result suggests that the change of topic model
by moving “health” from topic 4 to topic 0 is a good move.
This claim is further confirmed by a calculated metric of topic
coherence based on word context vectors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This metric has
been demonstrated to have the highest correlation with the
interpretability of topics [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. The topic coherence of topic
4 is increased from 0.453 to 0.555 after removing the word
intruder, and the overall topic coherence is increased from
0.431 to 0.443.
      </p>
      <p>Split a Multi-theme Topic
Observations show that the distribution of petitions of topic 6
is scattered in the reduced-dimensional space: there are
several small clusters of petitions. By sampling some of them for
detailed inspection of petition contents, we found that some
semantically irrelevant petitions are placed adjacently in the
visualization, e.g., “Prevent the FCC from ruining the Internet”
and “Put a fee on carbon-based fuels and return revenue to
households”, the former is about Internet and information
technology, while the latter is related to energy. This finding can
also be validated by examining topic words of topic 6:
“internet”, “information”, and “technology” are clearly incoherent
(a) Petitions of topic 0 and topic 4
(b) Petitions of topic 0’ and topic 4’
with “energy”, “fuel”, and “safety”. Therefore, we believe
topic 6 is of low quality since it contains several sub-topics
and needs to be diluted.
To address the quality concerns of topic 6, we split topic 6 into
two topics (by clicking on Topic 6 and choose "Split" button).
(b) Petitions of topic 6 (6.1) and topic 7 (6.2)
The modified version of the topic model is shown in Table 1
as 6-1 and 6-2 and Figure 4 as 6 and 7. The figure shows
that the weights of the first several topic words are increased,
indicating that these words can better represent the topics. It is
also apparent from Figure 5 that the distributions of petitions
for topic 6 and topic 7 become more focused, indicating that
the petitions documents within same clusters are more
topically homogeneous. After the new topic model applied, the
above example petitions are allocated to the correct topics
respectively, resulting in an increase of overall coherence value
from 0.431 to 0.441. Specifically, the original topic 6 has an
individual coherence score of 0.341, while the scores of newly
produced topic 6 and topic 7 are 0.594 and 0.419 respectively.
Merge Semantically Similar Topics
If the number of topics is set to a large number, CorEX
algorithm will generate topics in finer granularity of topics. This
could create situations where words that contribute to a single
theme end up in separate topics. Under such circumstance,
a merging operation is necessary to make sure that petitions
of similar topics are grouped together. In order to
demonstrate this situation, we trained another topic model by setting
the number of topics as 50 (relatively large) and the topic
words are shown in Figure 6a. By looking at the topic words,
topic 1 and topic 7 both describe “healthcare” but appear to be
different topics.</p>
      <p>The topic words after merging these two topics are shown
in Figure 6b. Petitions of these two topics are now grouped
into one cluster as well. Subsequently, these petitions can
be processed and analyzed as a whole, e.g., summarized and
forwarded to the Department of Health and Human Services.
Merging topics is also useful when a small number of topics
is used. Referring to the before-mentioned topic model of 20
topics, we found that topic 5 contains words “investigation”
and “justice” that may be related to topic 16. Therefore, we
performed a merging by joining on these two topics and it
leads to a more general topic denoted as 5+16. Although the
coherence value of merging the two topics remains almost
the same, it is noteworthy that a new word “corruption” is
prioritized as it could serve as a bridge to connect two topics
represented as “election” and “law enforcement” (e.g., a
petition titled “Arrest and prosecute officials who tried to suppress
the vote in the 2012 election”), showing that merging topics
has the potential of revealing latent relationship among them.
Topics that are difficult to interpret may still exist even after
several iterations of topic refinements. On the other hand,
some petitions are complicated in that they have multiple
equally important aspects and even people have difficulty in
identifying the most representative one. For those documents
that are related to "bad" topics and can not be fixed at this
round of analysis, the system can collect them into a subset of
data to be fed into the next round of analysis.</p>
      <p>DISCUSSION
Our work on analyzing the topic structures of online petitions
is still a work-in-progress, but we have gained several lessons
about interacting with topic modeling tools. First, users have
to deal with tremendous uncertainties when deciding what
is the proper strategy in tuning the topic model. Visualizing
the impact of multiple strategies and providing interaction
capabilities to assess the quality of topics and compare the
document clusters before and after the model tuning will be
critically important.</p>
      <p>Another finding from this exercise is that there is a need to
construct topic hierarchy from unsupervised topic models in
order to be aligned with the way political scientists perceive
the world of petition data. However, the topics discovered by
CorEx algorithm have a flat structure, and they tend to be
biased towards those topic branches that have more detailed data.
We will continue to explore our visual analytic approach for
incremental refinement of topic structures and demonstrated
how such an approach can be used to uncover topic
hierarchy of petitions that best reflects the human conception of
the domain. Further work is required to evaluate the usability
and effectiveness of this method. While we used dimension
reduction based visualization, other petition explorations and
analysis approaches should be investigated as well.
ACKNOWLEDGEMENT
The authors would like to acknowledge funding support from
National Science Foundation under award # IIS-1211059, and
from a grant funded by the Chinese Natural Science
Foundation under award 71373108.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Nikolaos</given-names>
            <surname>Aletras</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Stevenson</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Evaluating Topic Coherence Using Distributional Semantics</article-title>
          .
          <source>In Proceedings of the 10th International Conference on Computational Semantics (IWCS</source>
          <year>2013</year>
          ).
          <fpage>13</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>David</given-names>
            <surname>Andrzejewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Xiaojin</given-names>
            <surname>Zhu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Craven</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Incorporating Domain Knowledge into Topic Modeling via Dirichlet Forest Priors</article-title>
          .
          <source>In Proceedings of the 26th Annual International Conference on Machine Learning (ICML '09)</source>
          . ACM, New York, NY, USA,
          <fpage>25</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Sanjeev</given-names>
            <surname>Arora</surname>
          </string-name>
          , Rong Ge, Yonatan Halpern, David Mimno,
          <string-name>
            <given-names>Ankur</given-names>
            <surname>Moitra</surname>
          </string-name>
          , David Sontag,
          <string-name>
            <given-names>Yichen</given-names>
            <surname>Wu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A practical algorithm for topic modeling with provable guarantees</article-title>
          .
          <source>In International Conference on Machine Learning</source>
          .
          <fpage>280</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>David</surname>
            <given-names>M Blei</given-names>
          </string-name>
          , Andrew Y Ng, and
          <string-name>
            <given-names>Michael I</given-names>
            <surname>Jordan</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of machine Learning research 3</source>
          ,
          <string-name>
            <surname>Jan</surname>
          </string-name>
          (
          <year>2003</year>
          ),
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Jordan</given-names>
            <surname>Boyd-Graber</surname>
          </string-name>
          , David Mimno,
          <string-name>
            <given-names>and David</given-names>
            <surname>Newman</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Care and feeding of topic models: Problems, diagnostics, and improvements</article-title>
          .
          <source>In Handbook of Mixed Membership Models and Its Applications</source>
          . Chapman &amp; Hall, Chapter
          <volume>12</volume>
          ,
          <fpage>225</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Ajb</given-names>
            <surname>Chaney</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dm</given-names>
            <surname>Blei</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Visualizing Topic Models.</article-title>
          .
          <source>In Proceedings of the Sixth International AAAI Conference on Weblogs and Social Media</source>
          .
          <fpage>419</fpage>
          -
          <lpage>422</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Boyd-Graber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gerrish</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Reading tea leaves: How humans interpret topic models</article-title>
          .
          <source>In Proceedings of Advances in Neural Information Processing Systems</source>
          .
          <volume>288</volume>
          -
          <fpage>296</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Jaegul</given-names>
            <surname>Choo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Changhyun</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Chandan K Reddy,</article-title>
          and
          <string-name>
            <given-names>Haesun</given-names>
            <surname>Park</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>UTOPIAN: User-Driven Topic Modeling Based on Interactive Nonnegative Matrix Factorization</article-title>
          .
          <source>IEE Transactions of Visualization and Computer Graphics</source>
          <volume>19</volume>
          ,
          <issue>12</issue>
          (
          <year>2013</year>
          ),
          <fpage>1992</fpage>
          -
          <lpage>2001</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Jason</given-names>
            <surname>Chuang</surname>
          </string-name>
          , Sonal Gupta,
          <string-name>
            <surname>Christopher D Manning</surname>
            , and
            <given-names>Jeffrey</given-names>
          </string-name>
          <string-name>
            <surname>Heer</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Topic Model Diagnostics: Assessing Domain Relevance via Topical Alignment</article-title>
          .
          <source>In Proceedings of the 30th International Conference on Machine Learning</source>
          .
          <fpage>612</fpage>
          -
          <lpage>620</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jason</surname>
            <given-names>Chuang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Heer</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Termite : Visualization Techniques for Assessing Textual Topic Models</article-title>
          .
          <source>In Proceedings of the International Working Conference on Advanced Visual Interfaces - AVI '12</source>
          .
          <fpage>74</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Hal</given-names>
            <surname>Daumé</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Markov random topic fields</article-title>
          .
          <source>Proceedings of the ACL-IJCNLP 2009 Conference Short Papers August</source>
          (
          <year>2009</year>
          ),
          <fpage>293</fpage>
          -
          <lpage>296</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Ryan</surname>
          </string-name>
          J Gallagher, Kyle Reing, David Kale, and Greg Ver Steeg.
          <year>2016</year>
          .
          <article-title>Anchored Correlation Explanation: Topic Modeling with Minimal Domain Knowledge</article-title>
          .
          <source>arXiv preprint arXiv:1611.10277</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Brynjar</surname>
            <given-names>Gretarsson</given-names>
          </string-name>
          ,
          <string-name>
            <surname>John O'Donovan</surname>
            , Svetlin Bostandjiev, Tobias Höllerer, Arthur Asuncion, David Newman,
            <given-names>and Padhraic</given-names>
          </string-name>
          <string-name>
            <surname>Smyth</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>TopicNets: Visual Analysis of Large Text Corpora with Topic Modeling</article-title>
          .
          <source>ACM Trans. Intell. Syst. Technol. 3</source>
          ,
          <issue>2</issue>
          ,
          <string-name>
            <surname>Article 23</surname>
          </string-name>
          (
          <issue>Feb</issue>
          .
          <year>2012</year>
          ),
          <volume>26</volume>
          pages.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Loni</surname>
            <given-names>Hagen</given-names>
          </string-name>
          , Ozlem Uzuner, Christopher Kotfila,
          <string-name>
            <surname>Teresa M. Harrison</surname>
            , and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Lamanna</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Understanding Citizens' Direct Policy Suggestions to the Federal Government: A Natural Language Processing and Topic Modeling Approach</article-title>
          .
          <source>In 2015 48th Hawaii International Conference on System Sciences</source>
          , Vol.
          <fpage>2015</fpage>
          -March. IEEE,
          <fpage>2134</fpage>
          -
          <lpage>2143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Scott A Hale,
          <string-name>
            <given-names>Helen</given-names>
            <surname>Margetts</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Taha</given-names>
            <surname>Yasseri</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Petition growth and success rates on the UK No. 10 Downing Street website</article-title>
          .
          <source>In Proceedings of the 5th annual ACM web science conference. ACM</source>
          ,
          <volume>132</volume>
          -
          <fpage>138</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. David Hall,
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Studying the history of ideas using topic models</article-title>
          .
          <source>In EMNLP '08 Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          .
          <fpage>363</fpage>
          -
          <lpage>371</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Enamul</given-names>
            <surname>Hoque</surname>
          </string-name>
          and
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Carenini</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Interactive Topic Modeling for Exploring Asynchronous Online Conversations</article-title>
          .
          <source>ACM Transactions on Interactive Intelligent Systems</source>
          <volume>6</volume>
          ,
          <issue>1</issue>
          (feb
          <year>2016</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Yuening</surname>
            <given-names>Hu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            Boyd-Graber,
            <given-names>Brianna</given-names>
          </string-name>
          <string-name>
            <surname>Satinoff</surname>
            , and
            <given-names>Alison</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Interactive topic modeling</article-title>
          .
          <source>Machine Learning</source>
          <volume>95</volume>
          ,
          <issue>3</issue>
          (
          <year>2014</year>
          ),
          <fpage>423</fpage>
          -
          <lpage>469</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Hanseung</surname>
            <given-names>Lee</given-names>
          </string-name>
          , Jaeyeon Kihm, Jaegul Choo, John Stasko, and
          <string-name>
            <given-names>Haesun</given-names>
            <surname>Park</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>iVisClustering: An Interactive Visual Document Clustering via Topic Modeling</article-title>
          .
          <source>Computer Graphics Forum</source>
          <volume>31</volume>
          ,
          <issue>3pt3</issue>
          (
          <year>2012</year>
          ),
          <fpage>1155</fpage>
          -
          <lpage>1164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>Ralf</given-names>
            <surname>Lindner</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ulrich</given-names>
            <surname>Riehm</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Electronic petitions and institutional modernization. International parliamentary e-petition systems in comparative perspective</article-title>
          .
          <source>JeDEM-eJournal of eDemocracy and Open Government</source>
          <volume>1</volume>
          ,
          <issue>1</issue>
          (
          <year>2009</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Quoc V Le,
          <string-name>
            <given-names>and Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Exploiting similarities among languages for machine translation</article-title>
          .
          <source>arXiv preprint arXiv:1309.4168</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. David Mimno,
          <string-name>
            <given-names>Hanna M.</given-names>
            <surname>Wallach</surname>
          </string-name>
          , Edmund Talley, Miriam Leenders, and
          <string-name>
            <surname>Andrew McCallum</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Optimizing semantic coherence in topic models</article-title>
          .
          <source>Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          <volume>2</volume>
          (
          <year>2011</year>
          ),
          <fpage>262</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Sergey</surname>
            <given-names>I. Nikolenko</given-names>
          </string-name>
          , Sergei Koltcov, and
          <string-name>
            <given-names>Olessia</given-names>
            <surname>Koltsova</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Topic modelling for qualitative studies</article-title>
          .
          <source>Journal of Information Science</source>
          <volume>43</volume>
          ,
          <issue>1</issue>
          (
          <year>2017</year>
          ),
          <fpage>88</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>Paul</given-names>
            <surname>Hitlin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>'We the People': Five Years of Online Petitions</article-title>
          .
          <source>Technical Report</source>
          . Pew Research Center.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25. Daniel Ramage, Susan Dumais, and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Liebling</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Characterizing Microblogs with Topic Models</article-title>
          .
          <source>Proceedings of the Fourth International AAAI Conference on Weblogs and Social Media</source>
          (
          <year>2010</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Daniel</surname>
            <given-names>Ramage</given-names>
          </string-name>
          , David Hall,
          <string-name>
            <given-names>Ramesh</given-names>
            <surname>Nallapati</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2009a</year>
          .
          <string-name>
            <surname>Labeled</surname>
            <given-names>LDA</given-names>
          </string-name>
          :
          <article-title>A supervised topic model for credit attribution in multi-labeled corpora</article-title>
          .
          <source>In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing: Volume 1-Volume 1. Association for Computational Linguistics</source>
          ,
          <fpage>248</fpage>
          -
          <lpage>256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Daniel</surname>
            <given-names>Ramage</given-names>
          </string-name>
          , David Hall,
          <string-name>
            <given-names>Ramesh</given-names>
            <surname>Nallapati</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2009b</year>
          .
          <string-name>
            <surname>Labeled</surname>
            <given-names>LDA</given-names>
          </string-name>
          :
          <article-title>A supervised topic model for credit attribution in multi-labeled corpora</article-title>
          .
          <source>Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing August</source>
          (
          <year>2009</year>
          ),
          <fpage>248</fpage>
          -
          <lpage>256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Michael</surname>
            <given-names>Röder</given-names>
          </string-name>
          , Andreas Both, and
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Hinneburg</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Exploring the Space of Topic Coherence Measures</article-title>
          .
          <source>In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining (WSDM '15)</source>
          . ACM, New York, NY, USA,
          <fpage>399</fpage>
          -
          <lpage>408</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Amin</surname>
            <given-names>Sorkhei</given-names>
          </string-name>
          , Kalle Ilves, and
          <string-name>
            <given-names>Dorota</given-names>
            <surname>Glowacka</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Exploring Scientific Literature Search Through Topic Models</article-title>
          .
          <source>In Proceedings of the 2017 ACM Workshop on Exploratory Search and Interactive Data Analytics (ESIDA '17)</source>
          . ACM,
          <volume>65</volume>
          -
          <fpage>68</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30. Greg Ver Steeg and
          <string-name>
            <given-names>Aram</given-names>
            <surname>Galstyan</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Discovering Structure in High-Dimensional Data Through Correlation</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31. Greg Ver Steeg and
          <string-name>
            <given-names>Aram</given-names>
            <surname>Galstyan</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Discovering structure in high-dimensional data through correlation explanation</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          .
          <volume>577</volume>
          -
          <fpage>585</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <given-names>Sida</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Baselines and bigrams: Simple, good sentiment and topic classification</article-title>
          .
          <source>In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2. Association for Computational Linguistics</source>
          ,
          <fpage>90</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>