<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Dynamic Topic Model of Learning Analytics Research</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Derntl</string-name>
          <email>derntl@dbis.rwth-</email>
          <email>derntl@dbis.rwthaachen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikou Günnemann</string-name>
          <email>nikou@dbis.rwth-</email>
          <email>nikou@dbis.rwthaachen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralf Klamma</string-name>
          <email>klamma@dbis.rwth-</email>
          <email>klamma@dbis.rwthaachen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>RWTH Aachen University</institution>
          ,
          <addr-line>Advanced Community, Information Systems (ACIS), Aachen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Research on learning analytics and educational data mining has been published since the rst conference on Educational Data Mining (EDM) in 2008 and gained momentum through the establishment of the Learning Analytics and Knowledge (LAK) conference in 2011. This paper addresses the LAK Data Challenge from the perspective of visual analytics of topic dynamics in the LAK Dataset between 2008 and 2012. The data set was processed using probabilistic, dynamic topic mining algorithms. To enable exploration and visual analysis of the resulting topic model by LAK researchers and stakeholders we developed and deployed DVITA, a web-based browsing tool for dynamic topic models. In this paper we explore answers to the questions about past, present, and future of LAK posed in the Data Challenge based on a topic model of all papers in the LAK Dataset. We also brie y describe how users can explore the LAK topic model on their own using D-VITA.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Copyright 2013 by the authors
questions posed in the LAK Data Challenge into the
user's hands. D-VITA is a web-based tool that o ers
topic-based views on the LAK Dataset using a
pointand-click metaphor and simple visualizations.</p>
    </sec>
    <sec id="sec-2">
      <title>2. DATASET AND PREPROCESSING</title>
      <p>
        The LAK Dataset underlying the analyses presented in this
paper includes the EDM conference proceedings 2008{2012
(239 papers), the LAK conference proceedings 2011{2012
(66 papers), and the papers of the 2012 Special Issue on
Learning Analytics in the Educational Technology and
Society journal (10 papers; herafter referred to as ETS). The
RDF representation of the LAK Dataset was processed by a
script that extracted for each paper the identi er, venue 2
fLAK, EDM, ETSg, year of publication, title, authors,
abstract, full text, and hyperlink to the full RDF description
on data.linkededucation.org. The distribution of the 315
papers over time and venues is given in Figure 1.
In the next preprocessing step the paper records were cleaned
by removing stopwords and by applying stemming methods
on the included word sets. For word stemming we used the
Porter Stemming technique [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which is well established for
this purpuse. As a result, close to 5000 distinct word stems
were identi ed as being used in the 315 papers.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. DYNAMIC TOPIC MINING</title>
      <p>
        From a text mining perspective the LAK Dataset represents
a text corpus in which a set of words is used in a set of
papers. To identify what is relevant to LAK research, we used
the dynamic topic modeling approach described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to
obtain the distribution of words over a pre-de ned number of
topics. This is a probabilistic, unsupervised machine
learning approach that has been gaining increasing prominence
recently [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In these probabilistic topic models a topic is a
distribution of words, so each topic is typically represented
by its most frequently occurring words. Topic mining also
obtains the distribution of these topics over the papers in
the data set. Dynamic topic mining applies these
analysis steps using several consecutive time slices in the data
set. For the LAK Dataset, we chose the ve calendar years
2 f2008 : : : 2012g as time slices. The results will thus reveal
the evolution of topics over documents during those discrete
time slices, and the evolution of words used in the papers
for each topic over time.
      </p>
      <p>Dynamic topic mining requires the analyst to pre-set the
number of topics. Based on previous experiments with
varying numbers of topics in paper collections in well-de ned
subject areas, we decided to run the analysis of the LAK
Dataset with a set of 20 topics. This number, while
somewhat arbitrary, shall provide for su cient discriminatory
power for both the distribution of topics over papers and the
distribution of words over topics. With fewer topics, terms
like `learning', for instance, are more likely to be present
with relatively high relevance in many topics, while a larger
preset would increase the number of topics exposed in each
paper. Both situations would impede reasonable
interpretation and visualization of the results.</p>
      <p>A word of explanation regarding the labels used to refer to
topics in this paper: mathematically each topic is a
distribution over words. In a dynamic topic model this
distribution changes over time, i.e. a speci c word may rise or
fall in relevance for a topic. In the rest of the paper we
will therefore label each topic with an ordered tuple
representing those words with the highest mean relevance for this
topic over time. In topic modeling literature we found that
four words is a good number to form a topic label. For
instance, for topic \students model parameters skill" the
most relevant word on average is student followed by model,
parameters, and skill. For illustration, based on the word
distribution for this topic in 2008 only, the label would be
\model student skill learning". Often, such word tuples
are rephrased as more expressive labels; for instance \student
modeling" could be appropriate in our example.</p>
      <p>
        The obtained topic model including 20 topics was analyzed
to see whether the topics have su cient discriminatory power.
To this end, we used the ten most important words for
each topic and the corresponding probability distributions to
compute a dissimilarity measure of the distributions by
using the Jensen-Shannon divergence measure [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The matrix
with pairwise divergence values is displayed in Figure 2. The
maximum Jensen-Shannon divergence value is ln(2) :69.
The darker the cell color, the lower the divergence, thus the
higher the similarity. The matrix is generally \light-colored",
indicating that the topics' word distributions diverge to a
high degree. Topic pair (A; S) has the lowest dissimilarity
value, and Figure 3 reveals why: both topics are about
student modeling. Topic S generally appears to have several
loosely related topics.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. ANALYSIS OF LAK TOPIC DYNAMICS</title>
      <p>In this section we explore three questions about the LAK
Dataset, intending to shed some light on the past and present
topics of learning analytics research, along with a cautious
glimpse into the future.</p>
      <p>Question 1: What have been the most relevant topics
overall in the LAK data set?
This question addresses the LAK Data Challenge aspects
of roots and current state of learning analytics. Figure 3
shows an overview chart of the 20 topics identi ed in the
LAK Dataset. The horizontal axis re ects the rank of mean
relevance of each topic and the vertical axis re ects the rank
of stability3 over the ve time slices in the dataset. The size
of each bubble re ects the relevance of the topic in 2012, the
most recent period. We make several observations:</p>
      <p>
        The most relevant topics most prominently feature the
terms students/learners, model, and data. This aligns
well with SoLAR's de nition of learning analytics as
\the measurement, collection, analysis and reporting
of data about learners and their contexts, for purposes
of understanding and optimizing learning and the
environments in which it occurs,"[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] considering that
understanding and optimization is necessarily based on
models of learners and data.
      </p>
      <p>The topic with the highest mean relevance is \student
model parameters skill" (A); this topic also has the
highest variance in relevance.</p>
      <p>In the top-right quadrant we nd topic \model data
features prediction" (B) which has a strong
relevance in 2012, high mean relevance rank over all years
and a high stability. As such, it can be considered
as one of the core topics in the LAK Dataset. In
2012 the distribution of words in this topic would
advocate the label \prediction model data students",
i.e. prediction is currently most relevant for this topic.</p>
      <p>Topic \network community discussion analysis" (R)
is also worth looking at. While it is relatively
irrelevant and volatile, it is among the relevant topics in
2012 (cf. the bubble size). The topic evolution chart
in Figure 4 reveals that this topic accumulated most
of its relevance in 2011, the year of the rst LAK
conference. Also, 8 of the 10 papers with the strongest
focus on this topic in 2011 were published in the LAK
conference (see bottom portion of Figure 4) although
EDM published 2.5 times the number of papers in that
3Stability was computed by inverting the variance of the
topic's relevance over time
year. This topic, in 2012 represented by the word
order \network community social user", therefore
appears to be a genuine LAK topic which was previously
rather irrelevant for the EDM conference.</p>
      <p>Question 2: What changes in topic dynamics did the
first LAK conference in 2011 bring about?
This question aims to reveal whether and how the LAK
community relates to the EDM community in terms of topics
covered by their papers. To explore this we look (a) at at
the overall distribution of topics over time and (b) at the
relative change of topic relevance between 2010 and 2011.</p>
      <p>
        The evolution of the overall distribution of topics is
illustrated in the ThemeRiver in Figure 5. In a ThemeRiver [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
the horizontal axis represents the points in time to which the
documents in a dataset belong (in the LAK Dataset that is
the publication date), and the vertical axis represents the
relevance of the topic. Each current in the ThemeRiver
therefore presents the dynamic development of a selected
topic over time. The wider the current, the more relevant is
the topic, i.e. the more documents expose this topic. Since
each document exposes di erent topics to varying degrees
the relevance of topic k at time t is formally de ned as
relevance(k; t) := jD1tj Pd2Dt d[k], where Dt is the set of
documents belonging to time t, and d is the topic
distribution for document d. Observing the ThemeRiver in Figure 5
it is evident that there were some shifts in topic focus during
the years 2008 and 2010, where we have only the EDM
papers in the dataset. Between 2010 and 2011 we identify the
strongest turbulence, presumably based on substantial shifts
in topic foci introduced by the 2011 LAK conference.
Interestingly the topic distribution remains rather stable during
the last time slice, in which LAK 2012, EDM 2012 and the
ETS special issue are included. This might suggest that
these three publication venues propelled the convergence of
LAK research as represented in the LAK Dataset.
      </p>
      <p>To see which topics rose in relevance between 2010 and
2011 we lter for topics and zoom into the transition
between 2010 and 2011 as illustrated in Figure 6. Those three
topics that have their absolute highest relevance in 2011
are marked with an up-pointing triangle with a solid-black
outline. These are \model students data probability",
\network community discussion analysis", and \problem
students model types", indicating an increased focus on
student modeling as well as community and network
analysis through the rst LAK conference in 2011.</p>
      <p>Question 3: What topics rose the most in 2012, the
most recent time slice in the data set?
This question looks into what the dynamic topic model of
the LAK Dataset suggests as rising topics over the next
year(s). We try to answer this by identifying those ve
topics that had the highest rise in relevance between 2011 and
2012. The topic labels represent the word distribution in
2012, and the number in parentheses indicates the absolute
gain in relevance:
6. CONCLUSION
1. students data courses system (+.054) In a nutshell, we discovered the following: Regarding the
2. students interaction participants analysis (+.036) past, we found that LAK and EDM do have a substantial
3. learning analytics social learners (+.035) shared topic foundation including themes like student
mod4. students actions learning state (+.025) eling, data classi cation, and clustering. We also found that
5. data user learning dataset (+.013) the EDM conference series had some turbulence in topical
focus between 2008 and 2010, the time window when only
EDM papers are present in the dataset.</p>
      <p>The Document and Word Evolution Panel shows for the
selected topic an ordered list of the most relevant papers
in the \Relevant Documents" tab. The icons next to each
document allow showing the topic pie for the document
and its content, respectively. The \Similar Docs" icon will
bring up the Document Browser with a list of similar
documents. Under the \Word Evolution" tab the user will nd a
ThemeRiver illustrating the evolution of the distribution of
words in the selected topic over time.</p>
      <p>D-VITA also o ers a Document Browser to perform
keywordbased search, explore the topic distribution of documents,
and navigate documents based on similarity.</p>
      <p>In sum these ve topics have accumulated a share of 42% of
the topic distribution by 2012, starting from 11% in 2008 (cf.</p>
      <p>Figure 7). These developments indicate a strong increase in
focus on the students' activities and actions in courses as
well as social and interaction analytics.</p>
    </sec>
    <sec id="sec-5">
      <title>5. D-VITA TOPIC ANALYTICS TOOKIT</title>
      <p>Except for Figures 1 and 3 all gures were produced using
D-VITA, a web-based visual analytics tool we developed and
deployed for visual analytics of dynamic topic models. The
tool allows users to visually interact with the output of the
dynamic topic mining algorithms on the LAK Dataset. The
application window shown in Figure 8 has three panels:
The Topics Panel shows the list of topics obtained by the
dynamic topic modeling algorithm; topics can be sorted by
rising, falling and mean relevance, as well as variance of
relevance. The topics can be ltered using keywords; in the
screen shot the keyword \visual" is used as a lter. The topic
list thus only includes topics whose set of relevant words
includes this word stem. Topics checked by the user will be
visualized in the ThemeRiver in the Topic Evolution Panel.
The Topic Evolution Panel shows a ThemeRiver of
evolution of relevance of the topics selected in the Topics Panel.
Data points for each topic and time slice, respectively, can
be clicked, which will trigger the display of detailed
information on the clicked topic at the selected time slice in the
Document and Word Evolution Panel.</p>
      <p>Regarding the present we found that the LAK Dataset
exposes a strong emphasis on learner modeling, data
modeling, analysis and prediction. The rst LAK conference in
2011 also brought some considerable shifts in topic focus;
e.g. LAK 2011 has visibly strengthened network and social
analysis aspects on top of EDM topics.</p>
      <p>Regarding the near future we found that the shifts in the
topics' proportions in 2012 appear rather moderate, thus
indicating a phase of convergence of LAK research topics.
Projecting recent topic shifts into the future, we can expect
increased emphasis on social and interaction aspects and a
sustained, strong role of students as research subjects.</p>
    </sec>
    <sec id="sec-6">
      <title>7. ACKNOWLEDGMENTS</title>
      <p>This work was supported by the European Commission through
the the support action TEL-Map (FP7-257822) and the
integrated project Layers (FP7-318209).</p>
      <p>Figure 8: Application window showing ThemeRiver and document list (rotated image)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>About</surname>
            <given-names>SoLAR</given-names>
          </string-name>
          ,
          <year>2012</year>
          . http://www.solaresearch.org/mission/about/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          .
          <article-title>Probabilistic topic models</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>55</volume>
          (
          <issue>4</issue>
          ):
          <volume>77</volume>
          {
          <fpage>84</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          and
          <string-name>
            <surname>J. D.</surname>
          </string-name>
          <article-title>La erty. Dynamic topic models</article-title>
          .
          <source>In ICML</source>
          , pages
          <volume>113</volume>
          {
          <fpage>120</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Havre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. G.</given-names>
            <surname>Hetzler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Whitney</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. T.</given-names>
            <surname>Nowell</surname>
          </string-name>
          . Themeriver:
          <article-title>Visualizing thematic changes in large document collections</article-title>
          .
          <source>IEEE Trans. Vis. Comput. Graph.</source>
          ,
          <volume>8</volume>
          (
          <issue>1</issue>
          ):9{
          <fpage>20</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>Divergence Measures Based on the Shannon Entropy</article-title>
          .
          <source>IEEE Transactions on Information Theory</source>
          ,
          <volume>37</volume>
          (
          <issue>1</issue>
          ),
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>An algorithm for su x stripping</article-title>
          .
          <source>Program</source>
          ,
          <volume>14</volume>
          (
          <issue>3</issue>
          ):
          <volume>130</volume>
          {
          <fpage>137</fpage>
          ,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Taibi</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Dietze</surname>
          </string-name>
          .
          <article-title>Fostering analytics on learning analytics research: the LAK dataset</article-title>
          ,
          <source>Technical Report</source>
          , 03/
          <year>2013</year>
          ,
          <year>2013</year>
          . http://resources.linkededucation. org/
          <year>2013</year>
          /03/lak-dataset-taibi.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>