<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modelling student online behaviour in a virtual learning environment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lucie Vachovayy Annika Wolffy</string-name>
          <email>vachova@fm.vse.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Knowledge Media Institute, The Open University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Student Data</institution>
          ,
          <addr-line>Distance Learning, Predictive Models, Machine Learning, Information Visualisation</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Economics</institution>
          ,
          <addr-line>Prague</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>In recent years, distance education has enjoyed a major boom. Much work at The Open University (OU) has focused on improving retention rates in these modules by providing timely support to students who are at risk of failing the module. In this paper we explore methods for analysing student activity in online virtual learning environment (VLE) { General Unary Hypotheses Automaton (GUHA) and Markov chain-based analysis { and we explain how this analysis can be relevant for module tutors and other student support sta . We show that both methods are a valid approach to modelling student activities. An advantage of the Markov chain-based approach is in its graphical output and in the possibility to model time dependencies of the student activities.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>D.4.8 [Performance]: Modelling and Prediction;
H.2.8 [Database Applications]: Data Mining</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>The recent years have seen a massive growth of di
erent possibilities of online education, such as the well known
This work was carried out at the Knowledge Media
Institute.</p>
      <p>Permission to make digital or hard copies of all or part of this work for
personal or classroom use is granted without fee provided that copies are
not made or distributed for profit or commercial advantage and that copies
bear this notice and the full citation on the first page. To copy otherwise, to
republish, to post on servers or to redistribute to lists, requires prior specific
permission and/or a fee. Request permissions from Permissions@acm.org.
LAK ’14, March 24 – 28 2014, Indianapolis, IN, USA
Copyright 2014 ACM 978-1-4503-2664-3/14/03 ...$15.00.
massive open online courses (MOOCs) [Cormier,2008]. The
concept of distance education is however not new. The Open
University is an institution with over forty years of
experience with distance education, historically based on o -line
materials and nowadays making an increasing use of the
Internet. The great advantage of the online courses is in the
fact they are accessible to virtually anyone with Internet
access.</p>
      <p>The other side of the coin is that the retention rates in
these courses are often low. [Koller et al.,2013] mention, that
an average retention rate of a Coursera1 course is around
5%. The situation at traditional universities as well as at
The Open University is signi cantly better, however, there
is still a room for improvement.</p>
      <p>There might be many reasons for the low retention rates,
from the fact that the online courses are often o ered to
anybody interested to the fact that the performance of each
student depends almost exclusively on how much are they
willing to study on their own at home. Our work at The
Open University aims at analysing students' activities in the
online courses in order to gain insight into their behavioural
patterns, which can be utilised for building prediction
models.
1.1</p>
    </sec>
    <sec id="sec-3">
      <title>Problem Description</title>
      <p>The Open University2 is the biggest university in the
United Kingdom, o ering several hundred distance learning
modules, which can be studied both as standalone modules or as
part of a university degree. Anybody can sign up for a
module provided by The OU, without any previous education
whatsoever. The students receive their study material and
submit their assignments through an online virtual learning
environment.</p>
      <p>Students participating in a module are generally split into
smaller study groups of no more than few tens of students,
typically according to their geographic location. Each group
has an assigned tutor. The tutors grade the students'
assignments and exams, answer their questions in the online
forums, provide general advice and guidance, etc.</p>
      <p>In order to support the students who are at risk of
failing the module The OU also implements various
interven1https://www.coursera.org/, a well known and one of the
biggest platforms providing open online courses.
2http://www.open.ac.uk/
tions (such as phone calls from a specialised student support
teams) during the course of the module. Because the number
of students studying each module can reach several thousand
(the modules used in our analysis have enrolment of around
two thousand students) and the resources available for the
interventions are limited, the interventions have to be
carefully planned. Therefore, an important question one might
ask is how to identify students at risk of failing the module
so that the intervention is meaningful and e cient.</p>
      <p>Improving student retention through these focused
interventions and helping the tutors to focus on the students,
who require help, provides many bene ts, from improved
student satisfaction to nancial savings for the university.
In order to support the identi cation of students who are
currently at risk, we utilise several statistical and machine
learning methods. The available data contain both the
information about the students' activity in the VLE as well
as their demographic information. However, for modelling
student behaviour, only the data from the VLE was used.
A more detailed description of our data set can be found in
Section 3.
2.</p>
    </sec>
    <sec id="sec-4">
      <title>PREVIOUS WORK</title>
      <p>The current work builds on previous research done at
The Open University. The initial experiments with machine
learning techniques were using the VLE and assessment data
[Wol and Zdrahal,2012]. One of the main ndings of this
research was that decision trees generally outperform the
other methods [Wol and Zdrahal,2012]. This research also
included creation of a dashboard providing the university
sta with real-time information about student performance.</p>
      <p>Additional methods were tested in [Wol et al.,2013b] and
in [Wol et al.,2013a]. In the latter work, demographic data
were added to the predictions, however this research did
not con rm that this data provide a signi cant increase in
performance.</p>
    </sec>
    <sec id="sec-5">
      <title>DATA SPECIFICATION</title>
      <p>The analyses we performed were done using real data of
several modules from The Open University. We examined a
number of subsequent presentations of each module.</p>
      <p>The available data contain two types of information:
Information about the results of student assignments
(TMAs { tutor marked assignments). There are
several assignments in each module, typically between ve
and seven. Generally, the module is ended by a nal
exam.</p>
      <p>Data about student activity from the virtual learning
environment (VLE).</p>
      <p>The VLE data are aggregated by days and content type
(e.g. forum, wiki, resource, ...). This means that for each
day we know how many times did the student interact with
given content type. For our analysis we summarise the data
by weeks and content types. Summarising the data by weeks
seemed to be reasonable, it simpli es our analyses without
loosing too much detail. The features generated by the
summarisation and used in the methods are:
click counts aggregated by week,
click counts aggregated by week and content type,
binary ags indicating whether student was active in
the VLE and in various content types.
4.</p>
    </sec>
    <sec id="sec-6">
      <title>METHODS</title>
      <p>For analysis of student behaviour in the virtual
learning environments, we have used two di erent approaches {
GUHA [Hajek et al.,1966] and modelling based on Markov
chains [Norris,1998].
4.1</p>
    </sec>
    <sec id="sec-7">
      <title>Activity types analysis</title>
      <p>As mentioned in Section 3, the VLE data contain
information about the type of content the student accessed. The
content type can be for example forum, wiki, resource, quiz,
etc. In total there are 11 di erent content types. Using the
binary ags, indicating whether student was active in given
week and content type, we utilised Bayes Theorem [Bishop
and Nasrabadi,2006] for determining the probability that the
student will fail to complete the module. Moreover, we
analysed each of the content types in terms of mean number of
students succeeding in the module based on activity or
inactivity in the given content type. Based on this investigation
we have selected a set of content types which signi cantly
in uence students' performance, these were then used in the
further analyses.
4.2</p>
    </sec>
    <sec id="sec-8">
      <title>GUHA</title>
      <p>General Unary Hypotheses Automaton (GUHA), originally
published in [Hajek et al.,1966], is one of the oldest data
mining methods for automatic discovery of new interesting
hypotheses from the data. To achieve this goal, GUHA uses
various di erent procedures (ASSOC, IMPL, CORREL). The
choice of the procedure to use depends mainly on the user
needs and his experience. We have selected the ASSOC
procedure [Rauch and Simunek,2001], which allows to discover
interesting associations between attributes in the data. The
interestingness of the association is mostly based on their
co-occurrence. The ASSOC procedure allows to limit the
resulting rules by specifying constraints on the attributes.
This property is important for our eld of interest.</p>
      <p>For our research, we used the ASSOC procedure that is
implemented in the 4ft-Miner module within the Lisp-Miner
software tool 3. The speci cation of the constraints enabled
us to restrict the rules only to those, which cover students
that fail or succeed in the TMA. We used three basic types
of features for the analysis introduced in Section 3.</p>
      <p>The search space for both types of binary ags was
reasonable to perform analysis. For weekly aggregated counts,
it was necessary to reduce the search space via interval
discretisation. For this purpose, we utilized LISp-miner
and unsupervised equal frequency discretisation [Wong and
Chiu,1987].</p>
      <p>GUHA method produce large set of hypotheses. The
example of such results with the categorized binary ags are
depicted in the Figure 1. Unfortunately the information
contained in the output is di cult to interpret. Moreover, the
information of the time dimension is lost and this is even
worse when using various content types. For us, this was
the motivation to look for another modelling method.
4.3</p>
    </sec>
    <sec id="sec-9">
      <title>Markov chain-based analysis of student activity</title>
      <p>3LISp-Miner lispminer.vse.cz/ { software tool for
implementation of the GUHA method.</p>
      <p>In this part of analysis we examined the di erences in
intensity of student activity between students who were
successful in the rst TMA (TMA 1) and those who did not
submit TMA 1. Students who failed in TMA 1 are not
included in this analysis, due to the fact that they represent
only a small portion of students who submit TMA 1. In this
stage of research we analysed only students who at least once
had zero VLE activity in one of the weeks under
consideration. In VLE passive students represent the speci c group
important while looking for potentially at-risk students. On
this group we have studied di erent scenarios - the list of
them is displayed in Table 1. Moreover, the Table 1 shows
percentage of students who behaved according to given
Scenario and were successful in TMA 1, in comparison to those
who did not submit TMA 1, as well. Numbers in column
Scenario represents particular weeks of course.</p>
      <p>Based on the data in Table 1, we can identify behaviour of
at-risk students. This is evident especially in scenarios 2, 3,
4 and 7. This shows us that students who tend to reach zero
VLE activity in later weeks are more probable not tu submit
(scenarios 2, 3, 4). On the other hand those who have zero
VLE activity in earlier weeks and later start to show interest
represented by VLE activity, raise their chance to success in
TMA 1 (scenarios 6, 7, 8).</p>
      <p>Figures 2 and 3 specify more closely the situations from
the scenario 3 (with the TMA 1 not submitted, and TMA
1 passed, respectively). Colour tones of arrows di er (from
white to red) depending on the percentage of students who
moved in given direction. The more red the colour the bigger
the percentage of students it represents. The rows represent
the weeks in which VLE activity was measured. First row
shows activity before the beginning of the course (Week 0),
the other four rows capture the VLE activity in week 1, 2,
3 and 4 respectively. The columns represent the intervals of
VLE activity - di erent colour of the node mean di erent
interval. First column is zero VLE activity, while in other
columns the activity is divided into intervals with the cut
points in multiples of 30.</p>
      <p>This type of analysis enables us to look for speci c
patterns in students' behaviour. In the similar way we can
analyse the di erent types of VLE activities as shown in
Figure 4. In this case the nodes represent not the intensity
of student activity, but capture the interest of student in
speci c content type. And (unlike the previous two gures)
this directed acyclic graph as a whole depicts the Markov
chain [Norris,1998].</p>
      <p>In this paper we have examined two methods for analysing
activity of students in the online virtual learning
environment before the rst tutor marked assignment { GUHA and
Markov chain-based graphical models. Both methods
provide useful insights into the students' behaviour during their
studies. The bene t of the latter lies mostly in its
graphical output, which might be easier to interpret and could
potentially provide support in planning interventions, and
in the possibility to model time dependencies of the student
activities. We believe that the understanding of the student
behavioural patterns will also be useful for building better
predictive models of student performance.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Christopher M Bishop and Nasser M Nasrabadi</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Pattern recognition and machine learning</article-title>
          , volume
          <volume>1</volume>
          . springer New York.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Dave</given-names>
            <surname>Cormier</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>The CCK08 MOOC{connectivism course, 1/4 way</article-title>
          .
          <source>Dave's Educational Blog</source>
          ,
          <volume>2</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Petr</given-names>
            <surname>Hajek</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Havel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michal</given-names>
            <surname>Chytil</surname>
          </string-name>
          .
          <year>1966</year>
          .
          <article-title>The guha method of automatic hypotheses determination</article-title>
          .
          <source>Computing</source>
          ,
          <volume>1</volume>
          (
          <issue>4</issue>
          ):
          <volume>293</volume>
          {
          <fpage>308</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Daphne</given-names>
            <surname>Koller</surname>
          </string-name>
          , Andrew Ng, Chuong Do, and
          <string-name>
            <given-names>Zhenghao</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Retention and intention in massive open online courses: In depth</article-title>
          . EDUCAUSE, June.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>James R Norris</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Markov chains</article-title>
          .
          <source>Number 2008 in Cambridge series in statistical and probabilistic mathematics</source>
          . Cambridge university press.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Jan</given-names>
            <surname>Rauch</surname>
          </string-name>
          and
          <string-name>
            <given-names>Milan</given-names>
            <surname>Simunek</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Mining for association rules by 4ft-miner</article-title>
          .
          <source>In INAP</source>
          , pages
          <volume>285</volume>
          {
          <fpage>295</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Annika</given-names>
            <surname>Wol</surname>
          </string-name>
          and
          <string-name>
            <given-names>Zdenek</given-names>
            <surname>Zdrahal</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Improving retention by identifying and supporting "at-risk" students</article-title>
          .
          <source>EDUCAUSE Review Online, July/Summer.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Annika</given-names>
            <surname>Wol</surname>
          </string-name>
          , Zdenek Zdrahal, Drahomira Herrmannova, and
          <string-name>
            <given-names>Petr</given-names>
            <surname>Knoth</surname>
          </string-name>
          . 2013a.
          <article-title>Predicting student performance from combined data sources</article-title>
          . In Alejandro Pen~
          <fpage>a</fpage>
          -Ayala, editor,
          <source>Educational Data Mining: Applications and Trends, number 524 in Studies in Computational Intelligence</source>
          , pages
          <fpage>175</fpage>
          {
          <fpage>202</fpage>
          . Springer International Publishing, Cham.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Annika</given-names>
            <surname>Wol</surname>
          </string-name>
          , Zdenek Zdrahal, Andriy Nikolov, and
          <string-name>
            <given-names>Michal</given-names>
            <surname>Pantucek</surname>
          </string-name>
          . 2013b.
          <article-title>Improving retention: predicting atrisk students by analysing clicking behaviour in a virtual learning environment</article-title>
          .
          <source>In Third Conference on Learning Analytics and Knowledge (LAK</source>
          <year>2013</year>
          ).
          <source>ISBN 978-1-4503-1785-6.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Andrew K. C. Wong</surname>
            and
            <given-names>David K. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Chiu</surname>
          </string-name>
          .
          <year>1987</year>
          .
          <article-title>Synthesizing statistical knowledge from incomplete mixedmode data</article-title>
          .
          <source>IEEE Trans. Pattern Anal. Mach</source>
          . Intell.,
          <volume>9</volume>
          (
          <issue>6</issue>
          ):
          <volume>796</volume>
          {
          <fpage>805</fpage>
          , June.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>