<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Cognitive Presence Detection in Online Discussion Transcripts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vitomir Kovanovic</string-name>
          <email>vitomir_kovanovic@sfu.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Srecko Joksimovic</string-name>
          <email>sjoksimo@sfu.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marek Hatala</string-name>
          <email>mhatala@sfu.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dragan Gasevic</string-name>
          <email>dgasevic@acm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Athabasca University</institution>
          ,
          <addr-line>Edmonton, AB</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Simon Fraser University</institution>
          ,
          <addr-line>Vancouver, BC</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present the results of an exploratory study that examined the use of text mining and text classification for the automation of the content analysis of discussion transcripts within the context of distance education. We used Community of Inquiry (CoI) framework and focused on the content analysis of the cognitive presence construct given its central position within the CoI model. Our results demonstrate the potentials of proposed approach; The developed classifier achieved 58.4% accuracy and Cohen's Kappa of 0.41 for the 5-category classification task. In this paper we analyze different classification features and describe the main problems and lessons learned from the development of such a system. Furthermore, we analyzed the use of several novel classification features that are based on the specifics of cognitive presence construct and our results indicate that some of them significantly improve classification accuracy.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        One of the important aspects of modern distance education is the
focus on the social construction of the knowledge by the means of
asynchronous discussion groups [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Their increased usage in
distance education has produced an abundant amount of records on
the learning processes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Educational researchers recognized the
importance of this "gold-mine of information" [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] about the
learning process, and used it mainly for research, usually long after the
courses are over. Nowadays, there is a need to analyze this learners
generated data in automatic and continuous fashion in order to
inform instructors, and student about the current student performance
and possible learning outcomes. Learning Analytics, an emerging
research field that aims to make a sense of the large volume of
educational data in order to understand and improve learning [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], is a
promising area of research that could be successfully used to
analyze and understand the discussion transcript logs in their full
complexity. However, at the moment the majority of the approaches for
analysis of discussion transcripts are not based on the established
Permission to make digital or hard copies of all or part of this work for
personal or classroom use is granted without fee provided that copies are not
made or distributed for profit or commercial advantage and that copies bear
this notice and the full citation on the first page. Copyrights for components
of this work owned by others than ACM must be honored. Abstracting with
credit is permitted. To copy otherwise, or republish, to post on servers or to
redistribute to lists, requires prior specific permission and/or a fee. Request
permissions from Permissions@acm.org.
      </p>
      <p>
        LAK ’14 March 24 - 28 2014, Indianapolis, IN, USA
Copyright 2014 ACM 978-1-4503-2664-3/14/03 ...$15.00.
theories of educational research, and focus mostly on the
quantitative aspects of the trace and log data. Given the need to assess
the qualitative aspects of the learning products this is not enough.
To address this issue, we base our transcript analysis approach on
the well established Community of Inquiry (CoI) model of distance
education [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ] which is used for more than a decade to answer
this type of questions.
      </p>
      <p>In this paper we present the results of a study that focused on the
automation of the content analysis of discussion transcripts using
Community of Inquiry coding technique. We developed an
SVMbased classifier for automatic classification of the discussion
transcripts in accordance with the CoI framework, and we discuss in
detail the challenges and issues with this type of text classification,
most notably the creation of the relevant classification features.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>BACKGROUND WORK</title>
      <p>We based our work on the theoretical foundations of the
Community of Inquiry framework and previous work done in the field
of text classification. In this section we will present an overview of
the Community of Inquiry framework and the relevant findings in
text classification field that informed our approach.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Community of Inquiry (CoI) Framework</title>
      <p>
        Among the different techniques for assessment of quality of
distance education environments, one of the best-researched models
that comprehensively explain different dimensions of social
learning is Community of Inquiry (CoI) model [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. The model
consists of the three interdependent constructs that together provide
comprehensive coverage of distance learning phenomena [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]:
i) Social presence describes relationships and social climate in a
learning community [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], ii) Cognitive presence describes the
different phases of students’ cognitive engagement and knowledge
construction [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and iii) Teaching presence explains the
instructor’s role in the course planning and execution [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        For our study the most important is the Cognitive Presence
construct which is defined as “an extent to which the participants in
any particular configuration of a community of inquiry are able to
construct meaning through sustained communication.” [10, p .89].
The model defines four different phases of cognitive presence:
1. Triggering event: In this phase some issues, dilemma or
problem is identified. Often, in the formal educational context,
they are explicitly defined by the instructors, but also can be
created by any student that participates in the discussions [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
2. Exploration: In this phase students move between their
private reflective world and shared world where social
construction of knowledge happens [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
3. Integration: This phase is characterized by the synthesis of
the ideas that are generated in the exploration phase and
ultimately construction of meaning.
4. Resolution: In this phase students analyze practical
applicability of the generated knowledge, test different hypotheses,
and ultimately start a new cycle of knowledge construction
by generating a new triggering event.
      </p>
      <p>
        The framework comes with its own content analysis scheme and
it attracted a lot of attention in the research community resulting
in the fairly large number of replication studies and empirical
testing of the framework [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. However, even though Community of
Inquiry proved to be a viable model for assessing learning quality
in an online educational contexts, the practical issues of applying
CoI analysis and its coding scheme remain; It is still a manual, time
consuming process which makes the coding of the messages very
expensive. For example, for the study presented here, it took
approximately one month for the two coders to manually code the
1747 discussion messages. This need for manual coding has been
pointed as one of the main reasons why many transcript analysis
techniques had almost no impact on educational practice and never
moved out of the domain of educational research [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In order to
support broader adoption of CoI framework there is a need for an
automation of the coding process, and that is the exact purpose of
this study. We focus on the coding of the cognitive presence,
however the overall goal is to automate content analysis for all three
CoI presences in order to provide a comprehensive picture of the
learning process. This would allow instructors to adopt CoI
framework for guiding instructional interventions, and to provide
learners with feedback making them more aware of their own learning
and learning of their peers.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Text Classification and Automatic Content</title>
    </sec>
    <sec id="sec-5">
      <title>Analysis Approaches</title>
      <p>
        In order to automate the content analysis of discussion transcripts,
we adopted text mining classification techniques [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. As the
cognitive presence is a latent construct and not clearly observable, we
based our work on the previous work that focused also on mining
latent constructs. The work done on opinion mining of online
product reviews [
        <xref ref-type="bibr" rid="ref15 ref23 ref3">3, 15, 23</xref>
        ], gender style differences [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and sentiment
analysis [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] are some of the main areas of research that informed
our classification approach.
      </p>
      <p>
        The text classification tasks have been studied in the context of
several different areas. In general, the majority of the studies
extensively used lexical features such as N-grams, Part-of-Speech (PoS)
tags and word dependency triplets, or some mixture of them as their
main type of features. For example, for the problem of classifying
online product reviews as either based of qualified or unqualified
claims, Arora et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] used the combination of N-grams, PoS
bigrams and dependency triples with the approximation of syntactic
scope [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Authors achieved Cohen’s Kappa of .353 and
classification accuracy of 72.6% for their binary classification task. For
the similar problem, Joshi and Penstein-Rosé [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] used word
dependency triplets &lt;Rel, HeadWord,ModifierWord&gt; as
features where Rel is a grammatical relation between the words (e.g.,
Adjective), while HeadWord and ModifierWord are either a
concrete words (e.g., Camera, Great) or PoS classes (e.g., Noun,
Adverb). Their study showed that in the context of opinion
mining use of the PoS class as a HeadWord and the concrete word
as a ModifierWord provides a small, but statistically significant
improvement over baseline unigram model [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Another type of features that are also utilized are word pattern
features. For example, for sarcasm detection in online product
reviews Tsur and Davidov [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] used K-nearest neighbors (KNN)
classifier with the patterns of 1-6 content words and 2-6 high
frequency words (i.e., words that occur frequently in many reviews) as
classification features. In the context of stylistic differences among
genders, the idea of pattern features was further expanded by
Gianfortoni et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] with more complex notion of word patterns,
however, reported results showed very modest improvement in
classification accuracy achieving Cohen’s Kappa of only 0.18 in the best
case.
      </p>
      <p>
        Finally, there are other approaches as well, most notably the ones
which are based on the use of Latent Semantic Analysis (LSA) in
the context of automate assessment of student essays [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] or the use
of more complex features which make a use of genetic
programming [
        <xref ref-type="bibr" rid="ref18 ref4">4, 18</xref>
        ].
      </p>
      <p>
        In terms of the classification methods used, the majority of
approaches use K Nearest Neighbors (KNN) or Support Vector
Machines (SVM) algorithms. SVM is particularly popular algorithm
for text classification and according to Aggarwal and Zhai [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], “text
data is ideally suited for SVM classification because of the sparse
high-dimensional nature of text, in which few features are
irrelevant, but they tend to be correlated with one another and
generally organized into linearly separable categories”[pg. 195]. SVM
classifiers also work well with a large number of weak predictors,
which is the case of text classification where typically the majority
of features are very weakly predicting class label [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
3.
3.1
      </p>
    </sec>
    <sec id="sec-6">
      <title>METHODS</title>
    </sec>
    <sec id="sec-7">
      <title>Data set</title>
      <p>
        For the purpose of our study, we used the data set obtained from
a graduate level course in Software Engineering from a Canadian
fully distance learning university. The data set consists of 1747
messages which were coded by two human coders for the levels of
cognitive presence. Coders achieved excellent interrater agreement
(percent agreement=98.1%, Cohen’s Kappa=0.974) indicating the
quality of the content analysis scheme. The most frequent type of
messages were exploration messages occurring on average 39% of
the time (Table 1) while the least frequent were resolution messages
occurring on average in 6% of the cases. These large differences
in the category distributions are not surprising as they are shown
by the previous work in CoI research field. The reason for this is
that the majority of students are not progressing to the later stages
of cognitive presence [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] which in turn limits the potential for
development of their critical thinking skills. Thus, even though we
have 5 categories, the baseline accuracy using the simplest majority
vote classification is 39%.
N-grams
Part-of-Speech N-grams
Back-Off N-grams
Dependency Triplets
Back-Off Dependency Triplets
Named Entities
Thread Position Features
unigrams,
bigrams,
trigrams
pos-bigrams,
pos-trigrams
bo-bigrams,
bo-trigrams
dep-triplets
h-bo-triplets,
m-bo-triplets,
hm-bo-triplets
entity-count
is-first,
is-reply-first
      </p>
    </sec>
    <sec id="sec-8">
      <title>Feature Extraction</title>
      <p>
        Based on our literature review described in Section 2, we
extracted a wide variety of features that were frequently used in the
similar studies (Table 2). We extracted the commonly used
Ngram features (i.e., unigrams, bigrams and trigrams) and
Part-ofSpeech (PoS) bigrams and trigrams. In addition, similarly to the
works of Joshi and Penstein-Rosé [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] we extracted: i) back-off
versions of bigrams and trigrams by replacing one or more words
in a N-gram by the corresponding PoS tag, and ii) word dependency
triplets and their back-off versions. Finally, in addition to the
features found in the research literature, we extracted an additional set
of features which we thought might be useful given the specifics of
the cognitive presence construct.
      </p>
      <p>
        Given the difference among the phases of Cognitive Presence,
we extracted the entity-count feature, which shows the
number of named entities that were mentioned in the message using
DBPedia Spotlight [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] web service. The rationale behind this
feature is that different phases of cognitive presence could be
potentially characterized by the different number of constructs that were
discussed in the message. For example, it might be the case that
exploration messages contain on average a larger number of concepts,
as one of the key characteristics of exploration is brainstorming of
different problem solutions and ideas [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Another important aspect of cognitive presence is that it
develops over time through the communication with other students [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
In practice this means that triggering and exploration messages
are more likely to be observed in the early stages of discussions,
while integration and resolution messages are more likely in the
later stages of the discussions. To test this hypothesis, as the first
step we extracted two simple features: i) is-first, which
indicates whether a message is the first in the discussion topic, and
ii) is-reply-first which indicates whether a message is the
reply to the original discussion opening message.
3.3
      </p>
    </sec>
    <sec id="sec-9">
      <title>Classifier Implementation</title>
      <p>
        For the purpose of this study we decided to use SVM
classification as it is a well known and popular technique especially well
suited for the purpose of text classification as we described in
Section 2. In order to maximize classification quality and assess the
usefulness of different types of features, we experimented with the
several different sets of features and evaluated them using 10-fold
cross validation, which is considered a good compromise between
sizes of training and test data [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. We used only features that had
support threshold of 10 or more (i.e., occurring 10 or more times in
majority vote baseline
unigrams baseline
+ bigrams
+ trigrams
+ pos-bigrams
+ pos-trigrams
+ bo-bigrams
+ bo-trigrams
+ dep-triplets
+ h-bo-triplets
+ m-bo-triplets
+ hm-bo-triplets
+ entity-count
+ is-first
+ is-reply-first
      </p>
      <p>Additional Classification Cohen’s P -val</p>
      <p>
        Features Accuracy Kappa
the data) in order to keep the number of features reasonable and to
protect from overfitting the classifier to the noise in the data which
is captured by the low supported features. We used linear kernel
and default values of parameters (C = 1; gamma = 1=k). In
order to compare different set of features we used McNemar’s test [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
as it is shown to have low Type I error rate [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        To implement the classifier and feature extraction we used
several popular open source tools and libraries. In the feature
extraction step we used Stanford CoreNLP suite1 of tools for
tokenization, Part-of-Speech tagging [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and dependency parsing [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. We
used the popular Weka [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] data mining toolkit and LibSVM
library [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for developing the classifier, and to implement the
McNemar’s test we used Java Statistical Classes (JSC) library2.
4.
      </p>
    </sec>
    <sec id="sec-10">
      <title>RESULTS</title>
      <p>Table 3 shows the results of our classification experiment. The
baseline unigram model achieved 54.72% accuracy which is slightly
less than in the case of the more complex models with larger
number of features. The biggest improvement was observed by adding
the back-off version of trigrams which improved classification
accuracy to 58.38% and Cohen’s Kappa to 0.41 which is accompanied
with the largest increase in the feature space.</p>
      <p>
        Our results are similar to the ones of Arora et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and Joshi
and Penstein-Rosé [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] with our classifier having somewhat lower
absolute levels of accuracy and a slightly bigger values of Cohen’s
Kappa metric. Our results also show that adding both head-backoff
and modifier-backoff versions of dependency triplets improves the
classification accuracy, as well as the ordinary dependency triplets.
With respect to the three features that we proposed, the use of the
indicators for the number of named entities (i..e, entity-count)
and discussion starters (i.e., is-first) also showed statistically
significant improvement over the baseline unigram model. In
addition, the use of those features has an almost nonexistent impact
on the classifier feature space making the building of classification
model faster and more interpretable.
      </p>
    </sec>
    <sec id="sec-11">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>
        As our results show, the proposed approach for automating
content analysis seems promising. The current level of Cohen’s Kappa
1http://nlp.stanford.edu/software/corenlp.shtml
2http://www.jsc.nildram.co.uk/index.htm
is at the lower end of 0.4-0.7 range which is considered to be a fair
to good agreement [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. However, in order to replace manual
message coding, the Cohen’s Kappa should be above 0.7 level which is
still out of reach.
      </p>
      <p>One important aspect of coding discussion transcripts that we
observed and which does not affect the work that we reviewed is
message quoting. We observed many instances in which student
puts direct quotation of others’ message into his own which makes
a problem for classification based on lexical features such as
Ngrams, PoS tags or Dependency triplets. In our future works we
will look for a ways to address this issue and to estimate the impact
of quoting on classification accuracy.</p>
      <p>We also showed the potential of novel features which are based
on the deeper theoretical understanding of the latent construct
under interest and its coding instrument. They could provide a
significant improvement of the classification accuracy without a big
impact on the feature space complexity.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Aggarwal</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          .
          <source>Mining Text Data</source>
          . Springer, Feb.
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Anderson</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dron</surname>
          </string-name>
          .
          <article-title>Three generations of distance education pedagogy</article-title>
          .
          <source>The International Review of Research in Open and Distance Learning</source>
          ,
          <volume>12</volume>
          (
          <issue>3</issue>
          ):
          <fpage>80</fpage>
          -
          <lpage>97</lpage>
          , Nov.
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Rosé</surname>
          </string-name>
          .
          <article-title>Identifying types of claims in online customer reviews</article-title>
          .
          <source>In Proceedings of the HLT-NAACL</source>
          <year>2009</year>
          , page 37-40, Stroudsburg, PA, USA,
          <year>2009</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mayfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Penstein-Rosé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and E.</given-names>
            <surname>Nyberg</surname>
          </string-name>
          .
          <article-title>Sentiment classification using automatically extracted subgraph features</article-title>
          .
          <source>In Proceedings of the NAACL HLT 2010 Workshop on Computational Approaches to Analysis and Generation of Emotion in Text, page 131-139</source>
          , Stroudsburg, PA, USA,
          <year>2010</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.-C.</given-names>
            <surname>Chang</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.-J.</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>LIBSVM: A library for support vector machines</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          ,
          <volume>2</volume>
          :
          <issue>27</issue>
          :
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          :
          <fpage>27</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T. G.</given-names>
            <surname>Dietterich</surname>
          </string-name>
          .
          <article-title>Approximate statistical tests for comparing supervised classification learning algorithms</article-title>
          .
          <source>Neural Comput.</source>
          ,
          <volume>10</volume>
          (
          <issue>7</issue>
          ):
          <fpage>1895</fpage>
          -
          <lpage>1923</lpage>
          , Oct.
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Donnelly</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Gardner</surname>
          </string-name>
          .
          <article-title>Content analysis of computer conferencing transcripts</article-title>
          .
          <source>Interactive Learning Environments</source>
          ,
          <volume>19</volume>
          (
          <issue>4</issue>
          ):
          <fpage>303</fpage>
          -
          <lpage>315</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Duwairi</surname>
          </string-name>
          .
          <article-title>A framework for the computerized assessment of university student essays</article-title>
          .
          <source>Computers in Human Behavior</source>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <fpage>381</fpage>
          -
          <lpage>388</lpage>
          , May
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Everitt</surname>
          </string-name>
          .
          <article-title>The analysis of contingency tables</article-title>
          .
          <source>Chapman and Hall</source>
          ,
          <year>1977</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Garrison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Anderson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Archer</surname>
          </string-name>
          .
          <article-title>Critical inquiry in a text-based environment: Computer conferencing in higher education</article-title>
          .
          <source>The Internet and Higher Education</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          -3):
          <fpage>87</fpage>
          -
          <lpage>105</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Garrison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Anderson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Archer</surname>
          </string-name>
          .
          <article-title>Critical thinking, cognitive presence, and computer conferencing in distance education</article-title>
          .
          <source>American Journal of Distance Education</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <fpage>7</fpage>
          -
          <lpage>23</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Garrison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Anderson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Archer</surname>
          </string-name>
          .
          <article-title>The first decade of the community of inquiry framework: A retrospective</article-title>
          .
          <source>The Internet and Higher Education</source>
          ,
          <volume>13</volume>
          (
          <issue>1-2</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>9</lpage>
          , Jan.
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gianfortoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Adamson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Rosé</surname>
          </string-name>
          .
          <article-title>Modeling of stylistic variation in social media with stretchy patterns</article-title>
          .
          <source>In Proceedings of the First Workshop on Algorithms and Resources for Modelling of Dialects and Language Varieties</source>
          , page
          <volume>49</volume>
          -59, Stroudsburg, PA, USA,
          <year>2011</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Henri</surname>
          </string-name>
          .
          <article-title>Computer conferencing and content analysis</article-title>
          .
          <source>In A. R</source>
          . Kaye, editor,
          <source>Collaborative Learning Through Computer Conferencing</source>
          , pages
          <fpage>117</fpage>
          -
          <lpage>136</lpage>
          . Springer Berlin Heidelberg, Jan.
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Penstein-Rosé</surname>
          </string-name>
          .
          <article-title>Generalizing dependency features for opinion mining</article-title>
          .
          <source>In Proceedings of the ACLIJCNLP 2009 Conference, page 313-316</source>
          , Stroudsburg, PA, USA,
          <year>2009</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K. H.</given-names>
            <surname>Krippendorff</surname>
          </string-name>
          .
          <article-title>Content Analysis: An Introduction to Its Methodology</article-title>
          .
          <source>Sage Publications</source>
          ,
          <volume>0</volume>
          <fpage>edition</fpage>
          , Dec.
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>M.-C. d. Marneffe</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>MacCartney</surname>
            , and
            <given-names>C. D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Generating typed dependency parses from phrase structure parses</article-title>
          .
          <source>In Proceedings of the International Conference on Language Resources and Evaluation (LREC)</source>
          ,
          <source>page 449-454</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>E.</given-names>
            <surname>Mayfield</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Penstein-Rosé</surname>
          </string-name>
          .
          <article-title>Using feature construction to avoid large feature spaces in text classification</article-title>
          .
          <source>In Proceedings of the 12th annual conference on Genetic and evolutionary computation, page 1299-1306</source>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakob</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García-Silva</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>DBpedia spotlight: shedding light on the web of documents</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Semantic Systems, page 1-8</source>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>P.</given-names>
            <surname>Refaeilzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <article-title>Cross-validation</article-title>
          . In L. LIU and M. T. ÖZSU, editors,
          <source>Encyclopedia of Database Systems</source>
          , pages
          <fpage>532</fpage>
          -
          <lpage>538</lpage>
          . Springer US,
          <year>Jan</year>
          .
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G.</given-names>
            <surname>Siemens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gasevic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Haythornthwaite</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dawson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Shum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ferguson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Duval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verbert</surname>
          </string-name>
          , and R. S. d
          <string-name>
            <surname>Baker</surname>
          </string-name>
          .
          <article-title>Open learning analytics: an integrated &amp; modularized platform. Proposal to design, implement and evaluate an open platform to integrate heterogeneous learning analytics techniques</article-title>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Klein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer</surname>
          </string-name>
          .
          <article-title>Feature-rich part-of-speech tagging with a cyclic dependency network</article-title>
          .
          <source>In Proceedings of the HLT-NAACL 2003, page 252-259</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>O.</given-names>
            <surname>Tsur</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Davidov</surname>
          </string-name>
          .
          <article-title>Icwsm - a great catchy name: Semisupervised recognition of sarcastic sentences in product reviews</article-title>
          .
          <source>In Proceeding of the International AAAI Conference on Weblogs and Social Media</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          , E. Frank, and
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hall</surname>
          </string-name>
          .
          <source>Data Mining: Practical Machine Learning Tools and Techniques</source>
          ,
          <string-name>
            <given-names>Third</given-names>
            <surname>Edition</surname>
          </string-name>
          . Morgan Kaufmann, 3 edition, Jan.
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>