<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Short Answer Grading: A Simple Solution for a Difficult Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefano Meniniy</string-name>
          <email>menini@fbk.eu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Tonelliy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni De Gasperisz</string-name>
          <email>giovanni.degasperis@univaq.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierpaolo Vittoriniz</string-name>
          <email>pierpaolo.vittorini@univaq.it</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>1070</fpage>
      <lpage>1075</lpage>
      <abstract>
        <p>English. The task of short answer grading is aimed at assessing the outcome of an exam by automatically analysing students' answers in natural language and deciding whether they should pass or fail the exam. In this paper, we tackle this task training an SVM classifier on real data taken from a University statistics exam, showing that simple concatenated sentence embeddings used as features yield results around 0.90 F1, and that adding more complex distance-based features lead only to a slight improvement. We also release the dataset, that to our knowledge is the first freely available dataset of this kind in Italian.1</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Human grading of open ended questions is a
tedious and error-prone task, a problem that has
become particularly pressing when such an
assessment involves a large number of students, like in
an Academic setting. One possible solution to this
problem is to automate the grading process, so that
it can facilitate teachers in the correction and
enable students to receive immediate feedback.
Research on this task has been active since the ’60s
        <xref ref-type="bibr" rid="ref23">(Page, 1966)</xref>
        , and several computational methods
have been proposed to automatically grade
different types of texts, from longer essays to short text
answers. The advantages of this kind of automatic
assessment do not concern only the limited time
and effort required to grade tests compared with a
manual assessment, but include also the reduction
of mistakes and bias introduced by humans, as well
as a better formalization of assessment criteria.
      </p>
      <p>In this paper, we focus on tests comprising short
answers to natural language questions, proposing
1Copyright ©2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
a novel approach to binary automatic short
answer grading (ASAG). This has proven particularly
challenging because an understanding of natural
language is required, without having much textual
context, while grading multiple-choice questions
can be straightforwardly assessed, given that there
is only one possible correct response to each
question. Furthermore, the tests considered in this
paper are taken from real exams on statistical
analyses, with low variability, a limited vocabulary and
therefore little lexical difference between correct
and wrong answers.</p>
      <p>The contribution of this paper is two-fold: we
create and release a dataset for short-answer
grading containing real examples, which can be freely
downloaded at https://zenodo.org/record/
3257363#.XRsrn5P7TLY. Besides, we propose a
simple approach that, making use only of
concatenated sentence embeddings and an SVM classifier,
achieves up to 0.90 F1 after parameter tuning.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        In the literature, several works have been presented
on automated grading methods, to assess the
quality of answers in written examinations. Several
types of answers have been addressed, from
essays
        <xref ref-type="bibr" rid="ref12 ref28">(Kanejiya et al., 2003; Shermis et al., 2010)</xref>
        ,
to code
        <xref ref-type="bibr" rid="ref30">(Souza et al., 2016)</xref>
        . Here we focus on
works related to short answers, which are the
target of our tests. With short answers we refer to
open questions, given in natural language, usually
with the length of one paragraph, recalling external
knowledge
        <xref ref-type="bibr" rid="ref6">(Burrows et al., 2015)</xref>
        . When
assessing the grading of short answers we face two main
issues, i) the grading itself and ii) the presence of
appropiate datasets.
      </p>
      <p>
        ASAG can be tackled with several approaches,
including pattern matching
        <xref ref-type="bibr" rid="ref21">(Mitchell et al., 2002)</xref>
        ,
looking for specific concepts or keywords in the
answers
        <xref ref-type="bibr" rid="ref11 ref12 ref17 ref7">(Callear et al., 2001; Leacock and Chodorow,
2003; Jordan and Mitchell, 2009)</xref>
        , using bag of
words and matching terms
        <xref ref-type="bibr" rid="ref8">(Cutrone et al., 2011)</xref>
        or relying on LSA
        <xref ref-type="bibr" rid="ref14 ref8">(Klein et al., 2011)</xref>
        . Some other
solutions rely more heavily on NLP techniques, for
example by extracting metrics and features that can
be used for text classification such as the overlap
of n-grams or POS between student’s and teacher’s
answers
        <xref ref-type="bibr" rid="ref19 ref29 ref3 ref8">(Bailey and Meurers, 2008; Meurers et al.,
2011)</xref>
        . Some attempts have been made also to use
similarity between word embeddings as a feature
        <xref ref-type="bibr" rid="ref15 ref26 ref31">(Sultan et al., 2016; Sakaguchi et al., 2015; Kumar
et al., 2017)</xref>
        .
      </p>
      <p>
        Another aspect that can affect the performance
of different ASAG approaches is the target of
automated evaluation. We can for instance assess the
quality of the text
        <xref ref-type="bibr" rid="ref8">(Yannakoudakis et al., 2011)</xref>
        ,
its comprehension and summarization
        <xref ref-type="bibr" rid="ref18">(Madnani et
al., 2013)</xref>
        , or, as in our case, the knowledge of a
specific notion. Each task would therefore need a
specific dataset as a benchmark. Other dimensions
affecting the approach to ASAG and its performance
are also the school level for which an assessment
is required (e.g primary school vs. university) as
well as its domain, e.g. computer science (Gütl,
2007), biology
        <xref ref-type="bibr" rid="ref29 ref3">(Siddiqi and Harrison, 2008)</xref>
        or
math
        <xref ref-type="bibr" rid="ref12 ref17">(Leacock and Chodorow, 2003)</xref>
        . As for
Italian, we are not aware of existing automated grading
approaches, nor of available datasets specifically
released to foster research in this direction. These
are indeed the main contributions of the current
paper.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Task and Data Description</title>
      <p>
        The short grading task that we analyse in this paper
is meant to automatize part of the exam that
students of Health Informatics in the degree course of
Medicine and Surgery of the University of L’Aquila
(Italy) are required to pass. It includes two
activities: a statistical analysis in R and the explanation
of the results in terms of clinical findings. While
the evaluation of the first part has already been
automatized through automated grading of R code
snippets
        <xref ref-type="bibr" rid="ref1">(Angelone and Vittorini, 2019)</xref>
        , the
second task had been addressed by the same authors
using a string similarity approach, which however
did not yield satisfying results. Indeed, they used
Levenshtein distance to compute the distance
between the students’ answer and a gold standard
(i.e. correct) answer, but the approach failed to
capture the semantic equivalence between the two
sentences, while focusing only on the lexical one.
      </p>
      <p>For example, an exam provided students with
data about surgical operations, subjects, scar
visibility and hospital stay, and asked to compute
several statistical measures in R, such as the absolute
and relative frequencies of the surgical operations.
Then, students were required to comment in plain
text on some of the analyses, for example state
whether some data are extracted from a normal
distribution. For this second part of the exam, the
teacher prepared a “gold answer”, i.e. the correct
answer. Two real examples from the dataset are
reported below.</p>
      <p>Correct answer pair:
(Student) Poiché il p-value e maggiore di
0.05 in entrambi i casi, la distribuzione
è normale, procediamo con un test
parametrico per variabili appaiate.
(Gold) Siccome tutti i test di normalità
presentano un p&gt;0.05, posso utilizzare
un test parametrico.</p>
      <p>Wrong answer pair:
(Student) Siccome p&lt;0.05,la differenza
fra le due variabili è statisticamente
significativa.
(Gold) Siccome il t-test restituisce un
pvalue &gt; di 0.05, non posso
generalizzare alla popolazione il risultato
osservato nel mio campione, e quindi non c’è
differenza media di peso statisticamente
significativa fra i figli maschi e femmine.</p>
      <p>The goal of our task is, given each pair, to train
a classifier and label correct and wrong students’
answers. An important aspect of our task is that
the correctness of an answer is not defined with
respect to the question, which is not used for
classification. For the moment we also focus on binary
classification, to determine whether an answer is
correct or not, without providing a numeric score
on how much it is correct or wrong. With the data
organized into student-professor answers pairs, the
classification is done considering i) the semantic
content of the answers (represented through word
embeddings ii) features related to the pair
structure of the data such as the overlap or the distance
between the two texts. The adopted features are
explained in detail in Section 4.1.
3.1</p>
      <sec id="sec-3-1">
        <title>Dataset</title>
        <p>The dataset available at https://zenodo.org/
record/3257363#.XR5i8ZP7TLY has been
partially collected using data from real statistics exams
spanning different years, and partially extended by
the authors of this paper. The dataset contains the
list of sentences written by students, with a unique
sentence ID, the type of statistical analysis it refers
to (if either given for the hypothesis or normality
test), its degree in a range from 0 to 1, and its
fail/pass result, flanked with a manually defined gold
standard (i.e. the correct answer). The degree is a
numerical score manually assigned to each answer,
which takes into account whether an answer is
partially correct, mostly correct or completely wrong.
Based on this degree, the pass/fail decision was
taken, i.e. if degree &lt; 0:6 then fail, otherwise
pass.</p>
        <p>In order to increase the number of training
instances and achieve a better balance between the
two classes, we manually negated a set of correct
answers and reversed the corresponding fail/pass
result, adding a set of negated gold standard
sentences for a total of 332 new pairs. We also
manually paraphrased 297 of the original gold standard
sentences, so that we created some additional pairs.
Overall the dataset consists of 1,069 student/gold
standard answer pairs, 663 of which are labeled as
“pass” and 406 as “fail”.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Classification framework</title>
      <p>
        Although several works have explored the
possibility to automatically grade short text answers, these
attempts have mainly focused on English.
Furthermore, the best performing ones strongly rely on
knowledge bases and syntactic analyses
        <xref ref-type="bibr" rid="ref22 ref8">(Mohler et
al., 2011)</xref>
        , which are hard to obtain for Italian. We
therefore test for the first time the potential of
sentence embeddings to capture pass or fail judgments
in a supervised setting, where the only required
data are a) a training/test set and b) sentence
embeddings
        <xref ref-type="bibr" rid="ref4">(Bojanowski et al., 2017)</xref>
        trained using
fastText2.
4.1
      </p>
      <sec id="sec-4-1">
        <title>Method</title>
        <p>Since we cast the task in a supervised classification
framework, we first need to represent the pairs of
student/gold standard sentences as features. Two
different types of features are tested:
distancebased features, which capture the similarity of
the two sentences using measures based on lexical
and semantic similarity, and sentence embeddings
features, whose goal is to represent the semantics
of the two sentences in a distributional space.
2https://fasttext.cc/</p>
        <p>
          All sentences are first preprocessed by
removing the stopwords such as articles and prepositions,
and by replacing mathematical notations with their
transcription in plain language, e.g. “&gt;" with
“maggiore di" (greater than). We also perform
part of speech tagging, lemmatisation and affix
recognition using the TINT NLP Suite for Italian
          <xref ref-type="bibr" rid="ref13 ref2">(Aprosio and Moretti, 2018)</xref>
          . Then on each pair
of sentences the following distance-based features
are computed:
• Token overlap: a feature representing the
number of overlapping tokens between the
two sentences normalised by their length.
This feature captures the lexical similarity
between the two strings.
• Lemma overlap: a feature representing the
number of overlapping lemmas between the
two sentences normalised by their length.
Like the previous one, this feature captures
the lexical similarity between the two strings.
• Presence of negations: this feature represents
whether a content word is negated in one
sentence and not in the other. For each sentence,
negations are recognised based on the NEG
PoS tag or the affix ‘a-’ or ‘in-’ (e.g.
indipendente), and then the first content word
occurring after the negation is considered. We
extract two features, one for each sentence,
and the values are normalised by their length.
        </p>
        <p>
          Other distance-based features are computed at
sentence level, and to this purpose we employ
fastText
          <xref ref-type="bibr" rid="ref4">(Bojanowski et al., 2017)</xref>
          , an extension
of word embeddings
          <xref ref-type="bibr" rid="ref20 ref24">(Mikolov et al., 2013;
Pennington et al., 2014)</xref>
          developed at Facebook that is
able to deal with rare words by including subword
information, and representing sentences basically
by combining vectors representing both words and
subwords. To generate these embeddings we start
from the pre-computed Italian language model3
trained on Common Crawl and Wikipedia. The
latter, in particular, is suitable for our domain, since it
includes also scientific content and statistics pages,
therefore the language of the exam should be well
represented in our model. The embeddings are
created using continuous bag-of-word with
positionweights, a dimension of 300, character n-grams
of length 5, a window of size 5 and 10 negatives.
        </p>
        <p>
          3https://fasttext.cc/docs/en/crawl-vectors.
html
Then, the embedding of the sentences written by
the students and the gold standard ones are created
by combining the word and the subword
embeddings with the fastText library. Each sentence is
therefore represented through a 300 dimensional
embedding. Based on this, we extract four
additional distance-based features:
• Embeddings cosine: the cosine between the
two sentence embeddings is computed. The
intuition behind this feature is that the
embeddings of two sentences with a similar meaning
would be close in a multidimensional space
• Embeddings cosine (lemmatized): the same
feature as the previous one, with the only
difference that the sentences are first lemmatised
before creating the embeddings
• Word Mover’s Distance (WMD): WMD is a
similarity measures based on the minimum
amount of distance that the embedded words
of one document need to move to reach the
embedded words of another document
          <xref ref-type="bibr" rid="ref16">(Kusner et al., 2015)</xref>
          in a multidimensional space.
Compared with other existing similarity
measures, it works well also when two sentences
have a similar meaning despite having few
words in common. We apply this algorithm
to measure the distance between the solutions
proposed by the students and the ones in the
gold standard.
• Word Mover’s Distance (lemmatized): the
same feature as the previous one, with the only
difference that the sentences are first
lemmatised before creating the embeddings
        </p>
        <p>
          The sentence embeddings used to compute the
distance features are also tested as features in
isolation: a 600 dimensional vector is indeed created by
concatenating each sentence embeddings
composing a student answer – gold standard pair. This
representation is then directly fed to the classifier. We
adopt this solution inspired by recent approaches to
natural language inference using the concatenation
of premise and hypothesis
          <xref ref-type="bibr" rid="ref13 ref2 ref5">(Bowman et al., 2015;
Kiros and Chan, 2018)</xref>
          .
        </p>
        <p>
          As for the supervised classifier, we use support
vector machines
          <xref ref-type="bibr" rid="ref27 ref7">(Scholkopf and Smola, 2001)</xref>
          ,
which generally yield satisfying results in
classification tasks with a limited number of training
instances (as opposed to deep learning approaches).
We then proceeded to find the best C and
parameters by means of grid-search tuning
          <xref ref-type="bibr" rid="ref10">(Hsu et
al., 2016)</xref>
          , through a 10-fold cross-validation to
prevent to overfit the model. Finally, with the
parameters that returned the best performance, we
finalised the classifier and calculated its accuracy
and F1 score. The analyses were performed
using R 3.6.0 with caret v6.0-84 and e1071 v1.7-2
packages
          <xref ref-type="bibr" rid="ref25">(R Core Team, 2018)</xref>
          .
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Results</title>
        <p>Figure 1 shows the plot summarising the tuning
process. In summary, within the explored area, the
best parameters were found to be C = 104 and
= 2 6. The resulting tuned model produced the
following results:
• Accuracy = 0:891 (balanced accuracy =
0:876);
• F1 score = 0:914;</p>
        <p>With a similar approach, we also tuned the
classifier when fed with only the concatenated sentence
embeddings as features (i.e., without
distancebased features). With best parameters C = 103
and = 2 3, the results were:
• Accuracy = 0:885 (balanced accuracy =
0:870);
• F1 score = 0:909;</p>
        <p>To evaluate the quality of the model learned with
these two configurations, and make sure that it
does not overfit, we perform an additional test:
we collect a small set of students’ answers from a
different statistics exam than the one used to create
the training set. This is done on novel data by
collecting students’ answers from a small number
of new questions, and manually creating new gold
answers to be used in the pairs. Overall, we obtain
77 new answer pairs, consisting of 14 wrong and 63
correct answers. We then run the best performing
model with all features and using only sentence
embeddings (same C and as before). The results
are the following:
• Accuracy using all features = 0:7838
(balanced accuracy = 0:5965);
• F1 score 0:8710;
while the results achieved using only sentence
embeddings are:
• Accuracy = 0:7973 (balanced accuracy =
0:6349);
• F1 score = 0:8780;
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>The results presented in the previous section show
only a small increase in performance when using
the distance-based features in addition to the
sentence embeddings after tuning both configurations.
This outcome highlights the effectiveness of
using sentence embeddings to represent the semantic
content of the answers in tasks where student’s and
gold solutions are very similar to each other. In
fact, the sentence pairs in our dataset show a high
level of word overlap, and the only discriminant
between a correct and a wrong answer is
sometimes only the presence of “&lt;” instead of “&gt;”, or
a negation.</p>
      <p>The second experiment, where the same
configuration is run on a test set taken from a statistics
exam on different topics, shows an overall decrease
in performance as expected, but the classification
accuracy is still well above the most frequent
baseline. In this setting, using only the sentence
embeddings yields a slightly better performance than
including the other features, showing that they are
more robust with respect to a change of topic.</p>
      <p>In general terms, despite the accurate
parameter tuning, the classification approach seems to
be applicable to short answer grading tests
different from the data on which the training was done,
provided that the student’s and gold answer types
are the same as in our dataset (i.e. limited length,
limited lexical variability).
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this paper, we have presented a novel dataset
for short answer grading taken from a real
statistics exam, which we make freely available. To our
knowledge, this is the first dataset of this kind. We
also introduce a simple approach based on
sentence embeddings to automatically identify which
answers are correct or not, which is easy to
replicate and not computationally intensive.</p>
      <p>In the future, the work could be extended in
several directions. First of all, it would be interesting
to use deep-learning approaches instead of SVM,
but for that more training data are needed. These
could be collected in the upcoming exam sessions
at University of L’Aquila. Another refinement of
this work would be to grade the tests by assigning
a numerical score instead of a pass/fail judgment.
Since such scores are already included in the
released dataset (the degrees), this would be quite
straightforward to achieve. Finally, we plan to test
the classifier by integrating it in an online
evaluation tool, through which students can submit their
tests and the trainer can run an automatic pass/fail
assignment.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Anna</given-names>
            <surname>Maria</surname>
          </string-name>
          Angelone and
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Vittorini</surname>
          </string-name>
          .
          <year>2019</year>
          . The Automated Grading of R Code Snippets:
          <article-title>Preliminary Results in a Course of Health Informatics</article-title>
          .
          <source>In Proc. of the 9th International Conference in Methodologies and Intelligent Systems for Technology Enhanced Learning</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Alessio</given-names>
            <surname>Palmero</surname>
          </string-name>
          Aprosio and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Moretti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Tint 2.0: an all-inclusive suite for NLP in italian</article-title>
          .
          <source>In Proceedings of the Fifth Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2018</year>
          ), Torino, Italy,
          <source>December 10-12</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Stacey</given-names>
            <surname>Bailey</surname>
          </string-name>
          and
          <string-name>
            <given-names>Detmar</given-names>
            <surname>Meurers</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Diagnosing meaning errors in short answers to reading comprehension questions</article-title>
          .
          <source>In Proceedings of the Third Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , pages
          <fpage>107</fpage>
          -
          <lpage>115</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>5</volume>
          :
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Samuel R. Bowman</surname>
            , Gabor Angeli, Christopher Potts, and
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A large annotated corpus for learning natural language inference</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>632</fpage>
          -
          <lpage>642</lpage>
          , Lisbon, Portugal, September. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Steven</given-names>
            <surname>Burrows</surname>
          </string-name>
          , Iryna Gurevych, and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The eras and trends of automatic short answer grading</article-title>
          .
          <source>International Journal of Artificial Intelligence in Education</source>
          ,
          <volume>25</volume>
          (
          <issue>1</issue>
          ):
          <fpage>60</fpage>
          -
          <lpage>117</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>David H Callear</surname>
          </string-name>
          ,
          <article-title>Jenny Jerrams-Smith, and</article-title>
          <string-name>
            <given-names>Victor</given-names>
            <surname>Soh</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Caa of short non-mcq answers</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Laurie</given-names>
            <surname>Cutrone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Maiga</given-names>
            <surname>Chang</surname>
          </string-name>
          , et al.
          <year>2011</year>
          .
          <article-title>Autoassessor: Computerized assessment system for marking student's short-answers automatically</article-title>
          .
          <source>In 2011 IEEE International Conference on Technology for Education</source>
          , pages
          <fpage>81</fpage>
          -
          <lpage>88</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Christian</given-names>
            <surname>Gütl</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>e-examiner: towards a fullyautomatic knowledge assessment tool applicable in adaptive e-learning systems</article-title>
          .
          <source>In Proceedings of the 2nd international conference on interactive mobile and computer aided learning</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . Citeseer.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Chih-Wei</surname>
            <given-names>Hsu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chih-Chung Chang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Chih-Jen Lin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A Practical Guide to Support Vector Classification</article-title>
          .
          <source>Technical report</source>
          , National Taiwan University.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Sally</given-names>
            <surname>Jordan</surname>
          </string-name>
          and Tom Mitchell.
          <year>2009</year>
          .
          <article-title>e-assessment for learning? the potential of short-answer free-text questions with tailored feedback</article-title>
          .
          <source>British Journal of Educational Technology</source>
          ,
          <volume>40</volume>
          (
          <issue>2</issue>
          ):
          <fpage>371</fpage>
          -
          <lpage>385</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Dharmendra</given-names>
            <surname>Kanejiya</surname>
          </string-name>
          , Arun Kumar, and
          <string-name>
            <given-names>Surendra</given-names>
            <surname>Prasad</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Automatic evaluation of students' answers using syntactically enhanced lsa</article-title>
          .
          <source>In Proceedings of the HLT-NAACL 03 workshop on Building educational applications using natural language processing-</source>
          Volume
          <volume>2</volume>
          , pages
          <fpage>53</fpage>
          -
          <lpage>60</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Jamie</given-names>
            <surname>Kiros</surname>
          </string-name>
          and
          <string-name>
            <given-names>William</given-names>
            <surname>Chan</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Inferlite: Simple universal sentence representations from natural language inference data</article-title>
          .
          <source>In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source>
          , Brussels, Belgium,
          <source>October 31 - November 4</source>
          ,
          <year>2018</year>
          , pages
          <fpage>4868</fpage>
          -
          <lpage>4874</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Richard</given-names>
            <surname>Klein</surname>
          </string-name>
          , Angelo Kyrilov, and
          <string-name>
            <given-names>Mayya</given-names>
            <surname>Tokman</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Automated assessment of short free-text responses in computer science using latent semantic analysis</article-title>
          .
          <source>In Proceedings of the 16th annual joint conference on Innovation and technology in computer science education</source>
          , pages
          <fpage>158</fpage>
          -
          <lpage>162</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Sachin</given-names>
            <surname>Kumar</surname>
          </string-name>
          , Soumen Chakrabarti, and
          <string-name>
            <given-names>Shourya</given-names>
            <surname>Roy</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Earth mover's distance pooling over siamese lstms for automatic short answer grading</article-title>
          .
          <source>In IJCAI</source>
          , pages
          <fpage>2046</fpage>
          -
          <lpage>2052</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Matt</given-names>
            <surname>Kusner</surname>
          </string-name>
          , Yu Sun, Nicholas Kolkin,
          <string-name>
            <given-names>and Kilian</given-names>
            <surname>Weinberger</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>From word embeddings to document distances</article-title>
          .
          <source>In International Conference on Machine Learning</source>
          , pages
          <fpage>957</fpage>
          -
          <lpage>966</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Claudia</given-names>
            <surname>Leacock</surname>
          </string-name>
          and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Chodorow</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>C-rater: Automated scoring of short-answer questions</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>37</volume>
          (
          <issue>4</issue>
          ):
          <fpage>389</fpage>
          -
          <lpage>405</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Nitin</given-names>
            <surname>Madnani</surname>
          </string-name>
          , Jill Burstein, John Sabatini, and
          <string-name>
            <surname>Tenaha O'Reilly</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Automated scoring of a summarywriting task designed to measure reading comprehension</article-title>
          .
          <source>In Proceedings of the Eighth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , pages
          <fpage>163</fpage>
          -
          <lpage>168</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Detmar</given-names>
            <surname>Meurers</surname>
          </string-name>
          , Ramon Ziai, Niels Ott, and
          <string-name>
            <surname>Stacey M Bailey</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Integrating parallel analysis modules to evaluate the meaning of answers to reading comprehension questions</article-title>
          .
          <source>International Journal of Continuing Engineering Education and Life-Long Learning</source>
          ,
          <volume>21</volume>
          (
          <issue>4</issue>
          ):
          <fpage>355</fpage>
          -
          <lpage>369</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Tom</given-names>
            <surname>Mitchell</surname>
          </string-name>
          , Terry Russell,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Broomhead</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Aldridge</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Towards robust computerised marking of free-text responses</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Mohler</surname>
          </string-name>
          , Razvan Bunescu, and
          <string-name>
            <given-names>Rada</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Learning to grade short answer questions using semantic similarity measures and dependency graph alignments</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1, HLT '11</source>
          , pages
          <fpage>752</fpage>
          -
          <lpage>762</lpage>
          , Stroudsburg, PA, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Ellis B Page</surname>
          </string-name>
          .
          <year>1966</year>
          .
          <article-title>The imminence of grading essays by computer</article-title>
          .
          <source>The Phi Delta Kappan</source>
          ,
          <volume>47</volume>
          (
          <issue>5</issue>
          ):
          <fpage>238</fpage>
          -
          <lpage>243</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Proceedings of EMNLP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>R Core</given-names>
            <surname>Team</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>R: A Language and Environment for Statistical Computing</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Keisuke</given-names>
            <surname>Sakaguchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Heilman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Nitin</given-names>
            <surname>Madnani</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Effective feature integration for automated short answer scoring</article-title>
          .
          <source>In Proceedings of the 2015</source>
          conference
          <article-title>of the North American Chapter of the association for computational linguistics: Human language technologies</article-title>
          , pages
          <fpage>1049</fpage>
          -
          <lpage>1054</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Scholkopf and Alexander J Smola</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Learning with kernels: support vector machines, regularization, optimization, and beyond</article-title>
          . MIT press.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Mark D Shermis</surname>
            , Jill Burstein, Derrick Higgins, and
            <given-names>Klaus</given-names>
          </string-name>
          <string-name>
            <surname>Zechner</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Automated essay scoring: Writing assessment and instruction</article-title>
          .
          <source>International encyclopedia of education</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ):
          <fpage>20</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Raheel</given-names>
            <surname>Siddiqi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Harrison</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>A systematic approach to the automated marking of short-answer questions</article-title>
          .
          <source>In 2008 IEEE International Multitopic Conference</source>
          , pages
          <fpage>329</fpage>
          -
          <lpage>332</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Draylson M Souza</surname>
          </string-name>
          ,
          <string-name>
            <surname>Katia R Felizardo</surname>
          </string-name>
          , and Ellen F Barbosa.
          <year>2016</year>
          .
          <article-title>A systematic literature review of assessment tools for programming assignments</article-title>
          .
          <source>In 2016 IEEE 29th International Conference on Software Engineering Education and Training (CSEET)</source>
          , pages
          <fpage>147</fpage>
          -
          <lpage>156</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <given-names>Md</given-names>
            <surname>Arafat</surname>
          </string-name>
          <string-name>
            <surname>Sultan</surname>
          </string-name>
          , Cristobal Salazar, and
          <string-name>
            <given-names>Tamara</given-names>
            <surname>Sumner</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Fast and easy short answer grading with Helen Yannakoudakis, Ted Briscoe</article-title>
          , and
          <string-name>
            <given-names>Ben</given-names>
            <surname>Medlock</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A new dataset and method for automatically grading esol texts</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume</source>
          <volume>1</volume>
          , pages
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>