<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>I. Jones, M. Swan, and A. Pollitt. Assessing mathematical problem solving using comparative judge-
ment. International Journal of Science and Mathematics Education</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>High Stakes Automatic Assessments: Developing an Online Linear Algebra Examination</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Christopher J. Sangwin School of Mathematics University of Edinburgh</institution>
          ,
          <addr-line>Edinburgh</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Copyright c by the paper's authors. Copying permitted for private and academic purposes. In: O. Hasan, W. Neuper, Z. Kovcs, W. Schreiner (eds.): Proceedings of the Workshop CME-EI: Computer Mathematics in Education - Enlightenment or Incantation</institution>
          ,
          <addr-line>Hagenberg, Austria, 17-Aug-2018, published at</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>13</volume>
      <issue>1</issue>
      <abstract>
        <p>In this paper I investigate development of an automatically marked online version of a current paper-based examination for a university mathematics course, and the extent to which the outcomes are equivalent to a paper-based exam. An online examination was implemented using the STACK online assessment tool which is built using computer algebra, and in which students' answers are normally typed expressions. The study group was 376 undergraduates taking a year 1 Introduction to Linear Algebra course. The results of this experiment are cautiously optimistic: a signi cant proportion of current examination questions can be automatically assessed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Automation also has practical bene ts, reducing the marking load and potentially speeding up examination
processes. However, changing written examinations, with centuries of custom and practice, is a high-stakes and
high-risk undertaking. My previous joint work [SK16] examined questions set in school-level examination papers
with a view to developing automatically marked online versions. The results of [SK16] were cautiously optimistic
that a signi cant proportion of current questions could be automatically assessed. In this paper I extend this
work, and create examination questions and trial their use with a large group of university students.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>For this study I added a mock online examination to Introduction to Linear Algebra (ILA). This is a year 1,
semester 1, mathematics course worth 20 credits taken by mathematics, computer science and other
undergraduate students. Students normally take 120 credits per year, in two semesters. The course is de ned by [Poo11]
Chapters 1 to Chapter 6.2, with a selection of the applications included and selected topics omitted. ILA had
over 600 students, of whom 578 took the nal written examination and had a non-zero examination mark.</p>
      <p>Students had requested exam practice, but it was impractical to administer and mark students' attempts
(approximately 35 person-days for the genuine exam) in the short period between the end of teaching and
the scheduled examination. In context, a mock examination was likely to be taken seriously by a signi cant
proportion of the student cohort as a valuable practice and learning opportunity. Since the mock examination
did not contribute to the overall course grade there was no incentive for students to cheat, or to be impersonated.
Introduction to Linear Algebra, has an \open book" examination and so possible access to materials is less of a
threat to this experiment than would be the case for a closed-book examination. The lack of certainty over who
was sitting the online tests, the circumstances of participation, the potential use of internet resources and so on
is certainly a compromise. Such uncertainty does not a ect the extent to which I could produce questions at a
technical level, or the e ectiveness of the scoring mechanism in the face of students' attempts.</p>
      <p>The results consist of a report on the extent to which current questions can be faithfully automated, and I
give a preliminary report on students' attempts.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>The existing paper-based ILA examination takes 180 minutes and consists of Section A: compulsory questions
totalling 40 marks, and Section B: four questions each of 20 marks from which we take the student's best three
marks. Students may use any standard scienti c calculator but graphical calculators with matrix functions are
not permitted.</p>
      <p>
        The primary goal was to provide students with an online examination which was as close as possible to the
forthcoming paper-based summative course examination. ILA has been running for many years, with a stable
(but not invariant) syllabus, and I had access to examinations going back to
        <xref ref-type="bibr" rid="ref4">December 2011</xref>
        (two per year: the
main exam and an equivalent resit paper). I therefore decided to remove the oldest exam papers from easy access
      </p>
      <sec id="sec-3-1">
        <title>Online exam</title>
      </sec>
      <sec id="sec-3-2">
        <title>Study group paper Paper examination</title>
        <p>0
y 6
c
enu 04
q
reF 02
cy 60
n
e
u
q
reF 02
y
c
n
e
u
q
e
r
F
0
8
0
4
0
0
0
through the course website and base the online examination on those questions. Using as few papers as possible
helps provide a representative online examination. Technically it is di cult to operate a \best 3 out of 4" mark
scheme in the STACK online system and in any case for a formative mock exam this makes little sense.</p>
        <p>In deciding how to allocate marks I have taken a strict interpretation. Speci cally, where the original intention
of the examiners included \with justi cation", I only awarded a minimum number of marks for giving the answer
only online. For example, Q5 on our online exam asked the following.</p>
        <sec id="sec-3-2-1">
          <title>5. Is it possible for A and B to be 3</title>
          <p>3 rank 2 matrices with AB = 0? True/False.</p>
          <p>The original paper awarded 7 marks for the answer and justi cation, whereas only one mark was awarded for
the correct answer. I did ask students to provide typed free-text justi cations even though these would not be
marked and no feedback was provided.</p>
          <p>Ultimately I used two papers (120 marks each) to create the online exam with 59 marks of the online exam
coming from Dec-11 and 50 marks from Aug-12. I took one question from Dec-13 to add a mark to Section A
to make the online exam total 110 marks. Of the paper-based questions selected for the online exam, 44 marks
are not awarded online. These missing marks are for justi cation which cannot, at this time, be automatically
assessed. This resulted in Section A having fewer marks than would be the case with a paper based submission.
Of the 240 marks available on the Dec-11 and Aug-12 papers, 109/240 marks 45% were automated in a way
faithful to the original examinations. However, the online versions do lack some partial credit and do not (in
this experiment) implement follow on-marking, which in some Section B questions is substantial.</p>
          <p>An example question is shown in Figure 1, illustrating \validity" feedback which was available during the
exam. Validity feedback is normally available to a student, and provides information on syntax errors and other
input problems helping reduce the extent to which students are penalized on a technicality. For ILA, online
course work quizzes were already implemented using STACK. All students were expected to sit 30 online quizzes
using the STACK system as part of the ILA course before the mock examination, and would be thoroughly
familiar with how to enter answers into the system.</p>
          <p>
            The online examination was made available to students to do in their own time for a period of one week in
De
            <xref ref-type="bibr" rid="ref6">cember 2017</xref>
            , between the end of formal teaching and the scheduled paper-based exam. Students could choose
when to sit the online examination, but were given one attempt of 180 minutes to do so to simulate examination
practice. All data was downloaded from the online STACK system, and after rati cation by the exam board,
combined with overall achievement data. Students were assigned a unique number to ensure anonymity, and the
data loaded into R-studio for analysis.
          </p>
          <p>
            There were 395 attempts at the mo
            <xref ref-type="bibr" rid="ref6">ck online exam in December 2017</xref>
            . One student who was granted a second
attempt for technical reasons had their rst attempt disregarded, giving 394 attempts. There were no other
signi cant technical problems a ecting the conduct of the online examination. For the online exam (including
those who scored zero) the mean grade was 47:9% with standard deviation of 23:2%. The coe cient of internal
consistency (Cronbach Alpha) for the online exam was 0:87. There was a moderate positive correlation between
time taken (M=132 mins, SD=48.6 mins) and the online exam result (M=47.9%, SD=23.2) r(392) = 0:517,
p &lt; 10 16, as might be expected. Despite a small number of outlier questions, the mock online exam appears to
have operated successfully in its own right as a test.
          </p>
          <p>e
n
i
l
n
O
0
0
1
0
8
0
6
0
4
0
2
0
20
40
80</p>
          <p>100
60</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Paper Figure 3: Online mock exam grades vs paper exam grades for the study group Q4 Q6</title>
          <p>The nal mark for ILA is made up of coursework (20%) and a nal paper-based exam (80%). There were 394
attempts at the online mock examination, and all but one of these students also sat the paper-based examination.
Note that 17 students scored 0 for the online exam, perhaps indicating students who looked at the online questions
but made no serious attempt at them. Technically there is a di erence between students who never sat the online
exam, and those who opened the exam and scored 0. For the analysis I excluded the 17 students who scored 0
in the online exam: this leaves the study group of N = 376 students with paper and mock exam information.</p>
          <p>For the study group, the online exam results had (M=50.2, SD=21.3) and paper exam (M=68.0, SD=17.3).
For all students who sat the ILA paper exam (M=63.1, SD=21.6). Histograms of achievement in the online
mock and paper based examinations are shown in Figure 2. There is a signi cantly larger failure rate (score less
than 40%) in the online examination, and a signi cantly lower mean. These di erences could be explained by
the level of engagement: the online exam carried no credit, and students may have lost motivation when tired.</p>
          <p>A scatter plot of the online mock exam grades vs paper exam grades is shown in Figure 3, together with a linear
regression model. The dashed line shows the (ideal) linear relationship in which the online mock examination
has identical outcomes with the paper-based exam. Notice the online exam scores are clearly below those of the
paper exam, supporting the hypothesis that students may have lost motivation when tired and not performed
to their full potential in the online mock exam. The mock exam grades and paper exam grades were moderately
correlated, r(374) = 0:593, p &lt; 10 15.</p>
          <p>The number of non-empty free text responses to each of the \justify" questions is given in Table 1, together
with the mean and standard deviation of the response length (in number of characters). It is clear reading through
the free-text responses that over 200 students took the exercise seriously, providing sensible (and often correct)
justi cations in good English. For the Section A questions in paper there were 59 marks available, whereas in
the STACK exam only 24 marks were awarded. I did not expect students to make serious use of the free-text
entry. The fact students entered sensible justi cation to many of these questions, and received no marks, could
easily account for the di erence in mean scores between the paper-based and online exam. There were a large
number of empty responses (as there are on paper as well), together with some incoherent utterances, and some
plaintive messages. I did not assess these free-text responses, or subject them to comprehensive analysis for the
purposes of this paper. However, in a genuine online examination such responses could be assessed (1) manually
in the traditional way on-screen, (2) using automatic assessment technology such as described in [BJ10, Jor12],
or (3) using comparative judgement for longer passages, see [JSP14, Pol12].
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>The implementation of the mock online examination for linear algebra was a modest success. There were no
serious technical problem during the conduct, and no students complained of inaccurate or unfair marking. The
results of the online examination were broadly comparable with a paper-based exam, with the consistently lower
online performance explained by a combination of (1) potential disengagement in a low stakes setting, (2) lack
of assessment of students' justi cation, (3) lack of partial credit and follow through marking. Both partial credit
and follow through marking are technically possible in STACK, but are expensive (in sta time) to implement.
The results give us con dence to use such assessments in higher-stakes settings in the future.</p>
      <p>This research has done nothing to address serious practical problems associated with online examinations.
Problems include the need for invigilation to reduce plagiarism and impersonation, and security to eliminate
communication during the exam (such as answer sharing) or access to unauthorised resources. These examination
conduct problems must be solved, but they have nothing to do with mathematics.</p>
      <p>Automatic assessment is an area of mathematics which would particularly bene t from tools which automate
explanation, justi cation and reasoning. In particular \proof checking" software, as applied to students'
understanding, is necessary to move beyond assessing only a nal answer, as shown in Figure 1, to a full mathematical
answer. In this study, only students' nal answers were subject to automatic assessment which is a serious
limitation. However, progress is being made to assess working especially in the area of reasoning by equivalence
as discussed brie y in [SK16].</p>
      <p>I was surprised at the large extent to which existing questions could be automatically assessed with the current
tools, based on computer algebra, faithfully. However, there is nothing sacrosanct about current examination
questions. Why should the online examination be exactly the same as a paper-based examination? Current
questions are written explicitly for the paper-based format, and it is sensible to seek to write questions which
take advantage of the online format as appropriate. Many Section A are true/false, but the justi cation of a
\false" response is via appropriate examples. Computer algebra is ideally suited to assessing answers, such as
counter examples, which expect the teacher to perform some time-consuming and error-prone calculation. For
this research I did not rephrase questions to \give me examples, such that ...", but this would be one option.</p>
      <p>This analysis raises the question of whether we, as a mathematics community, believe current mathematics
examinations are a valid test of mathematical achievement. Do current examinations actually represent valid
mathematical practice, as undertaken by researchers, industrial mathematicians and for pure recreation as an
intellectual pursuit? Construct validity is a central educational concern, but it is not relevant to the research
question of whether we can actually automate current exams. My personal views about the nature of mathematics
broadly align with those expressed in [Pol54] and [Lak76]. That is, that setting up abstract problems and solving
them lies at the the heart of mathematics. [Pol62] identi ed four patterns of thought to help structure thinking
about solving mathematical problems. His \Cartesian" pattern is where a problem is translated into a system
of equations, and solved using algebra. Note that the algebraic manipulation is the technical middle step in
the process: setting up the equations and interpreting the solutions are essential parts to complete this pattern.
My previous work [SK16] examined questions set in school-level examination papers and found that line-by-line
algebraic reasoning, termed reasoning by equivalence [NBC04], is the most important single form of reasoning in
school mathematics. However, many examination questions do not relate to a problem at all, rather they instruct
students to undertake a well-rehearsed set of techniques, isolated from any problem. Many of the questions in
the ILA examinations also rely on predictable methods which can be well-rehearsed. Predictable methods
predominate in school examinations, such as those considered in my previous research in [SK16]. Current
examinations tend towards \incantation" by students, and there is a real danger that national examination
boards, universities, and others with responsibilities for examinations will replicate traditional examinations
online without a critical reassessment of the purpose of mathematics education.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The increasing use of software tools in online assessment will a ect mathematics education. It is likely that
automatic online examinations for mathematics in school, as well as for some methods-based university courses,
will become feasible and will be used in the very near future. A pragmatic combination of computer algebra
supported assessment and automatic assessment of short answer questions will assess a signi cant proportion of
current questions automatically. Traditional expert marking, and use of comparative judgement will potentially
widen the scope of exams at the expense of complete automation. A further pragmatic approach will be to
split courses into two summative assessment components: largely skill-based questions can be automatically
assessed online, with the justi cation and rhetorical discussion in a traditional written examination. Replication
of traditional examinations online without a critical reassessment of the purpose of mathematics education would
be a wasted opportunity to de ne the subject through valid assessments.
5.0.1</p>
      <p>Acknowledgements
The online questions were created with the help of Dr Konstantina Zerva, of the University of Edinburgh.
[BJ10]
[Jor12]</p>
      <p>P. G. Butcher and S. E. Jordan. A comparison of human and computer marking of short free-text
student responses. Computers and Education, 55(2):489{499, September 2010.</p>
      <p>H. Burkhardt. What you test is what you get. In I. Wirszup and R. Streit, editors, The Dynamics of
Curriculum Change in Developments in School Mathematics Worldwide. University of Chicago School
Mathematics Project, 1987.</p>
      <p>S. Jordan. Student engagement with assessment and feedback: Some lessons from short-answer free-text
e-assessment questions. Computers and Education, 58(2):818{834, 2012.
[NBC04] J. F. Nicaud, D. Bouhineau, and H. Chaachoua. Mixing microworlds and CAS features in building
computer systems that help students learn algebra. International Journal of Computers for Mathematical
Learning, 9(2):169{211, 2004.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Bur87]
          <string-name>
            <given-names>G.</given-names>
            <surname>Polya</surname>
          </string-name>
          .
          <source>Mathematics and Plausible Reasoning</source>
          . Vol.
          <volume>1</volume>
          : Induction and Analogy in Mathematics. Vol
          <volume>2</volume>
          . Patterns of Plausible Inference. Princeton University Press,
          <year>1954</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>G.</given-names>
            <surname>Polya</surname>
          </string-name>
          .
          <article-title>Mathematical discovery: on understanding, learning, and teaching problem solving</article-title>
          . Wiley, London, UK,
          <year>1962</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Pollitt</surname>
          </string-name>
          .
          <article-title>The method of adaptive comparative judgement</article-title>
          .
          <source>Assessment in Education: Principles, Policy &amp; Practice</source>
          ,
          <volume>19</volume>
          (
          <issue>3</issue>
          ):
          <volume>281</volume>
          {
          <fpage>300</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Poole</surname>
          </string-name>
          .
          <article-title>Linear Algebra: a modern approach</article-title>
          . Brooks/Cole, Cengage learning,
          <source>third edition</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Sangwin</surname>
          </string-name>
          .
          <source>Computer Aided Assessment of Mathematics</source>
          . Oxford University Press, Oxford, UK,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Sangwin</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Jones.</surname>
          </string-name>
          <article-title>Asymmetry in student achievement on multiple choice and constructed response items in reversible mathematics processes</article-title>
          .
          <source>Educational Studies in Mathematics</source>
          ,
          <volume>94</volume>
          :
          <fpage>205</fpage>
          {
          <fpage>222</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Sangwin</surname>
          </string-name>
          and
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Kocher</article-title>
          .
          <source>Automation of mathematics examinations. Computers and Education</source>
          ,
          <volume>94</volume>
          :
          <fpage>215</fpage>
          {
          <fpage>227</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>