<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Toward Better Training in Peer Assessment: Does Calibration Help?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julia Morris</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jennifer Kidd</string-name>
          <email>jkidd@odu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stacie Ringleb</string-name>
          <email>sringleb@odu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Darden College of Education Old Dominion University Norfolk, VA, U.S</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mechanical &amp; Aerospace Engineering Old Dominion University Norfolk, VA, U.S</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Yang Song, Zhewei Hu, Edward F. Gehringer Department of Computer Science North Carolina State University Raleigh, NC, U.S</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>For peer assessments to be helpful, student reviewers need to submit reviews of good quality. This requires certain training or guidance from teaching staff, lest reviewers read each other's work uncritically, and assign good scores but offer few suggestions. One approach to improving the review quality is calibration. Calibration refers to comparing students' individual reviews to a standard-usually a review done by teaching staff on the same reviewed artifact. In this paper, we categorize two modes of calibration for peer assessment and discuss our experience with both of them in a pilot study with Expertiza system. Educational peer review; peer assessment; calibration.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In educational peer-review systems, students submit their artifacts
and other students rate and/or give comments on artifacts
submitted by their peers. Previous research has shown that this
process benefits both reviewers and reviewees. The reviewers
benefit by seeing others’ work and thinking metacognitively about
how they can improve their own work. The reviewees profit from
receiving comments and advice from their classmates. That
feedback is both more timely and more copious than feedback
from teaching staff [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        The efficacy of peer assessment depends heavily on the quality of
the reviewing. Left to their own devices, students tend to examine
peers’ work uncritically, and make few suggestions on how to
The Peerlogic project is funded by the National Science Foundation under grants
1432347, 1431856, 1432580, 1432690, and 1431975.
improve it. When asked to rate it on a Likert scale, they gravitate
to the upper end of the scale, making little distinction between the
various artifacts that they review [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        One approach to improving the quality of peer review is to
interpose a calibration phase before the actual peer-review task.
“Calibration” refers to having students evaluate sample artifacts
that have already been rated by teaching staff. Then the online
peer-review system can use the comparison between students’
reviews and those of the teaching staff to calculate review
proficiency values for students. This approach was pioneered in
Calibrated Peer Review ™ [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and later adopted by other
systems as well (such as Coursera [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], EduPCR5.8 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Expertiza
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Mechanical TA [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Peerceptiv [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and Peergrade.io).
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. TWO MODES OF CALIBRATION</title>
      <p>
        We can divide calibrations into two modes. The first mode
separates the calibration from actual peer-review assignments, in
which students rate and on comment each other's work. We call
this stand-alone calibration. An example is Calibrated Peer
Review™. A calibrated assignment has a separate calibration
phase in which students need to rate three sample artifacts, one of
which is exemplary, and the other two of which have known
defects. The system uses their ratings to calculate the Reviewer
Competency Index, which is a measure of the student’s review
proficiency [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The motivation for this mode of calibration is
to train students to become proficient reviewers first before they
start to review each other’s artifacts. The resultant peer-review
grades should have greater validity and thereby, make grading
easier for the teaching staff.
      </p>
      <p>
        The other mode of calibration combines the calibration with
ordinary peer-review activity. In the peer-review phase, students
review both sample artifacts and artifacts submitted by their peers.
Usually, they are not aware of whether the artifact is a sample for
calibration or an actual peer submission. We call this approach
mixed calibration. An example is the Coursera system [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In a
calibrated assignment, the teaching staff grades only a small
number of artifacts, which are then used as sample artifacts in the
peer-review phase. When doing peer review, each student
evaluates four random artifacts and one sample artifact that has
already been graded by teaching staff. Just as in stand-alone
calibration, the review proficiency is determined by agreement on
the sample artifacts between students and teaching staff.
Comparing these two modes of calibration, we observe that
standalone calibration requires more work for teaching staff: they need
to locate sample artifacts (which they could take from earlier
semesters) and set up a calibration phase in the assignment.
Students are aware of the fact that they are rating some sample
artifacts, so they may pay more attention than they do in the actual
peer-review tasks, which also makes it harder to test the efficacy
of the calibration. However, stand-alone calibration fits in well
with in-class lecture. Instructors can give students time to do the
calibration in class as training. They can also explain how the
rating was done on sample artifacts so that students may have a
better understanding of the rating rubrics.
      </p>
      <p>Mixed calibration does not emphasize training — to make
students better peer-reviewers — but score aggregation — how to
identify the good reviewers and use their peer-review responses to
aggregate grades for each artifact. Therefore, students who did
poorly on the peer-review do not receive any pedagogical
intervention, though their identities are known. So the mixed
calibration is used more often by classes of massive sizes, e.g.
some courses in the Coursera system.</p>
    </sec>
    <sec id="sec-3">
      <title>2.1 Calibration in Expertiza</title>
      <p>Beginning in 2016, the Expertiza system has included a
calibration feature, which supports both stand-alone calibration
and mixed calibration. In setting up an assignment, an instructor
can designate an assignment as a calibrated assignment, and
submit sample artifacts and “expert” reviews. The instructor can
give students the right to do reviews, but not submit work. This
makes the assignment a stand-alone calibration assignment.
(Ordinarily, students are permitted both to submit and to review.)
The review was done in double-blind style in Expertiza. In neither
calibration mode did student reviewers see the expert review
before they finished reviewing an artifact. But, after a student
finishes reviewing an artifact that is a calibration sample that has
been reviewed by the instructor, Expertiza shows a comparison
between the student’s review and the expert review (see Figure 1
for an example). No update is allowed after the expert review is
displayed.</p>
    </sec>
    <sec id="sec-4">
      <title>3. ASSIGNMENT DESIGN</title>
      <p>Three instructors at two universities set up a total of four
calibration assignments using Expertiza. Those assignments used
calibration feature in Spring 2016 but did not have calibration in
Fall 2015. Other than the calibration, those four assignments were
of the same settings including review rubrics.
</p>
      <p>Assignment 1: Course: Foundations and Introduction to
Assessment of Education; Assignment: Grade Sample
Lessons. This assignment was a precursor to engage
students in evaluating peers’ writing before they assessed
each other’s work. Pre-service teachers were asked to grade
two different example lesson plans with a five-item rubric


by ranking the (1) importance, (2) interest, (3) credibility, (4)
effectiveness, and (5) writing quality of the lesson. They
were asked to consider what was effective and ineffective in
each lesson based on the strengths and weakness they
identified from the rubric. The artifacts were lessons created
by students of prior semesters whose lessons exemplified
both noteworthy achievements and pitfalls. By evaluating
these two lessons, students gain valuable insight into the act
of evaluating peers’ writing and are provided with a model
to guide their own submissions. The students’ completed the
calibration assignment, ranking each of the rubric categories
on a 1-5 scale. Their results were then compared with the
“expert” review completed by the course instructor.</p>
      <p>Assignment 2: Course: Project Design and Management I;
Assignment: Practice Introduction to Peer Review. This
assignment was designed to expose students to writing an
introduction for their senior project, to orient them to the
peer review process, and to understand the instructor’s
expectations for the peer review assignment. The calibration
exercise had the students peer review two introductions
from a previous class, one with a good grade and one that
received a poor grade. The calibration exercise was
performed before the introduction was drafted. The general
introduction assignment included a draft with an in class
peer review, a second draft peer review using Expertiza and
the submission of a final draft.</p>
      <p>Assignment 3: Course: Object-Oriented Design and
Development; Assignment: Calibration for reviewing
Wikipedia pages. This assignment was to get the students
ready to write and peer-review Wikipedia entries. The
instructor provided a list of topics on recent
softwaredevelopment techniques, frameworks, and products. Some
of these topics had pre-existing Wikipedia pages; some did
not. Where the pages existed, they were stubs or otherwise
in need of improvement. Students could choose one topic
and create the corresponding page. Then students were
required to review at least two others’ artifacts and provide
both textual feedback and ratings.</p>
      <p>We created a separate assignment for calibration. The
sample artifacts were chosen from a previous semester. The
instructor took two reviews done by good reviewers and
made further changes in an effort to make the review of
exemplary quality.</p>
      <p>Assignment 4: Course: Object-Oriented Design and
Development; Assignment: create and review CRC
(Classresponsibility-collaborator) cards. CRC cards are an
approach to designing object-oriented software. The
instructor’s students tended to make the same mistakes,
semester after semester. The goals of this calibration
assignment were to (1) allow students to submit their own
CRC-card design and (2) review some CRC-card designs
that contained common mistakes. In this assignment, each
student reviewed one of their peers’ designs, and two
designs arranged by the instructor to contain common
mistakes. These designs were created by merging the errors
made by previous students on an exam.</p>
      <p>Unlike the other three calibration assignments, this
assignment did not precede another assignment where the
students submitted their own work. Rather, it was done as
practice for the next exam.





We asked the instructors to identify a few good reviewers in the
actual peer-review assignments of exemplary quality to compare
the student performance on the calibration assignment and the
actual assignments for which they received training. To test
student performance on different assignments, we used the metrics
below:
</p>
      <p>Percentage of exact agreement on each criterion. All the
rubrics used in our experiments were scored on either a
0to-5 or a 1-to-5 scale. On each criterion, exact agreement
was when instructor and student gave exactly the same score.
Percentage of adjacent agreement on each criterion. On each
criterion, adjacent agreement means that the score assigned
by the student is within ±1 of the instructor’s score.</p>
      <p>Percentage of empty comment boxes. Some criteria asked
students to give both a score and textual feedback. In the
calibration, the instructors tried to give textual feedback on
all these criteria. If the sample artifact was in good shape,
the instructors commented why it was good; otherwise, if
the sample artifact needed improvement, the instructors
suggested changes for the author to consider. We hoped this
would encourage students to comment on more of the
criteria.</p>
      <p>Average non-empty comment length. We counted the words
in the non-empty responses. In calibration, the expert
reviews were usually longer than the average of students’
review (see Figure 1 for example).</p>
      <p>
        Average of number constructive comments. We tried to
measure how much constructive content was provided in the
non-empty responses. We used the same constructive
lexicon used by Hsiao and Naveed [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This lexicon
focuses mainly on assessment, emphasis, causation,
generalization, and conditional sentence patterns.
of words. The Flesch-Kincaid readability index rates work
between 0 (difficult to read) and 100 (easy to read).
Conversational English is usually between 80 and 90 on this
index. Text is considered to be hard to read (usually
requiring a college education or higher) if the index is lower
than 50.
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. HOW CALIBRATION AFFECTS</title>
    </sec>
    <sec id="sec-6">
      <title>STUDENT PERFORMANCE</title>
    </sec>
    <sec id="sec-7">
      <title>4.1 Results for stand-alone calibration</title>
      <p>The first three calibration assignments (Assignment 1, 2 and 3)
were followed by an actual assignment where the students carried
out the same kind of review on which they were calibrated. We
measured the percentage of empty comments, average comment
length, and number of constructive comments in the response to
each criterion, and the overall readability. In the following actual
assignment, we also measured the students’ agreement on
exemplary reviews (done by students). The results are shown in
Table 1.</p>
      <p>In all three classes, we found there was a similar amount of exact
agreement on calibration assignments and following assignment.
But we observed increases in the adjacent agreement on the
following assignment. The reason for that could be that the
calibration phase led students to become more skilled and more
polite as reviewers. The instructor of assignment 1 observed that
her students were critical or even bullying, in their peer reviews at
the very beginning of the semester. In the calibration phase,
students were able to see how the instructor reacted to various
issues and what the instructor grades were. This gave students
guidance on how to rate artifacts that still needed improvement.
The comment length between the calibration assignment and the
following assignment were almost the same. Two out of three
classes had a higher average comment length after they did
calibration, compared with corresponding assignments last
semester.</p>
      <p>From the amount of constructive content per response to each
criterion, we found that the students tended to give as many or
more constructive comments in the peer-review after the
calibration. Two out of three classes made more constructive
comments after calibration compared with corresponding
assignments last semester.</p>
      <p>In this study, we found that students tended to write more
complicated sentences in calibration tasks, but in the assignments
right after the calibration, their comments were a little easier to
read but close to college level, which was acceptable to instructors.</p>
    </sec>
    <sec id="sec-8">
      <title>4.2 Results for mixed calibration</title>
      <p>Assignment 4 was our only experiment with the mixed calibration
mode: each student reviewed two calibration submissions and one
submission from their classmates. Unlike Assignments 1–3, which
aimed to train students to become better reviewers on the actual
peer assessment, Assignment 4 was not followed with an “actual”
assignment on the same topic. Instead, Assignment 4 was
designed to give students the opportunity to see common mistakes
that others had made on a certain kind of question (on CRC-card
design) on exams in earlier semesters.</p>
      <p>On Assignment 4, the percentage of exact agreement was 52.2%
and percentage of adjacent agreement was 91.3%, which were
both very high. This was partially due to a review rubric that
asked students to count the number of errors of certain types (e.g.
the number of class names that are not singular nouns), instead of
ordinary rubric criteria that ask students to rate the artifact on
some aspect (e.g., the language usage of an article). This rubric
design reduces ambiguity and thereby increased the agreements.
The percentage of the empty comment was 77.0%, the average of
non-empty comment length was 5.4 and average of number
constructive comments was 0.13, which are all lower than
Assignment 1-3. The ostensible reason was that the review rubric
was not designed to encourage students to give textual comments,
but simply to count the errors. The review readability index was
60.1, which indicates that for those reviewers who gave textual
feedback, the feedback was not short and simple as we expected.
We hypothesized that after this calibration, student's’ average
score on related questions on the exam would be higher. We
compared the student performance on CRC-card related questions
in exams of this semester (with calibration as training) and last
semester (without training). However, we found that the students’
average grade was 85.3% on those questions in this semester, and
85.4% on last semester. We did not find any significant change
between this semester and last semester. Upon seeing those results,
we surmised this calibration assignment was done several weeks
before the next exam, and, without follow-up practice, students
forgot the training they received.</p>
    </sec>
    <sec id="sec-9">
      <title>5. WHAT SAMPLE ARTIFACTS WE</title>
    </sec>
    <sec id="sec-10">
      <title>SHOULD USE FOR CALIBRATION?</title>
      <p>After students finish the calibration, the instructor can see the
calibration reports for each artifact, as shown in Figure 2. Each
table shows the students’ grades on each question on a sample
artifact. The green color highlights the expert grade, and the
bolded number was the plurality of students’ grades.</p>
      <p>Figure 2 shows a sample artifact where the calibration was quite
successful, with exact agreement of more than 40% and adjacent
agreement of almost 80%. However, it is still not clear that if it
was related to the quality of the artifact. When we calculate the
percentages of agreements for each sample artifacts, we found that
the level of agreement is related to the quality of the artifact: the
higher grade that a sample had, the higher agreement that students
might achieve. This raises another question: what kind of artifacts
work better as samples in calibration?
We put the percentages of agreement and grades for the artifacts
together to compare the relationship between the agreement and
the grades that the sample artifacts received. We used both the
sample artifacts and the artifacts reviewed by the exemplary
reviewers. The distribution and fit line are shown below.
We find that the samples that received higher grades usually have
higher levels of agreement (on both exact agreement and adjacent
agreement). The lower quality a sample is, the lower agreement
we observed between teaching staff and students.</p>
      <p>We looked into the samples used in each assignment, and we
found that usually it is harder for students to make the same
judgment as teaching staff on an artifact of low quality. There
could be multiple reasons. The first reason is that teaching staff
has seen more artifacts, therefore they know the distribution of the
quality of the artifacts and thereby they made better judgments.
For student reviewers, they may be able to tell an artifact is of low
quality based on one criterion, but they could be more critical
than warranted since they have not seen even worse examples.
From this perspective, it is important for instructors to use at least
one or two low-quality sample artifact as a sample artifact to show
students how to rate poor work.</p>
      <p>
        Another factor that may lower the agreement between teaching
staff and students is the reliability of the criterion: some of the
criteria are not specific enough for the reviewers to make reliable
judgments [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. E.g. the criterion, “(On Likert scale) does the
author provide enough examples in this article?” is not reliable,
since “enough” is not well defined. To improve review rubrics,
instructors can create “advice” for each level (sometimes known
as an “anchored scale”). For example, “⅕ - No example
provided”, etc. From this perspective, the calibration can also be
used to test the instructor’s review rubric.
      </p>
    </sec>
    <sec id="sec-11">
      <title>6. CONCLUSION</title>
      <p>In this paper, we have described our experience with the
calibration in peer assessment in Expertiza. We first introduced
two modes of calibration that have been used in online peer
assessment systems, which are stand-alone calibration and mixed
calibration. Stand-alone calibration trains students to become
better reviewers, while mixed calibration finds credible reviewers
in the course of performing peer assessment. We also discussed
the pedagogical scenario in which each mode is suitable.
We calculated the agreement between students’ rating and
teaching staff’s rating on the sample artifacts. We found that
students in our assignments, on average agreed exactly with
teaching staff on more than 40% of ratings. This means that on
more than 40% of the ratings done by students during calibration
gave exactly the same scores given by teaching staff. In addition,
more than 70% of the ratings done by students gave the score
within the ±1 range to the scores given by teaching staff. To test if
students still perform as well on the actual peer assessment after
training, we asked the teaching staff to identify some good
reviewers in each course. Using their reviews as exemplars, we
found that, in the actual peer assessment phases, the agreement
was similar to that on the calibration assignments, sometime even
a little higher.</p>
      <p>We compared the volume of textual feedback from the semester
with calibration and the previous semester without calibration. We
found that after calibration, students tend to give more extensive
textual feedback, fill in more text boxes with comments, and give
more constructive feedback.</p>
      <p>We also found that the level of rating agreement between students
and teaching staff is related to the quality of the artifact; namely
students tended to agree less with teaching staff on artifacts of low
quality. To improve agreement, we suggested: (1) on the
calibration, an instructor can use both median-quality artifacts and
low-quality artifacts as samples and (2) the instructor can provide
“advice” for each level of each criterion.</p>
      <p>One future study we are interested in is to calibrate the textual
feedback. In this paper, we have only calibrated the numerical
scores. It is possible that both a student and the teaching staff
gave a ⅘ on one criterion on a sample artifact, but may not see
the same issue. This kind of agreement can only be measured by
calibration of textual feedback.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E. F.</given-names>
            <surname>Gehringer</surname>
          </string-name>
          , “
          <article-title>A Survey of Methods for Improving Review Quality,” in New Horizons in Web Based Learning</article-title>
          , Y. Cao,
          <string-name>
            <given-names>T.</given-names>
            <surname>Väljataga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. K. T.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Leung</surname>
          </string-name>
          , and M. Laanpere, Eds. Springer International Publishing,
          <year>2014</year>
          , pp.
          <fpage>92</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. F.</given-names>
            <surname>Gehringer</surname>
          </string-name>
          , “Closing the Circle:
          <article-title>Use of Students' Responses for Peer-Assessment Rubric Improvement,” in Advances in Web-Based Learning -</article-title>
          -
          <source>ICWL</source>
          <year>2015</year>
          ,
          <string-name>
            <given-names>F. W. B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Klamma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Laanpere</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. F.</given-names>
            <surname>Manjón</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. W. H.</given-names>
            <surname>Lau</surname>
          </string-name>
          , Eds. Springer International Publishing,
          <year>2015</year>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Robinson</surname>
          </string-name>
          , “Calibrated Peer ReviewTM,”
          <source>Am. Biol. Teach.</source>
          , vol.
          <volume>63</volume>
          , no.
          <issue>7</issue>
          , pp.
          <fpage>474</fpage>
          -
          <lpage>480</lpage>
          , Sep.
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Russell</surname>
          </string-name>
          , “
          <article-title>Calibrated peer review-a writing and criticalthinking instructional tool</article-title>
          ,” in Teaching Tips: Innovations in Undergraduate Science Instruction,
          <year>2004</year>
          , p.
          <fpage>54</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Piech</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Do</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Koller</surname>
          </string-name>
          , “
          <article-title>Tuned Models of Peer Assessment in MOOCs,” ArXiv13072579 Cs Stat</article-title>
          , Jul.
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>X.</given-names>
            <surname>Hao</surname>
          </string-name>
          , “
          <article-title>E-learningoriented incentive strategy: Taking EduPCR system as an example,”</article-title>
          <source>World Trans. Eng. Technol. Educ.</source>
          , vol.
          <volume>11</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>174</fpage>
          -
          <lpage>179</lpage>
          , Nov.
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Gehringer</surname>
          </string-name>
          , “
          <article-title>Expertiza: information management for collaborative learning</article-title>
          ,
          <source>” Monit. Assess. Online Collab. Environ. Emergent Comput. Technol. E-Learn. Support</source>
          , pp.
          <fpage>143</fpage>
          -
          <lpage>159</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thornton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Leyton-Brown</surname>
          </string-name>
          , “
          <string-name>
            <surname>Mechanical</surname>
            <given-names>TA</given-names>
          </string-name>
          :
          <string-name>
            <surname>Partially Automated High-Stakes Peer</surname>
            <given-names>Grading</given-names>
          </string-name>
          ,”
          <source>in Proceedings of the 46th ACM Technical Symposium on Computer Science Education</source>
          , New York, NY, USA,
          <year>2015</year>
          , pp.
          <fpage>96</fpage>
          -
          <lpage>101</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Schunn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Godley</surname>
          </string-name>
          , and S. DeMartino, “
          <article-title>The Reliability and Validity of Peer Review of Writing in High School AP English Classes,”</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Adolesc</surname>
          </string-name>
          . Adult Lit., p.
          <source>n/a-n/a, Apr</source>
          .
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Hsiao</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Naveed</surname>
          </string-name>
          , “
          <article-title>Identifying learning-inductive content in programming discussion forums,”</article-title>
          <source>in IEEE Frontiers in Education Conference (FIE)</source>
          ,
          <year>2015</year>
          . 32614
          <year>2015</year>
          ,
          <year>2015</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          , and E. Gehringer, “
          <article-title>An Experiment with Separate Formative and Summative Rubrics in Educational Peer Assessment,” in Submitted to IEEE Frontiers in Education Conference (FIE</article-title>
          ),
          <year>2016</year>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Kincaid</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Others</surname>
          </string-name>
          , “
          <article-title>Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel</article-title>
          .,” Feb.
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>