<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Essay Revision and Corresponding Grade Change as Captured by Text Similarity and Revision Purposes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sonia Cromp</string-name>
          <email>snc40@pitt.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diane Litman</string-name>
          <email>dlitman@pitt.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Pittsburgh</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Writing and revision are abstract skills that can be challenging to teach to students. Automatic essay revision assistants o er to help in this area because they compare two drafts of a student's essay and analyze the revisions performed. For these assistants to be useful, they need to provide useful information such as whether the revisions are likely to lead to an improvement in the student's grade. It is necessary to better understand the connection between revisions and grade change so that this information could be displayed in an assistant. So, this work explores the relationship between the tf-idf cosine similarity of two essay drafts and resulting essay grade change. Prior work has demonstrated that identifying the revisions between drafts, then labeling each revision with the purpose behind why the revision was performed is useful to predicting grade change. However, this process is expensive because this sort of annotation is time-consuming for humans. Moreover, classi ers achieve lower accuracy than humans when predicting purposes. Using similarity measures instead of or as supplement to revision purposes may correct these issues, as similarity can be computed automatically and without the issue of classication accuracy. As such, the correlations between grade change and the similarity measure are compared to the correlations between grade change and revision purposes with the potential use-case of an automatic writing assistant in mind. Findings suggest tf-idf cosine similarity captures overall essay and overall grade change while revision purposes capture lighter changes that x errors or cause the essay to read better.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;writing</kwd>
        <kwd>revision</kwd>
        <kwd>education</kwd>
        <kwd>similarity</kwd>
        <kwd>NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Copyright ©2021 for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0)
portant to investigate reliable methods to quantify revisions
as well as how these revisions relate to external measures
such as grade change across drafts.</p>
      <p>
        Prior work [
        <xref ref-type="bibr" rid="ref14 ref16">16, 14</xref>
        ] has focused on analyzing the purposes or
intentions behind the revisions performed, such as labeling
revisions as \Conventions" when they correct a spelling error
or \Evidence" when adding an example to support a claim.
Unsurprisingly, a greater quantity of revisions is associated
with a greater grade change. Further, most university-level
writing assignment rubrics place more importance on ideas,
reasoning and evidence than spelling or adherence to writing
conventions. As such, revisions to change the meaning of an
essay, such as Evidence, Claims or Reasoning revisions, are
associated with more grade change than minor changes such
as Conventions revisions. These works have demonstrated
revision purposes to be useful in assessing essay change.
However, revision purposes are time-consuming to obtain via
human annotation and classi ers achieve lower classi cation
accuracy than do human annotators. First, the sentences in
the two drafts must be aligned to indicate which sentences
in each draft correspond to each other. Second, revision
operations are determined. Sentences removed from the old
draft are marked as deleted, those inserted to the new draft
marked as added and those edited but present in both drafts
marked as modi ed. Third, each sentence that has been
modi ed, added or deleted must be labeled with a revision
purpose. For an example of an annotated essay from the
ArgRewrite V.2 corpus[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], see Table 1.
      </p>
      <p>Embedding-based similarity measures o er an alternative to
revision purposes that can be obtained in fewer steps,
automatically and without the issue of classi cation accuracy.
However, similarity measures can also be calculated in
additional ways when more annotation such as sentence
alignments is available. Similarity measures likely are also able
to capture additional information that cannot be detected
by revision counts alone.</p>
      <p>
        As an extreme example, consider two college students, A
and B, revising their essays that are both N sentences long.
Student A replaces one word in each sentence of their
essay with a synonym while Student B re-organizes their
essay to present their logic more clearly. Then, consider that
an automated revision assistant were to label revision
purposes using the popular binary schema from [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] that
distinguishes between revisions that change meaning (such as
      </p>
    </sec>
    <sec id="sec-2">
      <title>Old Draft</title>
      <p>Anything that can save lives is
good for society.</p>
      <p>Despite the limitations in technology,
self-driving will save lives.</p>
      <p>No other bene t matters.</p>
      <p>Most tra c fatalities should have been
prevented because the drivers simply
should not have been driving.</p>
    </sec>
    <sec id="sec-3">
      <title>Car accidents are the top cause of death for teenagers. New Draft</title>
    </sec>
    <sec id="sec-4">
      <title>Despite these technological limitations, self-driving will save lives.</title>
    </sec>
    <sec id="sec-5">
      <title>Most tra c fatalities should have been prevented because the drivers should not have been driving. Self-driving reduces fatalities a hundredfold.</title>
    </sec>
    <sec id="sec-6">
      <title>Operation</title>
    </sec>
    <sec id="sec-7">
      <title>Delete</title>
    </sec>
    <sec id="sec-8">
      <title>Modify</title>
    </sec>
    <sec id="sec-9">
      <title>Delete</title>
    </sec>
    <sec id="sec-10">
      <title>Modify Add</title>
    </sec>
    <sec id="sec-11">
      <title>Delete</title>
    </sec>
    <sec id="sec-12">
      <title>Claims</title>
    </sec>
    <sec id="sec-13">
      <title>Fluency</title>
    </sec>
    <sec id="sec-14">
      <title>Fluency</title>
    </sec>
    <sec id="sec-15">
      <title>Reasoning</title>
    </sec>
    <sec id="sec-16">
      <title>Evidence</title>
    </sec>
    <sec id="sec-17">
      <title>General Content</title>
      <p>
        citing a new source of evidence or changing the thesis
statement) and revisions that do not change meaning (such as
xing spelling errors or replacing a word with a synonym).
Neither student performed any meaning-changing revisions.
So, the revision assistant would see that Student A made
N non-meaning-changing revisions and Student B made up
to N non-meaning-changing revisions. To the revision
assistant, it may appear as if both students performed the same
types of revisions; the only di erence is that Student A made
many more revisions. However, Student A can likely expect
a smaller grade change than can Student B because Student
B's ideas are now more clearly presented and understood
whereas Student A's revisions may pass nearly unnoticed to
a reader of the old and new drafts. As such, the counts of
revision purposes was misleading in predicting grade change.
However, embedding cosine similarity could capture the
difference: the similarity of Student A's old and new drafts is
very close to 1 (identical) while the similarity between
Student B's old and new drafts would be noticeably lower (less
similar). So, a revision assistant using a similarity would be
able to capture the greater change in Student B's essay.
This present work explores using a tf-idf embedding cosine
similarity measure to quantify the relationship between the
similarity of essay drafts and grade change between drafts.
For comparison, the work also presents the relationship
between numbers of revision purposes and grade change.
Further, these relationships are analyzed at each subsequent
step of annotation shown at Table 1 - rst where no
annotation has been performed and the essay drafts are both just
raw strings of text (revision purposes are not yet available
at this stage, so only the similarity measure is performed
here), then with the rst annotation step of aligned
sentences (where the similarity measure and the number, but
not the purposes, of revisions is available), then with revision
purposes labeled using either a simple but commonly-used
two-class schema[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] or a ner-grained multi-class schema[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
presented in section 3.3.
      </p>
      <p>
        In situations where both the similarity measure and revision
counts are available, the correlation between similarity
measure and grade change when controlling for revision counts is
also considered. This information is provided because prior
related work ([
        <xref ref-type="bibr" rid="ref14 ref16">16, 14</xref>
        ], discussed in next section) has
demonstrated revision counts to be useful when assessing an essay's
revisions for grade change. Thus, it is desirable to determine
if the similarity measure provides additional information on
top of revision counts when evaluating grade change, or if the
similarity measure and revision counts provide the same
information and controlling for one renders the other
insignificant. Findings ultimately suggest that tf-idf embedding
cosine similarity does well at capturing deeper,
meaningaltering changes to the essay and overall essay grade change
while revision purposes capture how the essay changes with
respect to obeying conventions such as vocabulary choice
and grammar.
      </p>
      <p>Section 2 contains an overview of related work. Section 3
explains the dataset and tools employed in the analysis, which
is in Section 4. A discussion of ndings is in Section 5 and
the conclusions are listed in Section 6.</p>
      <sec id="sec-17-1">
        <title>2. RELATED WORK</title>
        <p>
          E ective writing can be a di cult skill to teach, so there
has been signi cant e ort towards analyzing methods and
developing tools to help in this goal. One center of attention
has been on discovering what makes a revision e ective[
          <xref ref-type="bibr" rid="ref12 ref14 ref2 ref4">4,
14, 2, 12</xref>
          ] and developing tools and writing assistants to help
students revise e ectively[
          <xref ref-type="bibr" rid="ref13 ref17">13, 17</xref>
          ]. Much of the work focuses
on revisions to Wikipedia[
          <xref ref-type="bibr" rid="ref8 ref9">9, 8</xref>
          ] and students' argumentative
essays[
          <xref ref-type="bibr" rid="ref16 ref7">16, 7</xref>
          ].
        </p>
        <p>
          There are three broad categories of information that can be
used to describe a revision. First, there is the type of
revision operation that was performed, such as adding, deleting
or modifying a sentence. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] analyzed which revision
operations are associated with essay improvement, as measured by
features such as lexical diversity and amount of rst, second
or third person.
        </p>
        <p>
          Second, there is the purpose or intent behind the revision
that was performed, such as to x a grammar mistake or
alter the meaning of a claim. There are many revision
schemata in common use. As a minimum amount of
granularity when classifying revision purposes, [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] propose a
binary schema of text-base revisions. One class is \Content"
revisions that alter text meaning, such as changing the
evidence cited or the ideas in the thesis statement, and the
other class is \Surface" revisions that do not alter meaning,
such as changing citation format or xing a spelling error in
the thesis statement. Both of these categories can be
further broken down into ne-grained categories. For instance,
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] used a set of 13 revision purposes on Wikipedia articles,
to compare the revision strategies between articles that are
featured on the Wikipedia homepage and those that have
not been featured.
        </p>
        <p>
          A third way of describing a revision is features that can
be automatically gathered about the revision such as edit
distance between the old and new draft or the change in
word count between drafts. For instance, [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] use an array
of features including Named Entity Recognition, word-level
edit distance and number of inserted or deleted characters
to build a classi er to distinguish between factual and
uency revisions. These sorts of features are used in classi ers
that aim to predict revision purposes. For instance, [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] uses
revision operation and statistics such as edit distance, word
count and presence of grammatical and spelling errors as
features to a revision purpose classi er for student
argumentative essays. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] used many features including number of
informal words, change in character counts and punctuation
to build a revision purpose classi er for Wikipedia edits.
While there has been application of similarity measures to
revision studies, such as Latent Semantic Analysis (LSA)[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
Levenshtein Distance[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and Kullback-Leibler divergence[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ],
these similarity measures are often used along the way to
predict other information about revisions. For instance, [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
used Levenshtein Distance as a feature for predicting
revision purpose. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] used LSA to analyze which revision
operations are associated with larger change in essay similarity.
The contribution of this present work is twofold. First, it
considers tf-idf cosine similarity as a method of quantifying
grade change between drafts, rather than using the
similarity measure to predict revision purposes and subsequently
using revision purposes to quantify grade change. See
Figure 1 and Figure 2 for a visualization of this change.
Searching through the literature, an application of similarity
measures in this way does not appear to have been done before.
This speci c similarity measure was chosen because
embedding cosine similarity is a very simple, not state of the art,
method of calculating similarity compared to methods like
LSA or Kullback-Leibler divergence. So, if cosine
similarity is shown to be a better predictor for grade change than
revision purposes, when using a high-quality corpus ([
          <xref ref-type="bibr" rid="ref1">1</xref>
          ],
introduced in Section 3) that has been human-annotated for
sentence alignments and revision purposes, then this
pattern may be able to also hold when using more advanced
similarity measures or lower-quality corpora such as
automatically annotated ones with less reliable revision purpose
labels. Further, only tf-idf is presented in this work, but
the same experiments were performed with sent2vec[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and
BERT[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] embeddings, which yielded similar results.
Because of the similar results, only the tf-idf embedding is
reported in this paper for simplicity and to demonstrate that
even this simpler embedding is able to perform favorably
compared to revision purposes. The second contribution of
this work is that it explores the information revealed by
tf-idf cosine similarity as subsequently more additional
information becomes available: rst on its own, then with
essay drafts that have had their sentences aligned, then with
coarse-grained revision purposes labeled, then nally with
ne-grained revision purposes labeled.
        </p>
      </sec>
      <sec id="sec-17-2">
        <title>3. METHODOLOGY 3.1 Data</title>
        <p>
          The ArgRewrite V.2 corpus[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] was used as the dataset for
this work. 86 recruited graduate and undergraduate
university students wrote argumentative essays about self-driving
cars and revised their essays two times for a total of three
drafts. Each draft has been graded as described in
Section 3.2, sentence purposes aligned between drafts and
revision purposes annotated as described in section 3.3. For the
present analysis, only revisions between drafts 1 and 2 were
used. An example section of one essay and its annotations
is given in Table 1.
        </p>
      </sec>
      <sec id="sec-17-3">
        <title>3.2 Grades</title>
        <p>
          Human graders assigned scores to each essay draft using a
10-category rubric, with each category being evaluated on a
scale from 1 (poor) to 4 (excellent), for minimum possible
score of 10 and maximum possible score of 40. The names of
these categories and the description for the 4-point/excellent
category are provided in Table 2. All drafts were scored
separately by two annotators and the Quadratic Weighted
Kappa was 0.537[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>Two additional grade categories beyond those in Table 2 are
also added in the analysis: Average, which is the average
score across all the rubric categories for some draft, and
Total, which is the total score for a draft on the 40-point
scale. For information on how grades changed between the
drafts, see Table 3.</p>
      </sec>
      <sec id="sec-17-4">
        <title>3.3 Revision Purposes</title>
        <p>
          The ne-grained revision purposes used by the ArgRewrite
V.2 corpus are able to capture more of the variation in
revisions by breaking the coarse categories of Surface and
Content as described by [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] each into nine smaller sub-categories,
        </p>
        <p>Requirement for
Maximum (4) Points
The author responds to all parts
of the prompt and the entire
essay is focused on the prompt.</p>
        <p>The author provided a clear,
nuanced and original statement
that acted as a speci c stance
for or against self-driving cars.</p>
        <p>The author makes multiple,
distinct claims that are clear,
and align with both their thesis
statement and the given reading.</p>
        <p>They fully support the author's
argument.</p>
        <p>The author provides speci c and
convincing evidence for each
claim, and most evidence is
given through detailed
examples, direct quotations, or
detailed examples from the
provided reading. The source of
the evidence is credible and
acknowledged/cited where
appropriate.</p>
        <p>All claims are supported with
clear reasoning that shows
thoughtful, elaborated analysis.</p>
        <p>The essay has an introduction,
body and conclusion and a
logical sequence of ideas. Each
paragraph makes a distinct claim.</p>
        <p>The essay explains a di erent
point of view and elaborates why
it is not convincing or correct.</p>
        <p>Throughout the essay, word
choices are speci c and convey
precise meanings(e.g.,
\Self-driving cars are dangerous
because the technology is still
not advanced enough to address
the ethical decisions drivers must
make.")
All sentences are clear because
of correct and appropriate word
choices and sentence structure.</p>
        <p>
          The author makes few or no
grammatical or spelling errors
throughout their piece, and the
meaning is clear.
with Surface containing Conventions, Organization and
Fluency, and Content containing Precision, Claims, Evidence,
Reasoning, Rebuttal and General Content [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Each
sentence pair with a revision operation of Add, Delete or
Modify is assigned one purpose. See Table 5 for details of the
ne-grained categories and Table 4 for information on the
number of occurrences of each revision purpose. Revision
        </p>
        <p>
          Better
9
23
18
19
19
25
26
12
14
13
59
59
purpose annotations of this corpus were performed by three
annotators, with a Fleiss' kappa[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] of 0.65 on a sample of
ve essays that all annotators labeled independently prior
to labeling disjoint sets of the remaining essays.
The revision purpose categories and grading rubric
categories bear some resemblance to one another, with most
revision purposes corresponding to one of the rubric
categories. In particular, the Conventions, Organization,
Fluency, Precision, Claims, Evidence, Reasoning and Rebuttal
revision purposes each align with the rubric categories of the
same names. The de nitions of these revision purposes and
the criteria for the rubric categories correspond such that
an Evidence-purpose revision, for instance, is more likely to
cause a change in the Evidence rubric category than any
other rubric category. The remaining rubric categories of
Response to Prompt and Thesis do not clearly align with
any revision purpose, although they can be thought of as
aligning most closely with Claims revisions. By the same
logic that the revision purposes can be sorted into Surface
versus Content categories, so too can the rubric categories:
rubric categories that align with Surface or Content revision
purposes respectively can be considered Surface-related or
Content-related rubric categories. Meanwhile, the General
Content revision purpose does not align with any certain
rubric category, but can be thought of as revisions that may
correspond to any of the Content-related rubric categories
such as Thesis or Evidence. While the rubric categories are
never grouped together in the tests performed (e.g.
considering \Content" and \Surface" grade categories), these
concepts are still useful for analysis and interpreting the results.
        </p>
      </sec>
      <sec id="sec-17-5">
        <title>3.4 Tf-idf Cosine Similarity Measure</title>
        <p>
          The tf-idf cosine similarity of two strings x and y in a
corpus containing W unique words is calculated in two steps:
rst, vector embedding representations ex and ey are
calculated for each string. The number of elements in each
vector equals W . To make the embedding for a string s,
the number of occurrences in the string s (Term Frequency,
TF) are counted for each word in the corpus. This results in
a length-W vector T Fs where the i-th element of T Fs
contains the number of occurrences in string s of the i-th unique
word of the corpus. Next, the number of documents
(Document Frequency) that each word occurs in are counted and
used inverse-proportionally to calculate the Inverse
Document Frequency (IDF) resulting in another length-W vector
IDF. Lastly, the tf-idf vector es to represent string s equals
the element-wise multiplication of T Fs and IDF .
After obtaining the embeddings ex and ey for two strings x
and y, the cosine similarity between the vectors is computed.
Cosine similarity is a measure of how the directions/angles
of vector ex and ey compare to one another, with 1 being
identical, 0 being orthogonal (signifying no correlation
between the meanings of the two strings) and -1 being opposite
(antonymous strings such as \up" and \down"). Because
tfidf embeddings never contain negative numbers, tf-idf cosine
similarity actually varies between 0 and 1 and does not
consider antonymy. Cosine similarity is calculated as
similarity = cos( ) = ex ey = exW 2ey :
jexjjeyj
In this sense, longer-length revisions are captured by
similarity measures as having lower similarity, because they change
more words and presumably the overall meaning of the
sentence. In this work, embeddings and similarities were
calculated using the Gensim package [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
      </sec>
      <sec id="sec-17-6">
        <title>ANALYSIS</title>
        <p>Each dataset and each educational data mining project has
di erent requirements and di erent resources available. For
instance, the numbers of revision purposes have been
demonstrated in prior works to be useful in assessing revision's
impacts on grade change. However, aligning sentence pairs
between old and new drafts and subsequently labelling
revision purposes requires time and e ort that may not be
possible for all datasets and all projects. Similarity
measures are able to be used with any essays dataset, without
needing any annotation, and capture slightly di erent
information than do numbers of revision purposes.</p>
        <p>As such, for each subsequent degree that annotation is
performed, this analysis explores what information is provided
by tf-idf cosine similarity or revision purposes when
assessing rubric category grade change. The di erent levels of
annotation are: (1) raw non-annotated drafts, (2) sentences
aligned, (3) revisions labeled with coarse-grained Surface/
Content revision purposes and (4) revisions labeled with
ne-grained 9-class revision purposes. In all cases, N = 86
for these tests because there are 86 essays, each associated
with one score in each of the rubric categories. Afterwards,
some patterns in which rubric categories are best assessed
at which level of dataset annotation will be highlighted.</p>
      </sec>
      <sec id="sec-17-7">
        <title>4.1 Non-annotated Data</title>
        <p>Document-level similarity is the simplest and least
expensive method of relating grade change to essay similarity
because it does not require sentences to be aligned. Each draft
is treated as one string, an embedding is created to
correspond to each of the two strings and the similarity between
the embeddings is calculated. Each draft is treated as one
unit and there is no need to align the sentences between the
drafts or identify sentences' revision purposes. A potential
use-case for document-level similarity might be a revision
assistant for students that gives real-time feedback as they
work, because this data can be computed quickly, accurately
and automatically as a student works on their essay.
First, the similarity between old and new draft are found for
each of the 86 essays, using the process described in Section
3.4. Then, the Pearson correlation was computed between
similarity and grade change in each of the rubric categories
(correlation between similarity and Claims grade change,
between similarity and Average rubric grade change, etc.).
This method shows a Pearson correlation of r = 0:2189
with Average rubric category grade change (p = 0:0428).
As such, essay drafts with embeddings that are more
similar to each other tend to experience less grade change. For
speci c rubric categories, only Reasoning grade change is
also signi cantly correlated (r = 0:2480; p = 0:0213) with
document-level similarity. See Table 6 for all signi cant
correlations with rubric categories.</p>
      </sec>
    </sec>
    <sec id="sec-18">
      <title>Grade Reasoning Average</title>
    </sec>
    <sec id="sec-19">
      <title>Essay-Level Similarity r p -0.2480 0.0213 -0.2189 0.0428</title>
      <p>Non-annotated data is where cosine similarity measures show
the greatest advantage over revision counts, because
revision counts cannot even be used without some amount of
annotation. Further, even if annotations are available, this
essay-level similarity is still a useful way to quickly assess
and summarize the revision of an essay as demonstrated by
the signi cant level of correlation between essay-level
similarity and Average rubric grade change.</p>
      <sec id="sec-19-1">
        <title>4.2 Sentence-aligned Data</title>
        <p>Aligning sentence pairs enables a further degree of
analysis between similarity measures and grades where the unit
of comparison is at the sentence level. Without using a
measure of similarity, it is possible to simply examine the
correlation between the total number of revised sentences
and rubric grade changes. The total number of revisions
between two drafts is signi cantly correlated with grade
change in the precision (r = 0:2495; p = 0:0205) and
uency (r = 0:2172; p = 0:0446) rubric categories. See Table 7
for all signi cant correlations with rubric categories.</p>
      </sec>
    </sec>
    <sec id="sec-20">
      <title>Grade</title>
      <p>Precision
Fluency</p>
      <p>Number of Revisions
r p
0.2495 0.0205
0.2172 0.0446</p>
      <p>These results can be contrasted with the ndings of the
similarity measure in the previous subsection. While the
essaylevel similarity measure is signi cantly correlated with
Reasoning grade change and Average category grade change,
the number of revision counts is signi cantly correlated with
Precision and the Surface-related category of Fluency. As
such, the revision purpose may be more useful in
applications where Surface-level information is desired, whereas the
essay-level cosine similarity may be better for gaining an
overall picture of how the essay has changed.</p>
      <p>
        Now that sentence alignments are available, the similarity
measure can also be calculated by nding the similarity
between each old-new sentence pair and then averaging over all
sentence pairs. An old draft-new draft sentence pair (x; y)
that is not modi ed between drafts (x == y) has a
similarity of 1, sentences deleted from the rst draft (y == ?)
or added to the second draft (x == ?) have a similarity of
0 (signifying no correlation between the empty string and
the added or deleted sentence) and similarity of a modi ed
sentence (x 6= y 6= ?) may be calculated by creating
embeddings for each sentence version and then nding the cosine
similarity between the embeddings. Using Gensim tf-idf
embeddings[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], this method demonstrates a Pearson
correlation of r = 0:2692 (p = 0:0122) with Average rubric grade
change, which is slightly stronger and more signi cant than
the essay-level correlation without using aligned sentence
pairs. Sentence-level similarity is also signi cantly
correlated with grade change in the Reasoning (r = 0:2962; p =
0:0056) and Claim (r = 0:2441; p = 0:0235) rubric
categories. See Table 8 for all signi cant correlations.
Controlling for the total number of revisions yields a
correlation between average sentence tf-idf cosine similarity and
Average rubric grade change of r = 0:2720 (p = 0:0113).
Further, when controlling for total number of revisions, there
are signi cant correlations between tf-idf cosine similarity
and Claim rubric grade change (r = 0:2261; p = 0:0363),
Reasoning grade change (r = 0:2723; p = 0:0112) and
Rebuttal grade change (r = 0:3352; p = 0:0016) even though
the Rebuttal category was not signi cantly correlated with
sentence-level similarity when not controlling for number of
revisions. Conventions grade change is almost signi cantly
correlated, with a correlation of r = 0:2120 and signi cance
of p = 0:0501. Interestingly, the correlation with
Conventions (as well as a few other, non-signi cant categories), is
positive. This positive correlation means that greater
cosine similarity (meaning more similar essay drafts) is
correlated with more grade improvement in these rubric
categories. See Table 8 for all signi cant correlations.
Perhaps students who focus more on revising for Conventions
see greater cosine similarity and greater grade increases in
the Conventions category, at the cost of less improvement
in other categories like Claims and Reasoning. As a result,
these other categories have negative correlations with
similarity. This pattern of Conventions-focused revision would
be similar to Student A in the example of Section 1.
      </p>
    </sec>
    <sec id="sec-21">
      <title>Grade</title>
      <p>Claim
Reasoning
Rebuttal
Average</p>
    </sec>
    <sec id="sec-22">
      <title>Sentence Pair</title>
      <p>Similarity
r
-0.2441
-0.2962</p>
      <p>At this level, there seems to be no overlap between the
rubric categories signi cantly correlated with number of
revisions (Precision and Fluency) and the rubric categories
signi cantly correlated with average sentence-level
similarity (Reasoning and Claims). Controlling for the number
of revisions allows similarities also to be signi cantly
correlated with Rebuttal grade change. Further, similarity
measures at this level are signi cantly correlated with Average
rubric category grade change. As such, while both revision
counts and similarity measures provide useful information
when sentences have been aligned across drafts, revision
counts may be better suited to getting a general, overall
preview of the degree of essay change and degree of grade
change. Revision counts are more signi cantly correlated
with Surface-related categories such as Conventions.
Meanwhile, similarity measures are more signi cantly correlated
with Content-level categories such as Reasoning.</p>
      <sec id="sec-22-1">
        <title>4.3 Coarse-grained Revision Purposes</title>
        <p>When annotating the dataset, the next step after aligning
sentences between drafts can be annotating the revisions
with coarse revision purposes, which means distinguishing
between Surface and Content. A potential use-case for this
level of annotation might be to applications where there is
more time available to do more dataset annotation, but a
higher degree of accuracy is desired than when annotating
for ne-grained revision purposes.</p>
        <p>Examining the number of coarse-grained revision purposes
per essay, the number of Surface revisions is not signi cantly
correlated with grade change in any rubric category. The
number of Content revisions has a Pearson correlation of r =
0:2588 (p = 0:0161) with Reasoning rubric grade change,
r = 0:2512 (p = 0:0917) with Precision grade change and
r = 0:2235 (p = 0:0386) with uency grade change. No
other rubric categories show a signi cant correlation of p
0:05. This is the rst level of annotation where revision
counts are signi cantly correlated with a Content-related
rubric category (Reasoning).</p>
      </sec>
    </sec>
    <sec id="sec-23">
      <title>Grade Reasoning Precision Fluency</title>
      <p>Reasoning grade change is signi cantly correlated with the
number of Content revisions, but is not correlated with the
total number of revisions as shown in Table 7. Meanwhile,
the number of Surface revisions is not signi cantly
correlated with any grade change, including in rubric categories
that correspond to ne-grained surface-level revision
purposes (i.e. Conventions, Fluency and Organization). The
number of Content revisions is signi cantly correlated with
Fluency, despite Fluency being a surface-level revision
purpose.</p>
      <p>New patterns arise when using cosine similarity but
controlling for the number of Surface or Content revisions. When
controlling for just for the number of Content revisions, the
correlation between cosine similarity (calculated as the
average of sentence pairs' embeddings' similarities) and grade
change is r = 0:2543 (p = 0:0181). When controlling for
the count of Surface revisions, the correlation is r = 0:2665
(p = 0:0131). Controlling for the number of Surface
revisions also results in signi cant correlation of similarity to
Claims grade and Reasoning grade. Controlling for the
number of Content revisions gives signi cant correlation of
similarity to Thesis grade and Rebuttal grade. See Table 10
for all signi cant correlations between cosine similarity and
rubric categories when controlling for number of Surface or
Content revisions.</p>
      <p>At this level, similarity measures are signi cantly correlated
with several content-oriented rubric categories such as
thesis and reason, as well as average rubric grade. Meanwhile,
the number of content revisions is signi cantly correlated
with two content-oriented categories (reasoning and
precision) and one surface-oriented category ( uency). The
number of surface revisions is not signi cantly related to any
rubric category grade change. As such, the similarity
measure seems to be doing the best at capturing deeper,
content</p>
    </sec>
    <sec id="sec-24">
      <title>Grade</title>
      <p>Thesis
Claim
Reasoning
Rebuttal
Average</p>
    </sec>
    <sec id="sec-25">
      <title>Control for</title>
      <p>Surface Count
r p
-0.2369
-0.2877
-0.4141
-0.2543
0.0001
0.0181
level changes and the number of content revisions captures
a general view of changes. However, the similarity measure
is also signi cantly correlated with average category grade
change, unlike the number of surface, content or total
revisions. As a result, at this level of annotation, the best overall
understanding of a student's revision pattern as a predictive
measure for grade change would be to calculate the number
of content revisions and the average sentence-level cosine
similarity, with or without controlling for number of content
revisions.</p>
      <sec id="sec-25-1">
        <title>4.4 Fine-grained Revision Purposes</title>
        <p>Labeling revisions for ne-grained revision purposes,
although more intensive and giving lower annotator
agreement, provides more information when correlating with grade
change. For counts of ne-grained revision purposes, the
results for all signi cant (p 0:05) correlations with rubric
grade changes are summarized in Table 11. Due to the great
number of combinations of 12 rubric categories and 9
revision purposes, only correlations between corresponding
categories (for instance, between Organization grade category
and Organization revision purpose) and correlations with
Average and Total rubric grade change were performed.
Only three revision purposes are signi cantly correlated with
grade change in their relevant rubric categories: Claims
(r = 0:3195; p = 0:0027), Evidence (r = 0:2578; p = 0:0166)
and General Content with Reasoning rubric category (r =
0:3226, p = 0:0024). Further, the Evidence revision purpose
is signi cantly correlated with Average rubric grade change
(r = 0:2642; p = 0:0140). This is the rst signi cant
correlation between any variety of revision count and Average
rubric grade change.</p>
        <p>The next test that was performed involved considering just
the correlation between a speci c rubric category and cosine
similarity between sentence pairs revised for a speci c
negrained revision purpose, when controlling for the number
of occurrences of that revision purpose. For instance, when
considering the subset of essays that contain at least one
sentence revised for Fluency, the correlation between
Fluency rubric grade change and average tf-idf cosine similarity
of old-new sentence pairs revised for Fluency is r = 0:1988
(p = 0:0790) when controlling for the number of Fluency
revisions. However, no correlations were found to be signi
cant in this complicated test. As such, this very high degree
of detail focusing on speci c rubric categories and related
revision purposes is not useful. Perhaps the test is so
focused on minuscule portions of essays that it loses sight of
the context in which the revision is situated. For instance,
adding a Reasoning sentence to a very long essay may make
a very small di erence to the overall comprehensibility of
the essay's reasoning, while adding a Reasoning sentence to
a very short essay may help to bridge a hole in the reasoning
that was caused by the essay's brevity and failure to give
detailed explanations. As such, this individual sentence may
contribute less to the rst essay's grade change than it does
to the second essay's grade change. Another consideration
about this test is that not all revision purposes occur in all
essays (see Table 4 for details), so the sample size of essays
included in these tests is often smaller than the full 86 essays
included in all other tests.</p>
        <p>At this level of annotation where the essays have been
annotated for ne-grained revision purposes, the most useful
indicator of grade change appears to be the counts of
revision purposes. However, the rubric categories (Reasoning,
Precision and Fluency) that are signi cantly correlated with
the counts of revision purposes are already signi cantly
correlated with other tests that do not require ne-grained
revision purposes. As a result, when looking to use revisions to
predict grade change, the additional e ort to annotate
negrained revision purposes instead of coarse-grained ones may
not be worthwhile.</p>
      </sec>
      <sec id="sec-25-2">
        <title>5. DISCUSSION</title>
        <p>Generally, the similarity measure is more signi cantly
correlated with Content-related rubric category grade change
and with Average rubric grade change, while numbers of
revisions are more signi cantly related with Surface-related
rubric category grade change. This is the case in Sections
4.1, 4.2 and 4.3. Signi cant correlations present at high
degrees of annotation, such as those between ne-grained
revision purposes and rubric categories in Section 4.4, tend
to be present at coarser degrees of annotation as well. As
such, when aiming to predict grade change, the most
productive level of annotation may be aligned sentences like in
Section 4.2 or coarse-grained revision purposes as in Section
4.3. These levels of analysis capture signi cant correlations
between revision and a wide range of the di erent rubric
categories for essays in this dataset, with revision counts
capturing Surface-related rubric category grade change and
similarity measures capturing Content-related and overall
average change.</p>
        <p>A caveat associated with this conclusion is that the
ArgRewrite corpus is entirely human-annotated and all
sentence alignments and revision purpose labels are the gold
standard between two annotators. However, many datasets
and applications do not have detailed human annotation
available. As such, the accuracy of this dataset's sentence
alignments and revision purpose annotations is at the upper
bound of possible accuracy. The correlations between
revision counts/purposes and rubric categories are, therefore, an
upper bound that may not be possible in datasets that have
been annotated by a classi er. This caveat also holds for
the similarity measure in cases where the measure is being
calculated on data that has some sort of annotation, such as
sentence alignments for average sentence-level similarity.
A second caveat is that this analysis does not take into
account that some students scored higher than others on
the rst draft. For instance, a student who nearly receives
a perfect score on the rst draft has little room to
improve whereas a student with a low score initially has
ample improvement opportunities. Potential ways to combat
this issue would be using some variety of corrected learning
gain score or nding the correlation between similarity score
and second draft score after controlling for rst draft score.
Lastly, it may be worthwhile to apply a post-hoc control
such as Bonferroni correction to the signi cance tests.</p>
      </sec>
      <sec id="sec-25-3">
        <title>6. CONCLUSION</title>
        <p>This work indicates that tf-idf embedding cosine similarity
captures overall essay grade change and essay revisions that
lead to rubric grade change in Content-related categories,
while revision purposes capture change in more
Surfaceoriented rubric categories. Future work is needed to
demonstrate whether this pattern extends to additional datasets,
particularly datasets where the sentence alignment and
revision purposes have been automatically labeled by a
classi er. Further, the dataset in this analysis contained only
argumentative essays about self-driving cars, so further work
would need to examine these ndings for datasets with other
writing styles and topics.</p>
      </sec>
      <sec id="sec-25-4">
        <title>7. ACKNOWLEDGMENTS</title>
        <p>Special thanks to the entire ArgRewrite group, and
particularly Tazin Afrin for her detailed comments on a draft of this
paper. This work is supported by National Science
Foundation (NSF) grant 1735752 to the University of Pittsburgh.
The opinions expressed are those of the authors and do not
represent the views of the Institute.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Afrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kashe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Olshefski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Litman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hwa</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Godley</surname>
          </string-name>
          .
          <article-title>E ective interfaces for student-driven revision sessions for argumentative writing</article-title>
          .
          <source>ACM Conference on Human Factors in Computing Systems (CHI)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Afrin</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Litman</surname>
          </string-name>
          .
          <article-title>Annotation and classi cation of sentence-level revision improvement</article-title>
          .
          <source>arXiv preprint arXiv:1909.05309</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bronner</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Monz</surname>
          </string-name>
          .
          <article-title>User edits classi cation using document revision histories</article-title>
          .
          <source>In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics</source>
          , pages
          <volume>356</volume>
          {
          <fpage>366</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Daxenberger</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Gurevych.</surname>
          </string-name>
          <article-title>A corpus-based study of edit categories in featured and non-featured wikipedia articles</article-title>
          .
          <source>In Proceedings of COLING 2012</source>
          , pages
          <fpage>711</fpage>
          {
          <fpage>726</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Faigley</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Witte</surname>
          </string-name>
          . Analyzing revision.
          <source>College composition and communication</source>
          ,
          <volume>32</volume>
          (
          <issue>4</issue>
          ):
          <volume>400</volume>
          {
          <fpage>414</fpage>
          ,
          <year>1981</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Fleiss</surname>
          </string-name>
          .
          <article-title>Measuring nominal scale agreement among many raters</article-title>
          .
          <source>Psychological bulletin</source>
          ,
          <volume>76</volume>
          (
          <issue>5</issue>
          ):
          <fpage>378</fpage>
          ,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. Y.</given-names>
            <surname>Yeung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zeldes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Reznicek</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Ludeling, and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Webster</surname>
          </string-name>
          .
          <article-title>Cityu corpus of essay drafts of english language learners: a corpus of textual revision in second language writing</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>49</volume>
          (
          <issue>3</issue>
          ):
          <volume>659</volume>
          {
          <fpage>683</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Max</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Wisniewski. Mining</surname>
          </string-name>
          naturally
          <article-title>-occurring corrections and paraphrases from wikipedia's revision history</article-title>
          .
          <source>In LREC</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          .
          <article-title>Crowdsourcing a wikipedia vandalism corpus</article-title>
          .
          <source>In Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <volume>789</volume>
          {
          <fpage>790</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Rehurek</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Sojka</surname>
          </string-name>
          .
          <article-title>Software Framework for Topic Modelling with Large Corpora</article-title>
          .
          <source>In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks</source>
          , pages
          <volume>45</volume>
          {
          <fpage>50</fpage>
          ,
          <string-name>
            <surname>Valletta</surname>
          </string-name>
          , Malta, May
          <year>2010</year>
          . ELRA. http://is.muni.cz/publication/884893/en.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          and
          <string-name>
            <given-names>I.</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R. D.</given-names>
            <surname>Roscoe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Jacovina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. K.</given-names>
            <surname>Allen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. S.</given-names>
            <surname>McNamara</surname>
          </string-name>
          .
          <article-title>Toward revision-sensitive feedback in automated writing evaluation</article-title>
          .
          <source>In EDM</source>
          , pages
          <volume>628</volume>
          {
          <fpage>629</fpage>
          .
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Shibani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Knight</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. Buckingham</given-names>
            <surname>Shum</surname>
          </string-name>
          .
          <article-title>Understanding students' revisions in writing: from word counts to the revision graph</article-title>
          .
          <source>Technical report, Technical report, Connected Intelligence Centre</source>
          , University of Technology Sydney,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Halfaker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kraut</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Hovy</surname>
          </string-name>
          .
          <article-title>Identifying semantic edit intentions from revisions in wikipedia</article-title>
          .
          <source>In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <year>2000</year>
          {
          <year>2010</year>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zesch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wojatzki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Scholten-Akoun</surname>
          </string-name>
          .
          <article-title>Task-independent features for automated essay grading</article-title>
          .
          <source>In Proceedings of the tenth workshop on innovative use of NLP for building educational applications</source>
          , pages
          <volume>224</volume>
          {
          <fpage>232</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , H. B.
          <string-name>
            <surname>Hashemi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hwa</surname>
            , and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Litman</surname>
          </string-name>
          .
          <article-title>A corpus of annotated revisions for studying argumentative writing</article-title>
          .
          <source>In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>1568</fpage>
          {
          <fpage>1578</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Litman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H. B.</given-names>
            <surname>Hashemi</surname>
          </string-name>
          .
          <article-title>Argrewrite: A web-based revision assistant for argumentative writings</article-title>
          .
          <source>In Proceedings of the 2016</source>
          conference
          <article-title>of the north american chapter of the association for computational linguistics:</article-title>
          <source>Demonstrations</source>
          , pages
          <volume>37</volume>
          {
          <fpage>41</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Litman</surname>
          </string-name>
          .
          <article-title>Annotation and classi cation of argumentative writing revisions</article-title>
          .
          <source>Grantee Submission</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>