<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Assessment of English CEFR Levels Using BERT Embeddings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Veronica Juliana Schmalz</string-name>
          <email>veronicajuliana.schmalz@kuleuven.be</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessio Brutti</string-name>
          <email>brutti@fbk.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>. Free University of Bozen-Bolzano</institution>
          ,
          <addr-line>Bolzano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>. KU Leuven, imec research group itec</institution>
          ,
          <addr-line>Kortrijk</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The automatic assessment of language
learners’ competences represents an
increasingly promising task thanks to recent
developments in NLP and deep learning
technologies. In this paper, we propose the
use of neural models for classifying
English written exams into one of the
Common European Framework of Reference
for Languages (CEFR) competence levels.
We employ pre-trained Bidirectional
Encoder Representations from Transformers
(BERT) models which provide efficient
and rapid language processing on account
of attention-based mechanisms and the
capacity of capturing long-range sequence
features. In particular, we investigate on
augmenting the original learner’s text with
corrections provided by an automatic tool
or by human evaluators. We consider
different architectures where the texts and
corrections are combined at an early stage,
via concatenation before the BERT
network, or as late fusion of the BERT
embeddings. The proposed approach is
evaluated on two open-source datasets: the
English First Cambridge open language
Database (EFCAMDAT) and the
Cambridge Learner Corpus for the First
Certificate in English (CLC-FCE). The
experimental results show that the proposed
approach can predict the learner’s
competence level with remarkably high accuracy,
in particular when large labelled corpora
are available. In addition, we observed
that augmenting the input text with
corrections provides further improvement in the
automatic language assessment task.</p>
      <p>Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).</p>
      <p>
        Finding a system which objectively evaluates
language learners’ competences is a daunting task.
Several aspects need to be considered, including
both subjective factors, like age, native language,
cognitive capacities of the learner, and
learningrelated factors, for example the amount and type
of received linguistic input
        <xref ref-type="bibr" rid="ref10 ref11 ref20 ref3 ref5">(James, 2005; Chapelle
and Voss, 2008; Jang, 2017)</xref>
        . Indeed, language
competences are not holistic, but concern
different domains, so that considering the mere formal
correctness of learners’ language has been shown
not to represent a proper assessment procedure
        <xref ref-type="bibr" rid="ref16 ref4 ref4 ref9">(Roever and McNamara, 2006; Harding and
McNamara, 2017; Chapelle, 2017)</xref>
        . Moreover,
human evaluators, despite having to adhere to a
predefined scale and guidelines, such as the CEFR
(Council of Europe, 2001), have proved to be
biased
        <xref ref-type="bibr" rid="ref12">(Karami, 2013)</xref>
        and inaccurate
        <xref ref-type="bibr" rid="ref7">(Figueras,
2012)</xref>
        . For these reasons, new language testing
methods and tools have been developed.
Current state-of-the-art models, such as
Transformers, allow to process numerous and complex
linguistic data efficiently and rapidly, by means of
attention-based mechanisms and deep neural
networks that capture the relevant features for the
targeted task. However, the creation and access to
necessary language examination resources
including annotations and metadata appear to date
limited. In this paper, we propose using a series of
BERT-base models to automatically assign CEFR
levels to language learners’ exams.
      </p>
      <p>
        Our aim is examining the possibility of
providing the system with previously generated
corrections, either by humans or automatically with a
language checker. Additionally, we want to
analyse the impact of the amount of data on the
accuracy of the model in the classification of
written exams taken from the English First
Cambridge Open Language Database (EFCAMDAT)
        <xref ref-type="bibr" rid="ref8">(Geertzen et al., 2013)</xref>
        and the Cambridge Learner
Corpus for the First Certificate in English
(CLCFCE)
        <xref ref-type="bibr" rid="ref21">(Yannakoudakis et al., 2011)</xref>
        . In this way,
a significant turning point could be made both in
improving the functioning of these automatic
systems and in the future collection of data from other
languages.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        Automatic language assessment methods concern
the creation of fast, effective, unbiased and
crosslinguistically valid systems that can both
simplify assessment and render it objective. However,
achieving such results represents a complex task
that researchers have been addressing for years
while experimenting with several methodologies
and techniques. The first developed tools used to
mainly deal with written texts and exploited
Partsof-Speech (PoS) tagging to grade students’ essays
        <xref ref-type="bibr" rid="ref2 ref8">(Burstein et al., 2013)</xref>
        , and latent semantic
analysis to evaluate the content, providing also short
feedback
        <xref ref-type="bibr" rid="ref13">(Landauer, 2003)</xref>
        . Advances in AI, NLP
and Automatic Speech Recognition (ASR) led to
the additional emergence of systems that assess
spoken language skills, such as the SpeechRater
        <xref ref-type="bibr" rid="ref20">(Xi et al., 2008)</xref>
        , which considers clarity of
expression, pronunciation and fluency. To date,
several other automatic language assessment tools
are applied in the domain of large scale testing,
for example Criterion
        <xref ref-type="bibr" rid="ref1">(Attali, 2004)</xref>
        , Project
Essay Grade
        <xref ref-type="bibr" rid="ref18">(Wilson and Roscoe, 2020)</xref>
        , MyAccess!
        <xref ref-type="bibr" rid="ref20 ref3 ref5">(Chen and Cheng, 2008)</xref>
        and Pigai
        <xref ref-type="bibr" rid="ref22">(Zhu, 2019)</xref>
        .
The first can detect grammatical and usage-based
errors, as well as punctuation mistakes,
providing also feedback. However, it requires being
trained on the specific topics to assess. The
second system exploits a training set of human-scored
essays to score unseen texts, evaluating diction,
grammar and complexity from statistical and
linguistic models. Similarly, MyAccess!, calibrated
with a large number of essays, can score
learners’ texts and measure advanced features such as
syntactic and lexical complexity, content
development and word choice, providing detailed
feedback. On the contrary, Pigai, exploits NLP to
compare the essays submitted by students with
those contained in its corpora, measuring the
distance between the two
        <xref ref-type="bibr" rid="ref22">(Zhu, 2019)</xref>
        . Despite the
extreme efficiency of these tools, to perform
accurately they generally need large amounts of
labelled and human-corrected training data.
Furthermore, a standard scale is needed, which can be
extended between different groups of learners. In
addition, powerful computational resources, and
in certain cases, significant memory, are required.
All these elements together constitute
fundamental pre-requisites which can be difficultly fulfilled.
For this reason, we present a distinct approach
to the previous ones which, starting from
different amounts of students’ original texts, provides a
classification within the different CEFR levels
exploiting BERT-base models and subsidiary
corrections.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Proposed Approach</title>
      <p>
        The approach we propose for the automatic
assessment of the language competences of adult
English language learners is based on the use of
Transformer-type architectures performing
multiclass classification. Among these, BERT-based
models, characterised by efficient parallel training
and the capacity of capturing long-range sequence
features, distinguish themselves for their size and
amount of training data
        <xref ref-type="bibr" rid="ref17">(Vaswani et al., 2017)</xref>
        .
Being pre-trained on generic large corpora, with
Masked Language Modelling (MLM) and Next
Sentence Prediction (NSP) strategies, they can be
conveniently employed in a wide range of tasks,
including text classification, language
understanding and machine translation.
      </p>
      <p>
        The models we use for our experiments are
grounded on the BERT-base-uncased architecture,
part of the Hugging Face Transformers Library
released in 2019
        <xref ref-type="bibr" rid="ref19">(Wolf et al., 2020)</xref>
        and inspired by
BERT
        <xref ref-type="bibr" rid="ref6">(Devlin et al., 2018)</xref>
        from Google Research,
that encodes input texts into low-dimensional
embeddings. Our baseline model maps these compact
representations into the CEFR levels using a
network with two fully connected layers. Fig. 1(a)
graphically represents the architecture. Note that
this approach requires training the final classifier
only. Retraining or fine-tuning the BERT model
would probably require very large datasets which
are not always available for this task. In order to
augment the input text with corrections (either
automatic or human) we investigate two possible
directions. The first one (Fig. 1(b)) concatenates the
two texts and applies the pre-trained BERT model.
The resulting embeddings are expected to encode
the information related to both texts. Conversely,
the second architecture extracts individual
embeddings for the original texts and the corrected ones.
(a)
(b)
(c)
These are then merged and processed by the
classifier, as shown in Fig. 1(c).
      </p>
      <p>We resort to these types of models to be able to
efficiently process texts capturing long-range
sequence features thanks to parallel word-processing
and self-attention mechanisms. Regardless of the
length of the texts, the architecture should be,
indeed, able to accurately categorise the
examinations according to the CEFR A1, A2, B1, B2
and C1 levels of competence. These, in fact, are
fed to the model as labels during the training
together with single contextual embeddings, or
concatenated ones if corrections are included. Note
that we do not provide the model with any
indication about the types of errors in the original text.
This information is directly extracted by the model
when processing the original text together with its
corrected version.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Analysis</title>
      <p>We evaluate the architectures described above,
using both automatic and human corrections, on
two English open-source datasets: EFCAMDAT
and CLC-FCE. We also experiment varying the
amount of training material. The performance of
the models is measured in terms of weighted
classification accuracy.
4.1</p>
      <sec id="sec-4-1">
        <title>EFCAMDAT Dataset</title>
        <p>
          The EFCAMDAT dataset constitutes one of the
largest language learners datasets currently
available
          <xref ref-type="bibr" rid="ref8">(Geertzen et al., 2013)</xref>
          . The version we use
contains 1,180,310 essays submitted by adult
English learners from more than 172 different
nationalities, covering 16 distinct levels compliant with
the CEFR proficiency ones. Each essay has been
corrected and evaluated by language instructors; in
addition to the original texts, their corrected
versions and annotated errors are also included.
        </p>
        <p>We considered a sub-set of the dataset
comprising 100,000 tests. Table 1 reports the
distribution of the exams across the different CEFR levels,
including also the average numbers of violations
identified by both humans evaluators and the
automatic tool, normalized by the average text length.
Note that the average errors per word decrease as
the level of competence increases. Observe also
that the automatic errors tend to be more numerous
than the human ones, in particular for low
competence levels. We use the official test partition
composed of 1,447 essays. The development set is a
20% subset of the training set.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>CLC-FCE Dataset</title>
        <p>
          The CLC-FCE dataset is a collection of texts
produced by adult learners for English as a Second
or Other Language (ESOL) examinations from the
First Certificate in English (FCE) written exam
to attest a B2 CEFR level
          <xref ref-type="bibr" rid="ref21">(Yannakoudakis et al.,
2011)</xref>
          . The learners’ productions, consisting of
two texts, have been evaluated with a score
between 0 and 5.3 and the errors have been classified
in 77 classes. Following the guidelines of the
authors, the average score of the two texts has been
mapped to CEFR levels, as shown in Table 2. Note
that only 4 levels are available in this dataset and
that the labels do not uniformly match the ones
present in EFCAMDAT. Table 2 reports also the
distributions of the texts across the 4 classes with
the error partitions. We notice that, in this case,
levels
        </p>
        <p>n. exams
A1
A2
B1
B2
C1
manual errors have been annotated more in
detail and they are indeed more numerous than the
automatic ones. In general, the number of
errors is higher than what observed in EFCAMDAT.
Also for this corpus the average amount of errors
per word, both automatic and manual, decreases
as the level increases. The total number of texts
within the corpus is 2,469. We employed a data
partition according to which 2,017 examinations
constituted the training set, whereas the
remaining 194 constituted the test set. Differently, 10%
of the training material represented the validation
set. From the entire corpus we had to exclude
10 texts since they were not provided with an
assigned score. Despite its small size, CLC-FCE
represents an important resource given its
systematic analysis of errors and the human corrections
provided.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>LanguageTool</title>
        <p>
          In both datasets, the content written by language
learners varies according to the levels of
competence they were supposed to demonstrate. In
addition to the human corrections provided with the
data, we have generated automatic corrections
using LanguageTool (Miłkowski, 2010), a language
checker capable of detecting grammatical,
syntactical, orthographic and stylistic errors to
automatically correct texts of different nature and length
          <xref ref-type="bibr" rid="ref15">(Naber and others, 2003)</xref>
          . The automatic checker
is based on surface text processing, does not use a
deep parser and does not require a fully formalised
grammar. By means of this, we have applied the
pre-defined rules for the English language to the
learners’ essays, generating new correct texts for
EFCAMDAT and for CLC-FCE. These were used
as additional input data for the experiments.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4 Implementation Details</title>
        <p>
          Our models have been implemented using
Keras and Hugging-Face’s pre-trained
BERTbase-uncased architecture
          <xref ref-type="bibr" rid="ref19">(Wolf et al., 2020)</xref>
          . The
models’ encoder module, consisting of a
MultiHead Attention and Feed Forward component,
receives as inputs the original learners’ exams,
together with additional possible human or
automatic corrections. The transformed contextual
embeddings are obtained applying Global
Average Pooling to the outputs of the pre-trained frozen
BERT Head. The classifier consists of a Dense
layer of 768 units, with activation function ReLu
and a Dropout rate of 0.2, followed by another
Dense layer with less units, 128, and the same
activation function and Dropout rate1.
        </p>
        <p>Lastly, the output layer consists of a Dense layer
with Softmax as activation function and the
models’ final logits correspond to the different CEFR
levels within which the texts are respectively
clas1https://www.kaggle.com/akensert/bert-base-tf2-0-nowhuggingface-transformer
N. Exams text only
sified. The selected loss is the Sparse Categorical
Cross-entropy and the evaluation metric is the
accuracy. The model is trained using Adam as
optimizer with learning rate 10− 5 for EFCAMDAT
and 10− 4 for CLC-FCE. The batch size is 32 and
the input text maximum length is set to 450 for
EFCAMDAT and 512 for CLC-FCE. These
hyperparameters were optimized on the related
development sets.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experimental Results</title>
      <p>Table 3 reports the classification accuracy on the
EFCAMDAT test set using the proposed
architectures in Fig. 1. Note that although EFCAMDAT
features more than 1 million samples, we limit our
analysis to 100K texts, due to memory issues and
performance saturation. The results include also
variations in the amount of training material,
considering 10K and 50K training exams. These
subsets have been obtained sampling in a uniform way
the training set, therefore the distribution of exams
per class does not change.</p>
      <p>First of all, it is worth noting that the best
approach reaches an extremely high classification
accuracy (almost 98%). In addition, performance
almost saturates with 50K essays, while with only
10K training samples the accuracy is well above
95%. The use of corrections, concatenated with
the original text, provides some improvements
over the model with original texts only.
Automatic corrections seem to be more effective with
less training data, while manual annotations
outperform the baseline with larger training sets. The
latter can, indeed, be more accurate, in particular
for high proficiency levels, but their inherited
variability makes the learning task more difficult. As
a consequence, more training samples are needed
to properly learn how to classify the input text.
This is evident in Table 3 where the manual
corrections are the worst for 10K samples, aligned
with the baseline with 50K training samples, and
the best performing when the 100K training texts
are used. Finally, the two-stream approach
averaging the BERT embeddings of the two texts, seems
to be less performing, although by a small margin.
Probably, the averaging operation does not
represent the most suitable one in this context as it tends
to generate embedding representations which are
somehow intermediate between those of the
original texts and those of the corrections and, hence,
less discriminative.</p>
      <p>Table 4 reports the results obtained on the
CLC-FCE corpus. With respect to EFCAMDAT,
this corpus is characterized by a smaller amount
of training material and by a less consistent
evaluation of the input text. These two facts lead to
a clear reduction of the classification accuracy, as
reported in the table. Due to the lower accuracy
and smaller size of the training set, the final
performance of each model has a certain degree of
variability, which dependents on the model
initialization and on the other random number generations
in the training process. Therefore, we performed
several runs varying the seed of the random
number generator. The average accuracy, as well as the
standard deviation, are also reported in Table 4.
model
text only
manual corr.
autom. corr.
two-streams</p>
      <p>accuracy
61.5% ± 2.0
60.7% ± 1.8
61.7% ± 1.8
61.5% ± 1.3</p>
      <p>Given the limited size of the training set, it is
not surprising to find rather similar results across
all the models. As expected, the manual
corrections are the worst performing, since they would
require large training sets to learn how to
handle human evaluations. It is worth pointing out
that the amount of errors per word in CLC-FCE
is much larger than in EFCAMDAT, which makes
the learning task even more complex.
Nevertheless, considering also the standard deviations, the
models based on automatic corrections are slightly
better than the model using the original texts only.
The two-streams model appears extremely close to
the concatenation model, but this could be related
to the fact that the overall accuracy is not that high.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this paper we presented an alternative approach
for the efficient and unbiased assessment of the
competences of English language learners using
pre-trained BERT-base models. We structured a
multi-class classification task to map the BERT
embeddings of written exams from the
EFCAMDAT and CLC-FCE open-source corpora to vfie
different levels of the CEFR scale. Alongside the
students’ original texts and the provided manual
corrections, we automatically generated additional
corrected versions with LanguageTool, a
multifaceted and versatile language checker . Thus, we
conducted several experiments varying both the
type and quantities of the models’ input, as well as
the typologies of models. Our results proved that
BERT-based architectures remarkably succeed in
classifying CEFR proficiency levels starting from
original texts, especially with numerically
significant data. Moreover, we observed that adding
automatic and manual corrections can contribute to
improve the quality of results.
Education Committee Council of Europe, Council for
Cultural Co-operation. 2001. Common European
Framework of Reference for Languages: learning,
teaching, assessment. Cambridge University Press.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Yigal</given-names>
            <surname>Attali</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Exploring the feedback and revision features of criterion</article-title>
          .
          <source>Journal of Second Language Writing</source>
          ,
          <volume>14</volume>
          :
          <fpage>191</fpage>
          -
          <lpage>205</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Jill</given-names>
            <surname>Burstein</surname>
          </string-name>
          , Joel Tetreault, and
          <string-name>
            <given-names>Nitin</given-names>
            <surname>Madnani</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The e-rater® automated essay scoring system</article-title>
          .
          <source>In Handbook of automated essay evaluation</source>
          , pages
          <fpage>77</fpage>
          -
          <lpage>89</lpage>
          . Routledge.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Carol A Chapelle</surname>
          </string-name>
          and
          <string-name>
            <given-names>Erik</given-names>
            <surname>Voss</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Utilizing technology in language assessment</article-title>
          .
          <source>Encyclopedia of language and education</source>
          ,
          <volume>7</volume>
          :
          <fpage>123</fpage>
          -
          <lpage>134</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Carol A</given-names>
            <surname>Chapelle</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Evaluation of technology and language learning. The handbook of technology and second language teaching and learning</article-title>
          , pages
          <fpage>378</fpage>
          -
          <lpage>392</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Chi-Fen Emily</surname>
          </string-name>
          Chen and Wei-Yuan Eugene Cheng Cheng.
          <year>2008</year>
          .
          <article-title>Beyond the design of automated writing evaluation: Pedagogical practices and perceived learning effectiveness in efl writing classes</article-title>
          .
          <source>Language Learning &amp; Technology</source>
          ,
          <volume>12</volume>
          (
          <issue>2</issue>
          ):
          <fpage>94</fpage>
          -
          <lpage>112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Neus</given-names>
            <surname>Figueras</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>The impact of the cefr</article-title>
          .
          <source>ELT journal</source>
          ,
          <volume>66</volume>
          (
          <issue>4</issue>
          ):
          <fpage>477</fpage>
          -
          <lpage>485</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Jeroen</given-names>
            <surname>Geertzen</surname>
          </string-name>
          , Theodora Alexopoulou,
          <string-name>
            <given-names>Anna</given-names>
            <surname>Korhonen</surname>
          </string-name>
          , et al.
          <year>2013</year>
          .
          <article-title>Automatic linguistic annotation of large scale l2 databases: The ef-cambridge open language database (efcamdat)</article-title>
          .
          <source>In Proceedings of the 31st Second Language Research Forum. Somerville</source>
          , MA: Cascadilla Proceedings Project, pages
          <fpage>240</fpage>
          -
          <lpage>254</lpage>
          . Citeseer.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Luke William</given-names>
            <surname>Harding and Tim McNamara</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Language assessment: The challenge of elf. In Routledge Handbook of English as a Lingua Franca</article-title>
          . Routledge.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Carl</given-names>
            <surname>James</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Contrastive analysis and the language learner. Linguistics, language teaching and language learning</article-title>
          ,
          <volume>120</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Eunice</given-names>
            <surname>Eunhee Jang</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Cognitive aspects of language assessment</article-title>
          .
          <source>Language Testing and Assessment</source>
          ,, pages
          <fpage>163</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Hossein</given-names>
            <surname>Karami</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The quest for fairness in language testing</article-title>
          .
          <source>Educational Research and Evaluation</source>
          ,
          <volume>19</volume>
          (
          <issue>2-3</issue>
          ):
          <fpage>158</fpage>
          -
          <lpage>169</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Thomas K</given-names>
            <surname>Landauer</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Automatic essay assessment. Assessment in education: Principles, policy</article-title>
          &amp; practice,
          <volume>10</volume>
          (
          <issue>3</issue>
          ):
          <fpage>295</fpage>
          -
          <lpage>308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Marcin</given-names>
            <surname>Miłkowski</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Developing an open-source, rule-based proofreading tool</article-title>
          .
          <source>Software: Practice and Experience</source>
          ,
          <volume>40</volume>
          (
          <issue>7</issue>
          ):
          <fpage>543</fpage>
          -
          <lpage>566</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Naber</surname>
          </string-name>
          et al.
          <year>2003</year>
          .
          <article-title>A rule-based style and grammar checker</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Carsten</given-names>
            <surname>Roever and Tim McNamara</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Language testing: The social dimension</article-title>
          .
          <source>International Journal of Applied Linguistics</source>
          ,
          <volume>16</volume>
          (
          <issue>2</issue>
          ):
          <fpage>242</fpage>
          -
          <lpage>258</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
          <string-name>
            <surname>Łukasz Kaiser</surname>
            , and
            <given-names>Illia</given-names>
          </string-name>
          <string-name>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Joshua</given-names>
            <surname>Wilson and Rod D Roscoe</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Automated writing evaluation and feedback: Multiple metrics of efficacy</article-title>
          .
          <source>Journal of Educational Computing Research</source>
          ,
          <volume>58</volume>
          (
          <issue>1</issue>
          ):
          <fpage>87</fpage>
          -
          <lpage>125</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Wolf</surname>
          </string-name>
          , Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison,
          <string-name>
            <given-names>Sam</given-names>
            <surname>Shleifer</surname>
          </string-name>
          , et al.
          <year>2020</year>
          .
          <article-title>Transformers: State-of-theart natural language processing</article-title>
          .
          <source>In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          , pages
          <fpage>38</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Xiaoming</given-names>
            <surname>Xi</surname>
          </string-name>
          , Derrick Higgins, Klaus Zechner, and David M Williamson.
          <year>2008</year>
          .
          <source>Automated scoring of spontaneous speech using speechratersm v1. 0. ETS Research Report Series</source>
          ,
          <year>2008</year>
          (2):
          <fpage>i</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Helen</given-names>
            <surname>Yannakoudakis</surname>
          </string-name>
          , Ted Briscoe, and
          <string-name>
            <given-names>Ben</given-names>
            <surname>Medlock</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A new dataset and method for automatically grading esol texts. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies</article-title>
          , pages
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Wenxin</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A study on the application of automated essay scoring in college english writing based on pigai</article-title>
          .
          <source>In 2019 5th International conference on social science and higher education (ICSSHE</source>
          <year>2019</year>
          ), pages
          <fpage>451</fpage>
          -
          <lpage>454</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>