<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Reading Comprehension Quiz Generation using Generative Pre-trained Transformers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ramon Dijkstra</string-name>
          <email>ramon.dijkstra@hotmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zu¨lku¨f Genc¸</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Prosus</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Amsterdam</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recent advances in AI have resulted in large pre-trained language models with superior performance on text generation tasks, prompting the question of whether we can use them to generate educationally useful text completions. This holds the potential to generate relevant quizzes for any educational text, greatly complementing current formative and summative tests from education professionals. We explore pre-trained language models for quiz generation on reading comprehension texts and propose EduQuiz, an end-to-end quiz generator based on a GPT-3 model fine-tuned on text-quiz pairs, able to generate a complete multiple-choice question, with the correct and distractor answers. We observed that the majority of generated quizzes is reasonable, and that generation of high-quality distractors is more challenging than question and answer generation. More generally, while it may be too early to replace manually generated tests for summative feedback and grading with automatic quiz generation, EduQuiz already has potential value for formative feedback and to increase engagement during the learning phase by enhancing textbooks with assessments.</p>
      </abstract>
      <kwd-group>
        <kwd>Quiz Generation</kwd>
        <kwd>Question Generation</kwd>
        <kwd>Question Answering</kwd>
        <kwd>Distractor Generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>This paper investigates the problem of quiz generation on reading
comprehension texts: given any arbitrary educational text as input, can we automatically
generate complete quizzes for potential use as formative or summative feedback
in an educational setting?</p>
      <p>Chinese young people love their smart phones. . . . But now, a book has let them
forget about their phones. The book is called Secret Garden. It is a colouring book
for adults’. . . . Many young people bought this book. They like it so much that they
stop playing games or surfing the Internet on their phones. . . . Drawing the book can
help them become less nervous. In the past, people thought colouring books are only
for children. With Secret Garden getting more and more popular, many young adults
also begin to paint as a hobby.</p>
      <p>(a) Prompt EQG-Race #649 (shortened)
. . . begin to paint as a hobby. An author from UK who came up with the idea told
Beijing Morning that he wanted to let people forget about the stressful world for a
while. He also thought people would feel relaxed while doing the work. The artist was
very busy. He had little time to meet and paint with his friends.</p>
      <p>(b) Text completion (unconstrained)
. . . begin to paint as a hobby. Question: Why do many young people like to colour
the book? True answer: Because it is a way to relax themselves. False answer:
Because they think it is interesting. False answer: Because they like to show their
talent. False answer: Because they can earn much money.</p>
      <p>
        (c) EduQuiz quiz generation (template in bold)
example reading comprehension text, demonstrating high levels of text quality
and relatedness to the text prompt, but also that unconstrained text completion
is very “creative” and the relation to the text at hand may be lost very quickly.
We propose an end-to-end quiz generator, EduQuiz, based on a GPT-3 model
ifne-tuned on text-quiz pairs to complete an educationally relevant quiz template
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Figure 1(c) shows the output for the example input, in this case successfully
generating a relevant question, with a correct answer, and three diferent false
answers. This paper addresses the quiz generation problem head-on, trying to
answer the question: Can generative pre-trained transformers learn to generate
a reading comprehension quiz?
EduQuiz is created for students and teachers and can be seen as an initial step
towards fully automatic quiz generation. This could reduce the burden of
manually creating quizzes for teachers or educational content creators. Moreover,
students could use this tool to test their knowledge while learning from
textbooks. This is beneficial for students as asking exam-like questions to learners is
proven to be the way to test the real knowledge of learners [
        <xref ref-type="bibr" rid="ref10 ref2">2, 10</xref>
        ]. Besides, active
learning and testing knowledge after reading an educational text have shown to
be beneficial for learning [
        <xref ref-type="bibr" rid="ref1 ref22">1, 22</xref>
        ].
      </p>
      <p>
        Quiz generation is a complex problem, with each quiz consisting of a question,
a true answer, and several false answers that are closely related to the true
answer. We will call these false answers distractors. We performed our research
on the EQG-RACE dataset, which contains examination-like questions from the
original RACE dataset [
        <xref ref-type="bibr" rid="ref13 ref15">13, 15</xref>
        ]. On this dataset, we compared a general-purpose
model called Macaw-11b and task-specific fine-tuned GPT-3 models on the tasks
in Table 1 [
        <xref ref-type="bibr" rid="ref27 ref4">4, 27</xref>
        ].
      </p>
      <p>As shown in Table 1, we introduce the tasks of Step-Wise Quiz Generation
(SWQG) and End-to-End Quiz Generation (EEQG). SWQG generates a
question based on the context, uses both the context and generated question to
generate the corresponding answer, and uses the context, generated question and
answer to generate distractors. EEQG fulfills all of this in one go by generating
a quiz directly from the context.</p>
      <p>
        To evaluate the performances, we used the common metrics BLEU-4, ROUGE-L,
and METEOR, against the human reference ground-truth [
        <xref ref-type="bibr" rid="ref18 ref21 ref3">3, 18, 21</xref>
        ].
Our main contributions can be summarized as follows:
– We propose the end-to-end quiz generation as a key research problem with
a large potential impact on the education domain, helping both learners and
educators to increase learning efectiveness in every situation where
professional quizzes are not readily available, and thereby also contributing to the
enhancement of assessments in textbooks.
– We propose an end-to-end quiz generator based on GPT-3, EduQuiz, where
we observed that the majority of generated quizzes is reasonable, and that
generation of high-quality distractors is more challenging than question and
answer generation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
– We encourage other researchers to reproduce and expand our results and share
all the used data and code on Github.1
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        Quiz Generation combines Question Generation (QG), Question Answering (QA),
and Distractor Generation (DG). QG, QA, and DG all have a comparable
research history. Early models within these domains were rule-based [
        <xref ref-type="bibr" rid="ref11 ref20 ref24">11, 20, 24</xref>
        ].
Later, the paradigm switched to neural methods [
        <xref ref-type="bibr" rid="ref17 ref32 ref8">8, 17, 32</xref>
        ]. With the rise of
transformers, the focus within QG, QA, and DG switched again [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. As of
      </p>
      <sec id="sec-2-1">
        <title>1 https://github.com/RamonDijkstra/EduQuiz</title>
        <p>
          current, large pre-trained language models such as BERT, T5, and GPT-3 have
shown superior performances on these text generation tasks [
          <xref ref-type="bibr" rid="ref23 ref4 ref6">4, 6, 23</xref>
          ]. A
generalpurpose model called Macaw can perform QG, QA, and DG as it is trained on
these diferent angles [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ].
        </p>
        <p>
          Generating full quizzes for educational purposes has been performed on
fillin-the-blank questions, knowledge bases, and listening comprehension [
          <xref ref-type="bibr" rid="ref12 ref19 ref25 ref26">12, 19,
25, 26</xref>
          ]. Previous research also aimed to generate assessments from textbooks
[
          <xref ref-type="bibr" rid="ref28 ref7">7, 28</xref>
          ]. Within the educational domain, strictly generating a quiz based on an
educational text using large pre-trained language models has not been done
before. Closely related, the research by Lelkes et al. generates quiz questions,
answers, and distractors on news articles [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. In their approach, they rfist apply
Question-Answer Generation (QAG) and then DG. Related to quiz generation is
a system proposed by Khan et al. where the user can generate assessment content
by interactively rating generated text [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. This approach requires the user to
have the domain knowledge to specify exactly what it wants to generate whereas
we focus on full automation of this process. To the best of our knowledge, SWQG
has not been done before by concatenating QG, QA, and DG and no previous
research has been performed on EEQG within the educational domain.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <p>
        GPT-3 is a generative pre-trained transformer that can be fine-tuned to
downstream tasks using the API2 of OpenAI [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The API allows users to submit
training and validation data files as JSONL documents where each
training/validation instance contains a prompt-completion pair. OpenAI takes care of the
training itself. In our case, the prompt is an educational text, and the
completion a quiz. Previous research has shown that fine-tuning with instructions or
with a template is needed to perform unseen tasks [
        <xref ref-type="bibr" rid="ref31 ref9">9, 31</xref>
        ]. For end-to-end quiz
generation, we propose to use the template specified below as our completion.
We propose an end-to-end quiz generator based on GPT-3, EduQuiz, which is
ifne-tuned on text-quiz pairs to generate quizzes on never-seen-before texts.
GPT-3 comes in four diferent versions called Ada, Babbage, Curie, and Davinci
which contain 350M, 1.3B, 6.7B, and 175B parameters respectively3. We used
Curie as manual tests have shown that similar results can be achieved with a
smaller model when fine-tuning. Lastly, we kept the default fine-tuning
hyperparameters.
      </p>
      <p>End-to-End Quiz Generation Template
Question: . . .</p>
      <p>True answer: . . .</p>
      <p>False answer: . . .</p>
      <p>False answer: . . .</p>
      <p>False answer: . . .</p>
      <sec id="sec-3-1">
        <title>2 https://beta.openai.com/docs/guides/fine-tuning</title>
      </sec>
      <sec id="sec-3-2">
        <title>3 https://blog.eleuther.ai/gpt3-model-sizes/</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Setup</title>
      <p>The main question that is addressed in this paper is: Can generative pre-trained
transformers learn to generate a reading comprehension quiz? We aim to answer
this question by evaluating SWQG and EEQG. In this section, we will elaborate
on the dataset used, our evaluation method, and the tested models.
4.1</p>
      <sec id="sec-4-1">
        <title>Dataset</title>
        <p>
          We perform our experiments on the EQG-RACE dataset3 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. It is a processed
RACE dataset where only examination questions are kept. As we are focusing
on the education domain, this dataset is a good fit for our purposes. During
processing, Jia et al. removed the distractors from the data and only kept the
questions and answers. We combine the questions from EQG-RACE with the
original data in RACE4 to extract the distractors again.
        </p>
        <p>The EQG-RACE dataset contains 18,501 train, 1,035 validation, and 950 test
instances. After connecting it to the original RACE dataset, each data instance
consists of a reading comprehension text and a quiz, containing a question, one
true answer, and three distractors, as specified in the template above.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Automatic Evaluation</title>
        <p>
          For automatic evaluation, we follow previous work on QG, QA, and DG and
use the existing evaluation methods BLEU-4, ROUGE-L, and METEOR as our
metrics [
          <xref ref-type="bibr" rid="ref18 ref21 ref3">3, 18, 21</xref>
          ]. BLEU-4 measures the 4-gram similarity between a
prediction and ground truth instances [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. ROUGE-L measures the longest common
sub-sequence between the prediction and ground truth instances [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. METEOR
is similar in comparison to BLEU-4 but also takes synonyms, stemming, and
paraphrasing into account [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. For BLEU-4 and ROUGE-L we use the
Huggingface implementation5. As the newest version of the METEOR metric is not in
Huggingface, we used the implementation from the original website6.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Tested Models</title>
        <p>
          To test our approach from Section 3, we compared both a general-purpose model
called Macaw-11b, and task-specific fine-tuned GPT-3 models [
          <xref ref-type="bibr" rid="ref27 ref4">4, 27</xref>
          ] on SWQG
and EEQG. For SWQG, we concatenate the QG, QA, and DG configurations
of Macaw-11b to generate a quiz in a step-wise manner. Macaw-11b could not
be used for EEQG7. To perform SWQG with GPT-3, we have fine-tuned the
QG, QA, and DG models according to the diferent prompt-completion pairs
from the earlier Table 1. We again concatenate these models to perform SWQG.
        </p>
        <sec id="sec-4-3-1">
          <title>3 https://github.com/jemmryx/EQG-RACE</title>
          <p>4 https://www.cs.cmu.edu/∼ glai1/data/race/</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>5 https://huggingface.co/metrics</title>
          <p>6 https://www.cs.cmu.edu/∼ alavie/METEOR</p>
          <p>Models
Macaw-11b SWQG
GPT-3 SWQG
EduQuiz
BLEU-4 ROUGE-L METEOR</p>
          <p>We compare the SWQG models with EduQuiz, an end-to-end quiz generator
we proposed in Section 3. For clarification, the name of each experiment will be
attached to the models. For example, Macaw-11b SWQG is the general-purpose
Macaw-11b model configured for SWQG and GPT-3 SWQG is a GPT-3 model
ifne-tuned on SWQG.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experimental Results</title>
      <p>In this section, we will discuss the results of the experiments. We have performed
SWQG and EEQG with Macaw-11b and fine-tuned GPT-3 models. Table 2
shows the automatic evaluation of the generated quizzes by the diferent models.
From this table, we would expect that EduQuiz outperforms Macaw-11b SWQG
and GPT-3 SWQG. We do however see that the BLEU-4 score is remarkably
low for all tested models. The ROUGE-L and METEOR score are higher as the
template specified in Section 3 boosts the scores. The next example showcases
the BLEU-4 score for generated quizzes.</p>
      <p>Test instance ID: #46
Original quiz:
Question: What is Edell’s strongest character?
True answer: Liking to try new things.</p>
      <p>False answer: Liking to remember things.</p>
      <p>False answer: Liking to help others.</p>
      <p>False answer: Liking to think of problems diferently.</p>
      <p>Macaw-11b SWQG generated quiz (BLEU-4 = 0.0):
Question: The writer’s mother told her that if she wanted to be successful in life,
she should be prepared for any kind of change. What did Edell do to prepare for the
Asian MTV Awards?
True answer: she read the biographies of the guests
False answer: she practiced her English with her friends
False answer: she made a list of the names of the guests
False answer: she made a plan to talk with the guests</p>
      <sec id="sec-5-1">
        <title>7 Macaw-11b did have the configurations to perform EEQG but our experiments</title>
        <p>showed that only 14% of the generated quizzes contained all the elements of the
quiz. Therefore, we will not use Macaw-11b for EEQG.
GPT-3 SWQG generated quiz (BLEU-4 = 0.0):
Question: What advice does Edell give us?
True answer: Get ready for any opportunity in life.</p>
        <p>False answer: Try to live a colorful life.</p>
        <p>False answer: Take any chance that comes up.</p>
        <p>False answer: Explore your potential talents.</p>
        <p>EduQuiz-generated quiz (BLEU-4 = 0.0):
Question: What advice does Edell give to young people?
True answer: Try to get yourself well-prepared in life.</p>
        <p>False answer: Have a rich collection of CDs.</p>
        <p>False answer: Never miss an opportunity to learn ballet.</p>
        <p>False answer: Be a hostess of the Asian MTV Awards.</p>
        <p>The example shows generated quizzes and the original quiz. We see that the
generated quizzes from the diferent models are diferent from the original quiz.
Therefore, when automatically comparing n-gram similarity with BLEU-4, there
is little overlap with the original quiz and the BLEU-4 scores are 0.0. Thus, the
automatic evaluation is problematic. The generated quizzes seem reasonable but
the BLEU-4 scores are uninterpretable as they only represent n-gram similarity.
Therefore, it is hard to evaluate the quality of generated quizzes with these
automatic scores. From this example, we can see that Macaw-11b SWQG first
generates a sentence regarding the prompt before asking the question. Also, the
punctuation from the GPT-3 models seems better in comparison to the
Macaw11b SWQG model. Detailed analysis is needed to conclude which model generates
the highest quality quizzes.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Analysis of generated quizzes</title>
      <p>
        The experimental results showed that automatic evaluation on quiz generation
is limited in expressing the quality of generated quizzes. We have seen that
generated quizzes seem reasonable even when they have little overlap with the
original quiz. In this section, we will analyse the experimental results with human
evaluation, as this is the golden standard to evaluate natural language generation
tasks [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
      </p>
      <p>
        In our analysis, we rely on previously used human evaluation approaches in QG
and DG to evaluate generated quizzes [
        <xref ref-type="bibr" rid="ref13 ref33">13, 33</xref>
        ]. As shown in Table 3, the question,
answer, and distractors are all evaluated on whether they are grammatical and
lfuent. Furthermore, the question is also rated on whether it is relevant to the
passage and if the passage contains the answer. The answer is also rated on
whether the generated text contains correct information, and whether the answer
contains all information to answer the question properly. The distractors are
rated on whether they are coherent with the text and whether they can mislead
the learner to choose a wrong answer. Whenever there is true information in the
distractors, the distractors fail in their distracting ability. Each of the metrics in
Table 3 is binary rated. Binary rating the metrics will result in a hard cut-of.
whether the question is grammatical and fluent.
whether the question is semantic relevant to the
passage.
whether the question can be answered by the right
answer.
whether the answer is grammatical and fluent.
whether the answer contains correct information.
      </p>
      <p>whether the answer is a correct answer to the question.</p>
      <p>Fluency whether the distractors are grammatical and fluent.</p>
      <p>Coherence whether the distractors are relevant to the article and
the question.</p>
      <p>Distracting Ability whether the distractors can mislead the learner and if
there is no true information in the distractors.
Also, as we have built upon previous work, there is a slight overlap between the
metrics. Lastly, we experienced that the output of Macaw-11b is often diferent
from GPT-3 which made double-blind evaluation impossible.</p>
      <p>
        The human evaluation scores are an average score of three human annotators
on 100 test instances for that specific task. The annotators rated 82.0% of all
instances similarly. The Cohen kappa (κ ) scores between editorial judges 1-2,
2-3, and 1-3 on all rated instances were 0.63, 0.53, and 0.49 respectively [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
6.1
      </p>
      <sec id="sec-6-1">
        <title>Step-Wise Quiz Generation (SWQG)</title>
        <p>Table 4 shows that the quizzes generated by Macaw-11b SWQG mostly contain
grammatical and fluent language. On all other metrics, Macaw-11b SWQG shows
low performance. In contrast, GPT-3 SWQG generates quizzes that only lack
the validness of the answer (73.7%) and the distracting ability of the distractors
(53.7%). Overall, GPT-3 SWQG seems to perform better (89.3%) in comparison
to Macaw-11b SWQG (62.7%) on the total average score.
Macaw-11b SWQG often first summarizes the text before it asks the question.
GPT-3 SWQG generates stand-alone questions. The following example
showcases where Macaw-11b SWQG and GPT-3 SWQG both generated a
highquality quiz. We do see that Macaw-11b SWQG lacks interpunction as it misses
a capital letter and does not have a period at the end of the sentence. In
contrast, GPT-3 SWQG does this correctly. When the models generate a relevant
and answerable question, it directly becomes easier to generate a high-quality
quiz. This is intuitive as the step-wise quiz generation starts with the task of
creating a high-quality question. When the question is of low quality, the whole
quiz will be of low quality.</p>
        <p>The next example showcases that both models also generate quizzes of low
quality. Here, Macaw-11b SWQG generated an irrelevant question which leads to
an unusable quiz. GPT-3 SWQG asks a really easy question as the answer is
in the question. This tricked GPT-3 SWQG to generate irrelevant answers and
distractors. Macaw-11b SWQG and GPT-3 SWQG both sufer from repetition.
Especially Mascaw-11b SWQG repeats the same answer option multiple times.
#55 Question: How did Jocelyn disappear?</p>
        <p>True answer: She disappeared from the spot where she was playing.</p>
        <p>False answer: She disappeared when she was playing with her friends.
False answer: She disappeared when she was getting her bike.</p>
        <p>False answer: She disappeared from her grandmother’s apartment.</p>
        <p>Macaw-11b SWQG generated low-quality quizzes whereas GPT-3 SWQG
generated quizzes of high quality. For GPT-3 SWQG, there is room for improvement
on the validity of the answer and the distracting ability of the distractors.
6.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>End-to-End Quiz Generation (EEQG)</title>
        <p>The evaluation results in Table 4 showcase that EduQuiz generates quizzes of
comparable quality to GPT-3 SWQG on total average scores. The highest scores
in the columns are interleaved between GPT-3 SWQG and EduQuiz. For SWQG,
we needed to fine-tune three diferent models. To achieve the same results, we
could just fine-tune one GPT-3 model which reduces the cost by a factor of
three.</p>
        <p>Table 5 shows generated quizzes by EduQuiz. The test instances #82 and #379
showcase that EduQuiz generates quizzes of high quality. The generated quizzes
are diferent from the ground truth as multiple quizzes can be valid for the same
piece of text. The test instances #685, #388, and #55 showcase examples where
EduQuiz generated low-quality quizzes. The first example of low-quality quizzes
contains a question that is irrelevant and unanswerable. Therefore, the full quiz
is of low quality. In the second row, EduQuiz generated a false answer in the
place of a true answer as it has dificulties with the negation in the question.
The third example shows that EduQuiz tends to switch true and false answers.
The question and true answer are of high quality. However, the distractors also
contain true information so they fail in distracting ability as the distractors could
have been the true answer. Referring back to our human evaluation in Table 4,
EduQuiz has the most dificulties with generating valid answers (77.7%) and
generating distractors with a distracting ability (60.3%). EduQuiz sometimes
lists facts about the question instead of generating a quiz by clearly separating
the answer and distractors.</p>
        <p>EduQuiz generates quizzes of comparable quality (90.5%) to GPT-3 SWQG
(89.3%) and can generate a complete quiz in one pass rather than the three
inference steps required by GPT-3 SWQG. It still has dificulties generating
valid answers and distractors with a distracting ability.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Discussion &amp; Conclusions</title>
      <p>In this paper, we explored pre-trained language models for quiz generation on
reading comprehension texts and propose an end-to-end quiz generator, EduQuiz,
where we observed that the majority of generated quizzes is reasonable, and
that generation of high-quality distractors is more challenging than question
and answer generation. We have performed a comparative study on Step-Wise
Quiz Generation (SWQG) and End-to-End Quiz Generation (EEQG). We
performed automatic evaluation and analysed generated quizzes using human
evaluation. We proposed two new tasks for text generation in the educational domain,
SWQG, and EEQG. The first task is a concatenation of QG, QA, and DG. The
latter is the generation of a full quiz only based on the context. Here, GPT-3
SWQG and EduQuiz outperformed Macaw-11b SWQG. This can be explained
by the fact that this task is far more dificult and fine-tuning is needed
otherwise it will not work. Besides that, GPT-3 is a strong larger pre-trained language
model. GPT-3 SWQG and EduQuiz generated quizzes on the same quality level.
Over all the domains, it is remarkable that almost all our generations contain
lfuent language. Traditional NLP pipelines often resulted in mixed quality text
generation, but the large pre-trained language models seem to handle this very
well with a lot of variation and expressiveness over rigid filled-out templates.
EduQuiz generated questions that are relevant and answerable. The generated
answer contains correct information but is not always a valid answer to the
question. The generated distractors are coherent to the text but sometimes lack
distracting ability.</p>
      <p>There are some limitations to our research. The used models are mostly trained
on the English language. Therefore, they will not fully generalize to other
languages. Some manual tests have shown that the models can fullfil the trick in
another language but the used language is not that expressive in comparison
to English. Another limitation is that the fine-tuned GPT-3 models are costly.
However, once the model is trained, it can be used for text generation on the fly
with a far smaller completion cost. Moreover, the models are domain-specific and
only create comparable questions and answers to the training dataset. There are
two solutions to this problem. The first is to train the model on a really broad
domain so that it can generate quizzes on all educational domains. Another
solution is to create domain-specific quiz generators. The latter would probably
generate better results for each domain specifically but it comes with a cost.
Lastly, one could also argue that we have not made a fair comparison between
the models. GPT-3 is fine-tuned on the task whereas Macaw-11b is used of the
shelf.</p>
      <p>While our experimental results are very encouraging and the model generates
many useful quiz questions, our evaluation also reveals that not every quiz is
perfect yet and the quality is sometimes lower than human-generated quizzes.
This is issuing a call to caution to replace education professionals in particular
for summative feedback and grading. This is also suggesting clear directions to
further improve quiz generation for education, both directly by further improving
the model and training regime, and indirectly in terms of the exact use case
(current models may be helpful for formative rather than summative feedback),
introducing ways of filtering out “bad” quizzes (as we can generate multiple
candidates), or using it in a human-in-the-loop setting in an educator support
system.</p>
      <p>We hope to encourage other researchers to work on quiz generation as a research
ifeld with large potential impact on students and teachers, and with many
applied research opportunities. We aimed to set the first step towards replacements
of labor-intensive quiz generation by automatic quiz generation, thereby also
contributing to the enhancement of textbooks with assessments.</p>
      <sec id="sec-7-1">
        <title>Acknowledgments</title>
        <p>We thank the reviewers for their insightful comments. Kamps is funded in part by the
Netherlands Organization for Scientific Research (NWO CI # CISC.CC.016), and the
Innovation Exchange Amsterdam (POC grant). Views expressed in this paper are not
necessarily shared or endorsed by those funding the research.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biddle</surname>
          </string-name>
          , W.B.:
          <article-title>On asking people questions about what they are reading</article-title>
          . In:
          <article-title>Psychology of learning and motivation</article-title>
          , vol.
          <volume>9</volume>
          , pp.
          <fpage>89</fpage>
          -
          <lpage>132</lpage>
          ,
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Andre</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Does answering higher-level questions while reading facilitate productive learning?</article-title>
          <source>Review of Educational Research</source>
          <volume>49</volume>
          (
          <issue>2</issue>
          ),
          <fpage>280</fpage>
          -
          <lpage>318</lpage>
          (
          <year>1979</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Meteor: An automatic metric for mt evaluation with improved correlation with human judgments</article-title>
          .
          <source>In: Proceedings of the acl workshop on</source>
          intrinsic and
          <article-title>extrinsic evaluation measures for machine translation</article-title>
          and/or summarization, pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Brown</surname>
          </string-name>
          , T.B.,
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.:
          <article-title>Language models are few-shot learners</article-title>
          . arXiv preprint arXiv:
          <year>2005</year>
          .
          <volume>14165</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A coeficient of agreement for nominal scales</article-title>
          .
          <source>Educational and psychological measurement 20(1)</source>
          ,
          <fpage>37</fpage>
          -
          <lpage>46</lpage>
          (
          <year>1960</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Dresscher</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alpizar-Chacon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sosnovsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Generation of assessment questions from textbooks enriched with knowledge models</article-title>
          .
          <source>In: CEUR Workshop Proceedings</source>
          , vol.
          <volume>2895</volume>
          , pp.
          <fpage>45</fpage>
          -
          <lpage>59</lpage>
          ,
          <string-name>
            <surname>CEUR WS</surname>
          </string-name>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cardie</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Learning to ask: Neural question generation for reading comprehension</article-title>
          .
          <source>arXiv preprint arXiv:1705.00106</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fisch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Making pre-trained language models better few-shot learners</article-title>
          . arXiv preprint arXiv:
          <year>2012</year>
          .
          <volume>15723</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Hamaker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The efects of adjunct questions on prose learning</article-title>
          .
          <source>Review of educational research 56(2)</source>
          ,
          <fpage>212</fpage>
          -
          <lpage>242</lpage>
          (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Heilman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          :
          <article-title>Good question! statistical ranking for question generation</article-title>
          .
          <source>In: Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics</source>
          , pp.
          <fpage>609</fpage>
          -
          <lpage>617</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Y.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tseng</surname>
            ,
            <given-names>Y.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>Y.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          :
          <article-title>Tedquiz: automatic quiz generation for ted talks video clips to assess listening comprehension</article-title>
          .
          <source>In: 2014 IEEE 14Th international conference on advanced learning technologies</source>
          , pp.
          <fpage>350</fpage>
          -
          <lpage>354</lpage>
          , IEEE (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Eqg-race: Examination-type question generation</article-title>
          . arXiv preprint arXiv:
          <year>2012</year>
          .
          <volume>06106</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Almeida</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Generate: A nlg system for educational content creation</article-title>
          .
          <source>In: EDM</source>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          , E.: Race:
          <article-title>Large-scale reading comprehension dataset from examinations</article-title>
          .
          <source>arXiv preprint arXiv:1704.04683</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Lelkes</surname>
            ,
            <given-names>A.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>V.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Quiz-style question generation for news stories</article-title>
          .
          <source>In: Proceedings of the Web Conference</source>
          <year>2021</year>
          , pp.
          <fpage>2501</fpage>
          -
          <lpage>2511</lpage>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dave</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wham</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pursel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          :
          <article-title>Distractor generation for multiple choice questions using learning to rank</article-title>
          .
          <source>In: Proceedings of the thirteenth workshop on innovative use of NLP for building educational applications</source>
          , pp.
          <fpage>284</fpage>
          -
          <lpage>290</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.Y.</given-names>
          </string-name>
          :
          <article-title>Rouge: A package for automatic evaluation of summaries</article-title>
          .
          <source>In: Text summarization branches out</source>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>81</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Sherlock: a semi-automatic quiz generation system using linked data</article-title>
          .
          <source>In: International Semantic Web Conference (Posters &amp; Demos)</source>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>12</lpage>
          ,
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Mitkov</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Le</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Karamanis</surname>
          </string-name>
          , N.:
          <article-title>A computer-aided environment for generating multiple-choice test items</article-title>
          .
          <source>Natural language engineering 12(2)</source>
          ,
          <fpage>177</fpage>
          -
          <lpage>194</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Papineni</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roukos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          , W.J.:
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          .
          <source>In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics</source>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Prince</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Does active learning work? a review of the research</article-title>
          .
          <source>Journal of engineering education 93(3)</source>
          ,
          <fpage>223</fpage>
          -
          <lpage>231</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Rafel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Narang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matena</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <surname>P.J.:</surname>
          </string-name>
          <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>
          . arXiv preprint arXiv:
          <year>1910</year>
          .
          <volume>10683</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Rilof</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thelen</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A rule-based question answering system for reading comprehension tests</article-title>
          . In: ANLP-NAACL 2000 Workshop: Reading Comprehension Tests as
          <article-title>Evaluation for Computer-Based Language Understanding Systems (</article-title>
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Rodıgr´ uez Rocha</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faron Zucker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giboin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Extraction of relevant resources and questions from dbpedia to automatically generate quizzes on specific domains</article-title>
          .
          <source>In: International Conference on Intelligent Tutoring Systems</source>
          , pp.
          <fpage>380</fpage>
          -
          <lpage>385</lpage>
          , Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Sakaguchi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arase</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Komachi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Discriminative approach to fill-in-theblank quiz generation for language learners</article-title>
          .
          <source>In: Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pp.
          <fpage>238</fpage>
          -
          <lpage>242</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Tafjord</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
          </string-name>
          , P.:
          <article-title>General-purpose question-answering with macaw</article-title>
          .
          <source>arXiv preprint arXiv:2109.02593</source>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Van</surname>
            <given-names>Campenhout</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Dittel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.S.</given-names>
            ,
            <surname>Jerome</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.G.</surname>
          </string-name>
          :
          <article-title>Transforming textbooks into learning by doing environments: an evaluation of textbook-based automatic question generation</article-title>
          .
          <source>In: Third Workshop on Intelligent Textbooks at the 22nd International Conference on Artificial Intelligence in Education</source>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Van Der Lee</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gatt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Miltenburg</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wubben</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krahmer</surname>
          </string-name>
          , E.:
          <article-title>Best practices for the human evaluation of automatically generated text</article-title>
          .
          <source>In: Proceedings of the 12th International Conference on Natural Language Generation</source>
          , pp.
          <fpage>355</fpage>
          -
          <lpage>368</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosma</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>V.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lester</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          :
          <article-title>Finetuned language models are zero-shot learners</article-title>
          .
          <source>arXiv preprint arXiv:2109.01652</source>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Neural generative question answering</article-title>
          .
          <source>arXiv preprint arXiv:1512.01337</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Co-attention hierarchical network: Generating coherent long distractors for reading comprehension</article-title>
          .
          <source>In: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , vol.
          <volume>34</volume>
          , pp.
          <fpage>9725</fpage>
          -
          <lpage>9732</lpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>