<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Questions⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Richard Glassey</string-name>
          <email>glassey@kth.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olle Bälter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Learnersourcing, Peerwise, Inter-rater Reliability</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>KTH Royal Institute of Technology</institution>
          ,
          <addr-line>Stockholm</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>Learnersourcing presents an eficient and economical pathway to producing more learning content, whilst engaging students more actively and deeply in their learning. However, it also presents new challenges to solve. Most of all, how best can we manage the variance in quality of content produced. In this work, we focus on the extent to which students and teachers agree on quality of learnersourced multiple choice questions. Students (n=30) were tasked with producing six questions over three weeks of an introductory programming course as part of their assessment. They also had to review 12 questions authored by their peers over the same period using a set of principles for good questions. After this period, four teaching staf involved with the course reviewed the student questions using the same process and principles. Inter-rater reliability statistics found overall positive agreement across principles, however this dropped to weaker agreement for principles aimed at more subjective and higher order concerns of question quality and quality of question feedback.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Learnersourcing can be succinctly defined as crowdsourcing content from students in a learning
context [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Diferent types of learnersourced content include: content annotation, resource
recommendation, explanation of misconceptions, content creation, and aspects of evaluation,
reflection and regulation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In all cases, students are active in adding value to content and
creating new content for the consumption of other students.
      </p>
      <p>
        Creating learning content is a higher-order activity for students [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, as with any
student activity there will be a spectrum of engagement and quality. This creates a challenge of
management as ideally students are exposed to the highest quality content, whilst lower quality
content is filtered out of the system. Both PeerWise [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and RiPPLE [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] have adopted content
quality management strategies as platform features. However, one question that bubbles up is
do students and teachers agree what quality is, given its subjective nature?
      </p>
      <p>
        Here, we approach this question by studying the agreement gaps that emerge when students
and teachers are asked to give their opinions about learnersourced MCQs. To give structure
to these opinions, we use a set of principles for producing better MCQs that were presented
during the production and review of questions [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Finally, we calculated inter-rater reliability
statistics [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to discover both the diferences within student and teacher reviews and between
student and teacher reviews.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Context of Study</title>
      <p>
        In response to the 2015 European Refugee Crisis, academics at KTH Royal Institute of Technology,
Stockholm, Sweden developed an intensive three month training to integrate newly arrived
into the local IT workforce [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. As this training was not attached to traditional course delivery,
there was much more freedom to innovate and try novel pedagogical interventions [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In the
ifrst three weeks, when students were covering introductory programming, assessment was
achieved by having students create multiple choice questions on the topics they were learning.
The main motivation was to quickly generate a lot of MCQs that the students could then answer
to increase their opportunities to practice. This was achieved and it was found that 50% of
students answered 100 questions or more, without any demand from teachers to do so [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Early iterations of this training found that, whilst students could produce good MCQs there
was a cold-start problem in MCQ quality - students gradually got better over time. Furthermore,
students were good at writing questions and answering alternatives, however when it came to
the explanation or feedback that accompanied the MCQ, students struggled to provide similar
quality [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In response, we developed a set of 12 principles for writing good MCQs, reflecting
the patterns of quality issue we detected in reviewing student MCQs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], rather than more
established guides developed for academics [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The principles are listed in table 1.
      </p>
      <p>In the most recent iteration, to help students understand and apply the principles, students
worked as a group to learn about writing good questions before they undertook the task
individually. Principles were used to create a question and then review a question created by
another group. In this way, students had a chance to discuss their interpretation of the principles
and apply them both in creating and reviewing MCQs. For each week of the three week course,
students were tasked with creating two questions and reviewing four questions. Once the
course had ended, four teachers reviewed 96 questions independently using the principles and
review process used by the students. For each question there were three student reviews and
four teacher reviews.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Findings and Discussion</title>
      <p>First, taking student and teacher IRR scores separately, teachers had a higher average
agreement over all principles (0.82) versus students (0.65). Also, students had a greater distribution
of scores (from 0.20 to 0.97) versus teachers (from 0.54 to 1.00). These findings can be partly
explained by the teachers both generating the principles and also discussing their possible
interpretations. Students on the other hand had much less opportunities to discuss the interpretation
of the principles, other than within the scheduled group creation and review sessions, which
only occurred twice in the three week course.</p>
      <p>Second, looking at the largest diferences in IRR scores by principle, “ Only the feedback to
the correct alternative reveals the answer ” had a magnitude of 0.53. This suggests a diference
in attitude with teachers seeing more value for MCQ feedback that helps correct student
misconception without taking away the chance of a second attempt. Another large diference
relating to the feedback (0.37), “Feedback is unique and provided for each answer alternative”,
continues this theme, whereas students did not have similar strong agreement. This is interesting
as it is an objective measure - each answering alternative should have its own unique feedback
and it only requires a cursory glance to confirm. Much like the smallest diference (0.03) for
“Three or more answer alternatives are provided”, where it is quite easy to count the number of
answering alternatives. One limitation of PeerWise is that there is only a single text area for
‘explanation’ and this may have contributed to this anomaly.</p>
      <p>Third, looking at the smallest diferences that are not clearly objective or easy to determine
(like “Question is from the course domain”), both principles: “All answer alternatives are plausible
and related to a misconception” and “Feedback is suficient for you to understand why each
alternative was incorrect or correct” are interesting as these are quite challenging aspects of
quality to agree upon, even for teachers who understand the intent of each principle. This
suggests a limitation of the depth of quality one can hope to measure when rating learnersourced
content, however the IRR results here are still promising with medium agreement between
students (0.59 and 0.48) and between teachers (0.69 and 0.54).</p>
      <p>Finally, when combining both student and teacher reviews together, creating a pool of
seven reviewers, the most agreement can be found in perhaps the most objectively answerable
principles. This is not surprising and acts as a nice control for agreement on the basics of
MCQs. Of more concern is that where there is weak agreement, two principles are concerned
with feedback (“Feedback is suficient for you to understand why each alternative was incorrect
or correct” and “Only the feedback to the correct alternative reveals the answer ”) and the other
concerning if the question targets higher order thinking. This represents a challenge that
warrants deeper investigation as the value of feedback is well known and accepted in efective
education, but what if our diferent points of view on it never actually meet and the feedback
that teachers feel is suficient is not exactly (or even close to) what students need?</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>Learnersourcing generates more content, but comes with the challenge of how to determine
quality. Part of solving the challenge is finding out where students and teachers agree (or not)
about quality. Use of agreement statistics, such as inter-rater reliability, are a potentially useful
metric, but as mentioned here, caution is advised as there are many to choose from and not all
work as expected. The work presented here shows that there are areas we can find agreement,
however we need to find better ways to solicit impressions of quality at deeper levels and then
ifnd ways to integrate and automate them within learnersourcing platforms.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Learnersourcing: improving learning with collective learner activity</article-title>
          ,
          <source>Ph.D. thesis</source>
          , Massachusetts Institute of Technology,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schlagwein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Benatallah</surname>
          </string-name>
          ,
          <article-title>A review on crowdsourcing for education: State of the art of literature and practice</article-title>
          .,
          <source>PACIS</source>
          (
          <year>2018</year>
          )
          <fpage>180</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Abdi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Khosravi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sadiq</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Demartini, Evaluating the quality of learning resources: A learnersourcing approach</article-title>
          ,
          <source>IEEE Transactions on Learning Technologies</source>
          <volume>14</volume>
          (
          <year>2021</year>
          )
          <fpage>81</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Denny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Luxton-Reilly</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Hamer,</surname>
          </string-name>
          <article-title>The peerwise system of student contributed assessment questions</article-title>
          ,
          <source>in: Proceedings of the tenth conference on Australasian computing education-</source>
          Volume
          <volume>78</volume>
          ,
          <year>2008</year>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Khosravi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gyamfi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. E.</given-names>
            <surname>Hanna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lodge</surname>
          </string-name>
          ,
          <article-title>Fostering and supporting empirical research on evaluative judgement via a crowdsourced adaptive learning system</article-title>
          ,
          <source>in: Proceedings of the Tenth International Conference on Learning Analytics &amp; Knowledge</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Glassey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bälter</surname>
          </string-name>
          ,
          <article-title>Put the students to work: Generating questions with constructive feedback</article-title>
          ,
          <source>in: 2020 IEEE Frontiers in Education Conference (FIE)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K. L.</given-names>
            <surname>Gwet</surname>
          </string-name>
          ,
          <article-title>Computing inter-rater reliability and its variance in the presence of high agreement</article-title>
          ,
          <source>British Journal of Mathematical and Statistical Psychology</source>
          <volume>61</volume>
          (
          <year>2008</year>
          )
          <fpage>29</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiggberg</surname>
          </string-name>
          , E. Gobena,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kaulio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Glassey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bälter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hussain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guanciale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Haller</surname>
          </string-name>
          ,
          <article-title>Efective reskilling of foreign-born people at universities-the software development academy</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>24556</fpage>
          -
          <lpage>24565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Glassey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bälter</surname>
          </string-name>
          ,
          <article-title>Sustainable approaches for accelerated learning</article-title>
          ,
          <source>Sustainability</source>
          <volume>13</volume>
          (
          <year>2021</year>
          )
          <fpage>11994</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>T. M. Haladyna</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Downing</surname>
            ,
            <given-names>M. C.</given-names>
          </string-name>
          <string-name>
            <surname>Rodriguez</surname>
          </string-name>
          ,
          <article-title>A review of multiple-choice item-writing guidelines for classroom assessment</article-title>
          ,
          <source>Applied measurement in education 15</source>
          (
          <year>2002</year>
          )
          <fpage>309</fpage>
          -
          <lpage>333</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>