<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LOGANShiny: An app for illustrating process data analysis from international large-scale assessments</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Denise Reis Costa Centre for Educational Measurement University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes a Shiny application for the R package LOGAN, LOGANShiny. This app was built to provide researchers and education stakeholders an overview of basic tools for starting their analysis of process data from international largescale assessments. Using the log file data from one item displayed at the PISA 2012 creative problem-solving assessment, the app is divided in three modules: (a) Data Preparation, (b) Response times, and (c) Respondent's actions. In each module, the user can interact with the app by analyzing students' performance on the item or comparing specific groups of students (e.g., gender or cross-country analyses). The exploration of such tools can not only illustrate the potential and limitation of process data analysis from these assessments but can also advance one's understanding of how students from 44 countries and economies interact with a problem-solving item on an international survey.</p>
      </abstract>
      <kwd-group>
        <kwd>Computer-based assessment</kwd>
        <kwd>Log data</kwd>
        <kwd>Digital items</kwd>
        <kwd>R package</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>International large-scale assessments have received widespread
attention by measuring key cognitive skills and gathering
information and data on how individuals use their knowledge in
different contexts. For example, since 2012, two important
assessments conducted by the Organisation for Economic
Cooperation and Development (OECD), the Programme of
International Student Assessment (PISA) and the Programme for
International Assessment of Adult Competencies (PIAAC), not
only started the administration of computer-based formats for a
large number of participating countries but also made a number of
items with respondent’s log file information publicly available.</p>
      <p>These log data contain a record of the interactions between the
respondents and the computer testing application during the
assessment.</p>
      <p>
        Process data from these kinds of data (e.g., response times and
respondent’s actions) are of potential relevance to researchers and
can provide a better understanding of a range of issues related to
test-taking behavior (e.g., engagement [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], navigation behavior
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). Despite these potentialities, research on this field is still not
well developed due to the challenges and obstacles associated
with the management of such data [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        To overcome this difficulty, an open-source R package was
developed: LOG file ANalysis in international large-scale
assessments (LOGAN [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]). This package is intended to present a
set of user-facing functions, and the user does not need to be
knowledgeable of the details of the underlying code or extensively
work on the data management to conduct specific analysis of the
log files from these assessments.
      </p>
      <p>
        To demonstrate the functionalities of the LOGAN package for use
by researchers and education stakeholders interested in process
data analysis, a web-based application using the Shiny app [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
was created, the LOGANShiny app. Hosted at the
https://loganpackage.shinyapps.io/shiny/ page, this interactive
platform brings to the users examples of analysis for one released
PISA 2012 creative problem-solving item, the Climate Control
(CP025Q01).
      </p>
      <p>To answer this item, students were first presented a stimulus
(Figure 2) where they needed to manipulate input variables (top,
central, and bottom controls/sliders) to understand how an air
conditioner changes the temperature and humidity of a room.</p>
      <p>Then, students had to draw arrows on a diagram that represent the
relationship between the three controls and the two outputs
(temperature and humidity). Full credit was given to students who
correctly completed this diagram (i.e., top control impacts
temperature and central and bottom controls on humidity).</p>
      <p>
        The available log file data from this item captured the student’s
time on the task, and their exploration on applying and resetting
the input variables using the sliders, the associated temperature
and humidity values, and the state of the diagram at each
exploration. There was no restriction on the number of times a
student could manipulate these features, and they did not change
Copyright © 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
by themselves without the student's interaction. Example of
studies using process data from this item are: [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
and [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>The following sections of this paper intend to showcase the
features of the LOGANShiny app regarding data management and
statistical analysis from process data for this item.</p>
    </sec>
    <sec id="sec-2">
      <title>2. DATA PREPARATION (MODULE 0)</title>
      <p>In tab “Module 0” from LOGANShiny app, a user will be
presented with the particularities of the Climate Control item and
its related log-file data. An interactive table displays all the 13
variables and 951,481 entries existing in the data. It represents
how students from 44 countries and economies interacted with
this item. The description of each variable and the three log events
of one student from the United Arab Emirates (code = “ARE”) is
illustrated in Table 1.
The first log event for the student ID = “04852” indicates when
the student was exposed to the item for the first time (event
status= “START_ITEM”). For this case, the registered time was
1288.1 seconds since the beginning of the assessment, and no
interactions with the item features were recorded (e.g.,
top_setting= “NULL”).</p>
      <p>Less than four seconds after starting the item, this student clicked
on the “RESET” button. In this scenario, all the input variables
were set as 0 (indicated by a triangle in Figure 2), the output
indicated “25”, and no arrows were drawn on the diagram as the
default.</p>
      <p>About 46 seconds after resetting the task, the same student moved
all sliders in one position to the right (i.e., top_setting= “1”,
central_setting= “1”, and bottom_setting= “1”) and clicked on
“APPLY”. In this case, the temperature value automatically
changed to “27” and the humidity to “28”. Again, the status of the
diagram was still in its initial setting.</p>
      <p>When the student interacted with the diagram, the information
displayed in the “diag_state” variable was represented as a binary
number (e..g, “'000001“) with each digit associated with one input
and one output variable (e.g., top control and temperature).
After looking at the information that one can extract from the
available log file data, analytical tools are presented at the
LOGANShiny app. For example, a summary of the total number
of event actions (including "START_ITEM" and "END_ITEM")
can be performed in the app. Figure 3 illustrates the log events
from 1,015 students from Bulgaria (code=”BGR”). The same
analysis can also be done for data from other countries.</p>
      <p>From the provided summary statistics, one can verify issues
related to the OECD log data. For example, a student from this
country has only one entry in the log file data (i.e., one event
action). Since it is expected to have at least two events for this
dataset ("START_ITEM" and "END_ITEM"), the app emphasizes
the importance of a closer look at the data and acknowledges
researcher’s freedom to review/filter/delete such inconsistencies.
After data management, two analytical tools are provided in the
subsequent tabs: Time (Module 1) and Actions - Cognitive related
(Module 2).</p>
    </sec>
    <sec id="sec-3">
      <title>3. RESPONSE TIMES (MODULE 1)</title>
      <p>On this tab, the amount of time students spent on the Climate
control item is analyzed. First, a user should decide if the analysis
of the total time will be conducted by item performance
(CP025Q01=0: incorrect answer; CP025Q01=1: correct answer)
or gender (ST04Q01=1: female; ST04Q01=2: male). Later, the
user can choose if the analyses will consider all countries or select
a specific country, as illustrated in Figure 4.
After these choices, two types of descriptive statistics are
provided: a summary table and a density plot.</p>
    </sec>
    <sec id="sec-4">
      <title>3.1 Summary table of response times</title>
      <p>Figure 5 displays the information that one can gather from the
LOGAN package for the analysis of the overall time and by item
performance. A total of 30,345 students from all PISA 2012
participating countries and economies was analyzed for the
Climate control item. The maximum amount of time spent on this
item was 26 minutes from a student who got an incorrect answer.
For the group of students who got a correct answer, the maximum
was 16 minutes.</p>
      <p>In general, students spent an average of 2 minutes on the task.
However, this estimate is not precise since negative response
times were observed in this sample (i.e., the minimum amount of
time for those who got an incorrect answer was equal to -0.43).
Although one could remove such cases from the dataset as they
lower the average values, it is displayed in the LOGANShiny to
reveal another inconsistency in this log-file dataset.</p>
      <p>Even though this type of discrepancy could possibly be detected
at the data preparation stage of analysis, it was left to the “Module
1: Time” tab for the LOGANShiny to highlight again the
importance of further inspection of process data and proper data
manipulation of the files from large-scale assessments.</p>
    </sec>
    <sec id="sec-5">
      <title>3.2 Time density plot</title>
      <p>Figure 6 illustrates the distribution of time by the performance on
the task obtained with the LOGANShiny app. One could also plot
the distribution of time by gender.</p>
    </sec>
    <sec id="sec-6">
      <title>4. RESPONDENT’S ACTION (MODULE 2)</title>
      <p>
        To illustrate how to explore the actions recorded in the log files,
LOGANShiny describes two respondent’s action strategies
discussed at [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] based on the vary-one-thing-at-at-time (VOTAT)
strategy. In the case of the Climate control item, the VOTAT
strategy consists of a student varying one specific variable (i.e.,
put the top control on "++"), while keeping all other variables
constant (i.e., put the central and bottom controls on the delta
symbol), and clicking on "apply". To operationalize the VOTAT
strategy, [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]’s authors suggest:
(a) VOTAT 1: a dichotomous variable with “1” to students who
applied VOTAT for all input variables; and
(b) VOTAT 2: incorporated four categories for no isolated
variation at all (category 0), isolated variation of one input
variable (for example, only the top control), isolated variation of
two input variables (for example, the top and bottom controls),
and isolated variation of all three input variables (category 3).
One must note that VOTAT 2 category 3 is the same as the
VOTAT 1 = “1”. Category 0, on the other hand, indicates the case
where the student did not vary any slider or vary all the sliders at a
time before clicking the “apply” button.
      </p>
      <p>To illustrate how to derive these VOTAT variables from the log
data, the third log event from Table 1 shows the case where the
student selected top_setting= “1”, central_setting= “1”, and
bottom_setting= “1” before clicking on apply. In this scenario,
both VOTAT 1 and VOTAT 2 would receive the value “0” where
no isolated variation on all the controls were found.</p>
      <p>Based on these categories, one can investigate how performance
outcomes and VOTAT strategies are related by country, item level
performance (CP025Q01=0: incorrect answer; CP025Q01=1:
correct answer), and problem-solving overall performance (first
plausible value, PV1CPRO). To do this on LOGANShiny, one
should select the type of VOTAT strategy they are interested in,
followed by each participating country the analyses will be related
to (Figure 7):
After these choices, two types of descriptive and correlational
statistics are provided: a summary report and a frequency plot.</p>
    </sec>
    <sec id="sec-7">
      <title>4.1 Summary report of student’s strategies and performance</title>
      <p>On LOGANShiny, it is possible to conduct a statistical summary
of students’ exploration via the “VOTAT 1” strategy and its
relationship with performance. This analysis is presented as a
report divided in three parts: (1) frequency table, (2) measures of
association between strategy and item performance, and (3)
summary of test performance (considering the first plausible value
from the PISA 2012 problem-solving assessment) by VOTAT
strategy. Figures 8 and 9 show an example of this report.
From Figure 8, it is possible to see that about half of the students
from this sample applied the VOTAT 1 strategy at least one time.
For the group of students who got a correct answer in the
CP025Q01 item, the majority (12,404 out of 15,076 students)
applied this strategy at least once during the item evaluation.
Correlational measures (i.e., chi-square statistic and phi
coefficient) are also provided to evaluate the strength of the
association between these variables.</p>
      <p>
        Imputation methods are used in PISA to generate plausible values
to report students’ overall performance [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In a scale with a mean
score among OECD countries of 500, five plausible values were
defined for the PISA 2012 creative problem-solving assessment.
In the LOGANShiny, an analysis using one plausible values is
illustrated in Figure 9.
      </p>
      <p>Based on the provided statistics, it is possible to note that
students who used the VOTAT 1 strategy on the Climate Control
item received, on average, more than 100 score points on the
PISA 2012 creative problem-solving assessment in contrast to
those who did not use this strategy.</p>
    </sec>
    <sec id="sec-8">
      <title>4.2 Frequency Plot</title>
      <p>
        In PISA, student’s scores in the assessments are also divided into
proficiency scale levels to provide a substantive meaning of the
overall performance. For PISA 2012 creative problem-solving
assessment, seven levels of proficiency were created where level 1
(358 &lt; PV1CPRO ≤ 423) corresponds to an elementary level of
problem-solving skills and level 6 (PV1CPRO &gt;= 683) the
highest level. A complete description of these levels is presented
in Figure V.2.2 from the OECD report [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>In LOGANShiny, these proficiency levels are plotted in relation
to the use of the VOTAT strategy. Figure 10 illustrates this
relationship. Here, percentages within the categorized proficiency
score are provided in parenthesis for each PISA proficiency level.
Findings from this analysis indicate that students on the high level
of the scale tend to use “VOTAT 1” more than those on the lower
levels of the PISA 2012 creative problem-solving proficiency
scale.</p>
    </sec>
    <sec id="sec-9">
      <title>5. CONCLUSION</title>
      <p>
        In this paper, LOGANShiny is presented as an illustrative tool for
showcasing the functionalities of the LOGAN R package
functions for the analysis of process data from international
largescale assessments. Interactive tables and graphical displays
intended to shed light on the potentialities and limitations of the
use of log-file data regarding data management and analysis of
response times and student’s actions. This app can be a valuable
tool to deepen researchers’ and education stakeholder’s
knowledge on the item features and provide insights on students’
cognitive process. The understanding of how process data can be
extracted and analyzed may not only inspire the development of
new item features that could enrich one’s experience with digital
environments, but also has the potential to improve the
assessment’s results by, for instance, incorporating process data
into the scoring procedure [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , W., Cheng, J.,
          <string-name>
            <surname>Allaire</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McPherson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <source>shiny: Web Application Framework for R. R package version 1.4.0</source>
          .2. https://CRAN.R-project.org/package=shiny
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ying</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Statistical Analysis of Complex Problem-Solving Process Data: An Event History Analysis Approach</article-title>
          . Frontiers in Psychology,
          <volume>10</volume>
          , 486. https://doi.org/10.3389/FPSYG.
          <year>2019</year>
          .00486
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Goldhammer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martens</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Lüdtke</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Conditioning factors of test-taking engagement in PIAAC: an exploratory IRT modelling approach considering person and item characteristics</article-title>
          .
          <source>Large-Scale Assessments in Education</source>
          ,
          <volume>5</volume>
          :
          <fpage>18</fpage>
          . https://doi.org/10.1186/s40536-017-0051-9
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Greiff</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wüstenberg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Avvisati</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Computergenerated log-file analyses as a window into students' minds? A showcase study based on the PISA 2012 assessment of problem solving</article-title>
          .
          <source>Computers and Education</source>
          ,
          <volume>91</volume>
          ,
          <fpage>92</fpage>
          -
          <lpage>105</lpage>
          . https://doi.org/10.1016/j.compedu.
          <year>2015</year>
          .
          <volume>10</volume>
          .018
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Hahnel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldhammer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kröhne</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Effects of linear reading, basic computer skills, evaluating online information, and navigation on reading digital text</article-title>
          .
          <source>Computers in Human Behavior</source>
          ,
          <volume>55</volume>
          ,
          <fpage>486</fpage>
          -
          <lpage>500</lpage>
          . https://doi.org/10.1016/j.chb.
          <year>2015</year>
          .
          <volume>09</volume>
          .042
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          , &amp; von
          <string-name>
            <surname>Davier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Predictive Feature Generation and Selection from Process Data in PISA Simulation-Based Environment: An Implementation of Treebased Ensemble Methods</article-title>
          . Frontiers in Psychology,
          <volume>10</volume>
          , 2461. https://doi.org/10.3389/fpsyg.
          <year>2019</year>
          .02461
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>He</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          , &amp; von
          <string-name>
            <surname>Davier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Analyzing Process Data from Problem-Solving Items with N-Grams</article-title>
          .
          <source>In Handbook of Research on Technology Tools for Real-World Skill Development</source>
          (pp.
          <fpage>750</fpage>
          -
          <lpage>777</lpage>
          ). https://doi.org/10.4018/978-1-
          <fpage>4666</fpage>
          -9441-5.
          <fpage>ch029</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>OECD.</surname>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>PISA 2012 Results: Creative Problem Solving</article-title>
          (Volume V): Vol. V. OECD Publishing. https://doi.org/https://doi.org/10.1787/9789264208070-en
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Pejic</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Molcer</surname>
            ,
            <given-names>P. S.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Exploring data mining possibilities on computer based problem solving data</article-title>
          .
          <source>SISY 2016 - IEEE 14th International Symposium on Intelligent Systems and Informatics, Proceedings</source>
          ,
          <fpage>171</fpage>
          -
          <lpage>176</lpage>
          . https://doi.org/10.1109/SISY.
          <year>2016</year>
          .7601491
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Reis</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            , &amp;
            <surname>Leoncio</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>LOGAN: An R package for log file analysis in international large-scale assessments</article-title>
          . R Package. https://cran.rproject.org/web/packages/LOGAN/index.html
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Reis</given-names>
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Bolsinova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tijmstra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            , &amp;
            <surname>Andersson</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>Improving the Precision of Ability Estimates Using Time-On-Task Variables: Insights From the PISA 2012 Computer-Based Assessment of Mathematics</article-title>
          . Frontiers in Psychology,
          <volume>12</volume>
          . https://doi.org/10.3389/fpsyg.
          <year>2021</year>
          .579128
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ying</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Latent Class Analysis of Recurrent Events in ProblemSolving Items</article-title>
          . Applied Psychological Measurement,
          <volume>42</volume>
          (
          <issue>6</issue>
          ),
          <fpage>476</fpage>
          -
          <lpage>498</lpage>
          . https://doi.org/10.1177/0146621617748325
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>