<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>DigiTransfEd</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Developing and validating an integrated three-tier multiple choice test for smartphone-based experiments in physics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Frank Angelo A. Pacala</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of San Carlos</institution>
          ,
          <addr-line>Cebu City</addr-line>
          ,
          <country country="PH">Philippines</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>3</volume>
      <fpage>23</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>The increasing use of smartphone-based physics experiments has become a theme of many physics educationrelated research and publications, including its actual implementation inside the classroom. Measuring this scientific process and teachers' and students' corresponding conceptual understanding is needed. This paper aims to create a well-designed, reliable, validated integrated instrument to measure the intervention's impact on the student's conceptual understanding and science process skills (SPS) of various physics topics. The reliability measure was internal consistency using Cronbach's alpha, and the validity measures were content and construct validity. The instrument was found reliable with an alpha point estimate value of 0.766, which is considered to be acceptably good. The content validity index revealed some items as inappropriate, so they were removed, and 30 items were appropriate. The exploratory factor analysis (EFA) revealed ten components that were measured in the instrument. These components were classified by their factor loads using orthogonal rotation in varimax. The manual analysis of these factor loads revealed that the EFA was measuring the teachers' cognition level based on the question's science process skills level. The most dominant scores were from remembering and lowly from evaluating level. This revealed the power of the instrument to integrate the SPS and conceptual understanding constructs into one instrument. The instrument is now validated; therefore, this paper suggests using this instrument to measure the level of integrated physics SPS and conceptual understanding during a smartphone-based physics experiment.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;physics education</kwd>
        <kwd>smartphone physics experiment</kwd>
        <kwd>three-tier multiple tests</kwd>
        <kwd>instrument development</kwd>
        <kwd>instrument validation</kwd>
        <kwd>exploratory factor analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Technology integration in education has revolutionized learning, particularly in physics. Smartphones
have emerged as powerful tools for experiments, but the need for standardized assessment tools poses
a challenge. To bridge this gap, this research was conducted to develop and validate robust evaluation
instruments that adapt to the evolving educational technology landscape, enhancing the quality of
physics education.</p>
      <p>
        The traditional method of teaching physics involves conducting experiments in a lab using specialized
equipment, which may not be easily accessible to all students. This results in unequal exposure to
practical concepts, especially in resource-limited environments. Smartphone-based experiments ofer a
potential solution to this issue. However, the need for standardized assessment tools is a significant
obstacle. Several studies (e.g., [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]; [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) were conducted to measure the efect of smartphone-based
experiments in Physics but could not produce standardized tests to measure this efect. The current
evaluation methods may not fully capture smartphone-based experiments’ unique features and outcomes,
underscoring the need for tailored assessment instruments.
      </p>
      <p>
        Moreover, the studies conducted by Cai [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and Hochberg [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] explored the efect of various elements
of smartphone experiments, like augmented reality, on the cognitive loads, self-eficacy, and conceptions
of students’ learning. They used validated research instruments to measure the diferent constructs of
students’ learning. Also, Nikou and Economides [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] measured the impact of mobile-based experiments
on the topic of electric currents. They successfully concluded that the instrument has an internal
consistency of 0.85 on average. These sources used one-tier tests to measure the efect of smartphone
experiments on student learning. This present research is focused on developing and validating a
three-tiered multiple-choice test in the mechanics section of the physics curriculum.
      </p>
      <p>
        Most of the test instruments in the literature measured a single construct like conceptual
understanding, critical thinking, and science process skills (SPS). For instance, Tiruneh et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] measured
the critical thinking of students in electricity and magnetism, and He et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] devised a test to
diagnose students’ conception of aqueous solutions in chemistry. This research created an integrated
three-tier multiple-choice test (ITTMCT) measuring the students’ conceptual understanding and SPS
simultaneously. The approach discussed here aims to assess student’s knowledge of physics principles
theoretically and through their ability to design experiments, analyze data, and draw conclusions
based on empirical evidence. This method combines conceptual understanding with science process
skills, which allows students to gain a deeper understanding of the scientific method and its practical
applications. Ultimately, this prepares them to face complex challenges in physics and related areas
with greater confidence and competence.
      </p>
      <p>
        Furthermore, multiple choice can be created in various ways. Single-tier and two-tier
multiplechoice tests difer in assessing knowledge and understanding. In a single-tier test, each question is a
standalone query with a fixed set of options, evaluating typically factual recall, conceptual understanding,
or problem-solving skills. On the other hand, two-tier multiple-choice tests incorporate follow-up
questions into each primary query, introducing an additional layer of complexity. The first tier resembles
a traditional multiple-choice question, requiring selecting the correct answer from the provided options.
However, the second tier prompts participants to justify their initial choice or reasoning, requiring
them to apply critical thinking skills and provide a rationale for their selection. He et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] argued that
two-tier tests ofer a more comprehensive assessment of understanding, probing factual knowledge and
the thought processes and reasoning behind participants’ responses. Hanson [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] noted that although
they may take more time to administer and grade, two-tier tests ofer deeper insights into participants’
comprehension and analytical abilities. They are, therefore, valuable tools for assessing higher-order
thinking skills and conceptual understanding.
      </p>
      <p>
        A three-tier multiple-choice test has a unique feature called the certainty of response, which adds
an extra layer of diagnostic assessment. In the first tier, participants answer a question, followed
by justifying the second tier. The third tier prompts them to assess their confidence or certainty in
their response. Laeli [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] emphasized that this feature provides valuable insight into their responses’
reliability and accuracy, metacognitive awareness, and self-assessment skills. The third tier helps
measure participants’ knowledge and understanding while comprehensively evaluating their confidence
levels [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. By using this three-tier test in the diagnostic test, educators and evaluators gain a deeper
understanding of the degree of certainty associated with participants’ answers, ultimately contributing
to informed decision-making in educational and diagnostic settings.
      </p>
      <p>Therefore, developing and validating a three-tier multiple-choice test to measure the efect of
smartphone-based experiments in mechanics is necessary. This research aims to create a well-designed,
reliable, and validated instrument to measure the intervention’s impact on the student’s conceptual
understanding of various physics topics. The research questions are as follows: What is the instrument’s
point estimate of Cronbach’s alpha? How many factors this instrument can measure? What is the level
of content validity of the instrument?</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <sec id="sec-2-1">
        <title>2.1. Defining the constructs and formulating table of specifications</title>
        <p>
          The first phase in developing the ITTMCT is to define and determine the constructs under which items
will fall and formulate the table of specifications. For instance, if question 1 is under the construct of
measuring or observing SPS. This paper used Vitti and Torres’s [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] practical science process skills.
Their handbook provided several activities that can measure SPS, but this paper utilized their description
of the process skills. These skills are observing, measuring, sorting and classifying, making inferences,
predicting, experimenting, graphing, communicating, and controlling variables. All of these process
skills were utilized.
        </p>
        <p>Moreover, the test items followed Anderson and Krathwohl’s [12] revised Bloom’s taxonomy of
learning cognition, which diferentiates conceptual knowledge. In this taxonomy, the cognitive process
categorizes cognitive skills into six levels that increase in complexity from Remembering to Creating.
Unlike the original taxonomy, which only focused on the cognitive domain, this revision emphasizes
the role of metacognition in the learning process. In this research, the cognition level started from
remembering to evaluation.</p>
        <p>The instrument followed the Philippines’ Department of Education (DepEd) curriculum guide for
the covered topics. However, the experiment is almost similar to CAIE practical exams, and it was
made sure that the instrument would follow the Science Curriculum Guide of DepEd. This would
ensure horizontalization of the skills and knowledge to be measured, plus the instrument is made to be
localized.</p>
        <p>Napal et al. [13] said science process skills could be seen as a progression or hierarchy, but they
maintain that these skills are interconnected. Hence, both the science process skills and the taxonomy of
cognition can be merged. Both of these constructs were combined to create the ITTMCT. The construct
level in SPS is matched with the taxonomy of learning cognition. For instance, observing is matched
with the remembering skill.</p>
        <p>Each of the levels of cognition is given a total of five items. Observing, making inferences
(understanding), experimenting, graphing, and controlling variables have three items each. The SPS of sorting
and classifying, predicting, measuring, communicating, and making inferences (evaluating) is given
two items each. There are 30 items in these instruments after the revision has been made.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Instrument development</title>
        <p>The items were developed based on the principles of assessment design of Trends in International
Mathematics and Science Study (TIMSS) and Cambridge Assessment International Education (CAIE).
TIMSS consists of two main categories of questions. First is the multiple-choice format, where students
are presented with response options, and second is the constructed response format, where students are
expected to come up with their answers [14]. The CAIE also consists of multiple-choice tests, usually
paper 1, and structured questions, generally papers 2 and 4.</p>
        <p>The physics topics covered in this test are measuring acceleration due to gravity, magnetic fields,
Doppler efect, and momentum and collision. These topics are commonly taken by students aged 16-18
under the Advanced Subsidiary (AS) and A-Level programs. These topics mostly come with practical
activities under the Cambridge 9702 Curriculum. Although the topics do not represent the whole 9702
syllabus, these experiments are vital because they are core experiments.</p>
        <p>These question structures benefit this three-tier test. The first part is the common multiple-choice
C. To stabilize the readings on the magnetometer
D. To measure the length of the solenoid
2. When recording the time taken for the experiment
using the Phyphox mobile application, what would be
the most appropriate unit of measurement?
A. Volts
B. Seconds
C. Tesla
D. Meters
3. Explain how you measured the fractional uncertainty
in magnetic field strength (B).</p>
        <p>A. dividing 0.2 seconds to the value of B
B. multiplying 0.2 seconds by the value of B
C. dividing the value of B to 0.2 seconds
D. multiplying 0.02 seconds by the value of B
Describe the other [ ] Sure
units you did not
choose.</p>
        <p>State the reasoning [ ] Sure
for your answer.</p>
        <p>[ ] Unsure
[ ] I guessed my
answer
[ ] Unsure
[ ] I guessed my
answer
[ ] Unsure
[ ] I guessed my
answer</p>
        <p>Tier 1 Tier 2 Tier 3
1. Explain the significance of waiting for 60 seconds after Write a reason for [ ] Sure
placing the phone near the solenoid in the experiment. your answer.</p>
        <p>What does this duration aim to achieve?
A. To allow the conducting wire to cool down</p>
        <p>B. To synchronize with the voltage generator
question with a stem-option structure. The second tier is the structured/response question, which is a
segment of papers 2 and 4 of CAIE and the constructed response of TIMSS. The addition is the certainty
level of student response. The number of items in multiple choices and constructed responses is almost
similar [15]. This paper followed this. The guiding principles of CAIE and TIMSS were followed as
much as possible during this item development.</p>
        <p>The third tier is the certainty level of the student’s response; each question contains this level. This
third tier asked the students to reveal the extent of their certainty in their answers. There are two
options: sure, unsure, and guessed the answer.</p>
        <p>Initially, five questions were constructed. Three colleagues with at least four years of teaching
experience in the CAIE curriculum and experience dealing with TIMSS checked these five questions
for consistency with the assessment design of the said two frameworks. Then, their comments and
suggestions were incorporated, and the question-making process continued until the questions reached
37. In table 5, the questions were only 30 because this is after the disqualification of some items due to
lower internal consistency.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Scoring guide</title>
        <p>Each item is given a full five marks. Table 3 shows the full-scoring guide, which is borrowed from
Pacala [16]. The provided rubric describes a methodical way to assess students’ conceptual knowledge
at diferent levels. Students who score at the lowest level, “No Understanding”, give wholly inaccurate
answers, demonstrating a lack of background knowledge or a conceptual misunderstanding. As they
advance to “Alternative Conception”, pupils could show signs of incomplete comprehension and false
beliefs or explanations. “Partial Understanding with Alternative Conception” denotes an answer that
contains both true and false information, frequently together with ambiguities or diferent
interpretations. On the other hand, “Partial Understanding without Alternative Conception” denotes an answer
that is partially accurate but lacks precision or clarity. When students reach the highest level, “Complete
Understanding”, they can demonstrate a thorough understanding of the idea and confidence in their
answers by providing entirely accurate solutions.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Pilot testing</title>
        <p>The test instrument was piloted to a small circle of science teachers in the Division of Catbalogan City.
They were contacted via Facebook Messenger and email to ask if they would participate (N = 8) in the
pilot testing. Once they approved it, the researcher sent the test instrument using a Google form to their
accounts in Messenger or email. At the end of this form is a statement about their comments/suggestions
towards the instrument.</p>
        <p>This pilot testing aims to evaluate the instrument’s validity and reliability before the complete
administration of the test. This pilot testing phase allows researchers to determine if the test items
efectively measure the intended constructs or skills. Any issues with the test items can be identified
and corrected at this stage to ensure that the final test is reliable and valid for assessing the desired
outcomes.</p>
        <p>Based on this pilot testing, it is revealed that teachers left some items in tier two blank because they
did not have an introductory statement that would guide them to the reason. This was mended by
providing a command structure to the statements in tier 2. During this pilot testing, the Cronbach’s
Alpha was 0.62.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Instrument revision and final administration</title>
        <p>The comments, suggestions, and insights from the pilot testing were incorporated into the updated
instrument. The participants noted that the value of the constant (e.g., Planck’s constant) should be
present in the question’s stem. They commended the standard length of the question and the separation</p>
        <p>Relevance Clarity Simplicity Ambiguity
1 = not relevant 1 = not clear 1 = not simple 1 = doubtful
2 = item needed some 2 = item needed some 2 = item needed some 2 = item needed some
revision revision revision revision
3 = relevant but need 3 = clear but need mi- 3 = simple but need 3 = no doubt but need
minor revision nor revision minor revision minor revision
4 = very relevant 4 = very clear 4 = very simple 4 = meaning is clear
of each paragraph to enhance readability.</p>
        <p>The revised instrument was finally administered to the science teachers (N = 30) in one of the schools
in the Catbalogan City Division. These teachers taught DepEd’s science curriculum and had at least
three years of teaching experience. The teachers were not participants in the pilot testing. During
the day of testing, the tables and chairs were arranged following the CAIE standards as stated in the
What to Say to Candidates document. For instance, they were told to fill out the instrument with their
names, school names, testing date, and candidate numbers. The time given was one hour; however,
many teachers finished 15-20 minutes before the hour. Once they were done, they were told that they
could leave the testing room.</p>
        <p>The teachers were sent a letter if they agreed to join this activity. Once they signed up, they were
included in the pool of participants. The teachers were told that the data collected was for research
purposes only. Their names, school names, and scores were not published on the Internet or social
media to preserve their anonymity and confidentiality.</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.6. Data analysis</title>
        <p>The ITTMCT results were subjected to validity and reliability testing. The validity measures were content
validity and construct validity using exploratory factor analysis (EFA). The instrument’s reliability was
measured using the internal consistency value of Cronbach’s alpha.</p>
        <p>The content validity index (CVI) was used to measure content validity. The CVI is a way to assess
the validity of a questionnaire or test. Polit et al. [17] argued that CVI determines how well the items in
the instrument represent the content being measured. A panel of experts judges the relevance of each
item using a 4-point scale. The higher the CVI score, the stronger the instrument’s content validity,
indicating that the items efectively measure the intended construct.</p>
        <p>The instrument used was adapted from Waltz and Bussel [18]. This scale contains four sections:
relevance, clarity, simplicity, and ambiguity. For each section, the expert rated the instrument from 1 to
4. As shown in table 4, the expert rated the instrument as 1 for not relevant and 4 for very appropriate
under the relevance section.</p>
        <p>The researcher calculated the average rating for each criterion to determine an item’s CVI based on
relevance, clarity, simplicity, and ambiguity. This was done by adding up the ratings given by all ten
experts and then dividing by the total number of experts. Then, the overall CVI was taken from the
average of these four sections.</p>
        <p>Another measure to ensure the validity of the study’s instrument is the EFA. It is a statistical method
widely used in research to investigate the underlying structure of a set of variables and discover
the relationships between them [19]. They added that this technique is beneficial in evaluating a
measurement tool’s validity by exploring the data’s dimensionality and uncovering the hidden factors
that may afect participants’ responses. EFA ofers valuable insights into the structure of the construct
being measured, which can aid researchers in refining their measurement tools, creating theories, and
guiding future research eforts.</p>
        <p>This study has only 30 participants (N = 30). Some author recommends using a formula where the
number of samples should be at least five times the number of variables to determine the sample size
[20]. This study’s combination of SPS and conceptual understanding is only one variable. Hence, N =
30 is suitable for EFA.</p>
        <p>The data analysis for the EFA and Cronbach’s Alpha were conducted using the JASP free software,
while the descriptive statistics like mean and standard deviation were from Microsoft Excel.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results and discussion</title>
      <sec id="sec-3-1">
        <title>3.1. Reliability of the instrument</title>
        <p>The instrument’s reliability was measured by internal consistency using Cronbach’s Alpha. A new
instrument is said to have good internal consistency if the Alpha value is higher than 0.60 [21]. The
values of the alpha per item are found in table 5.</p>
        <p>The data were uploaded to Jasp Software. Reliability analysis and unidimensional reliability were
chosen, and Cronbach’s Alpha was selected. The study revealed the overall point estimate of alpha is
0.766 when items 8, 10, 11, 18, 28, 30, and 36 were removed.</p>
        <p>The data in table 5 provides information regarding the participants’ preliminary test scores. Most
obtained a score of 2, meaning their concepts contain alternative conceptions. Very few got a score of 3,
meaning they have a partial understanding with alternative conceptions.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Content validity of the instrument</title>
        <p>The ten experts who judged the instruments were teachers of the CAIE curriculum for at least three
years. They are familiar with the CAIE content and how TIMSS questions are structured. They were
given a printed copy of the instrument and the scoring guide for content validity. It took the judges
four days to finish all the scoring and writing additional comments for the instrument.</p>
        <p>The experts appraised most items with a score between 3.00 and 3.99, meaning minor revision was
needed. Only a tiny number were rated perfect. On the other hand, some items were not relevant, not
clear enough, not simple or very dificult, and too doubtful, while some needed significant revision.
This means that most of the instrument’s contents are relevant, clear, and simple, and the meaning is
clear.</p>
        <p>Zamanzadeh et al. [22] recommended that new instruments have 80 percent agreement or higher
among experts when developing them. The CVI of each item should be considered to determine its
appropriateness. An item is considered appropriate if the CVI is greater than 79 percent. However, it
requires revision if it falls between 70 and 79 percent. The item is eliminated if the CVI is less than 70
percent. The relevant items mean that experts rated it by 3 or 4. This CVI was computed for all four
criteria to carefully examine the items needing revision. The data in table 4 and table 5 also match
and agree with one another. Both ideas in table 4 and table 6 were considered which item to eliminate
or revise. When the item is lower than 2 in table 6 and is to be eliminated in its CVI, it is completely
removed.</p>
        <p>Table 7 shows the sample distribution of items and their CVI under the relevance section. The CVI of
each item is calculated as the number of 3 or 4 ratings divided by the number of experts (N = 10). The
researcher eliminated all items with an average rating of 1.00-1.99 and those with CVI below 0.50. The
ifnal number of items with appropriate and considerable content validity and reliability was 30. The
items that needed revisions were revised based on the comments of the ten experts.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Construct validity of the instrument</title>
        <p>The Kaiser–Meyer–Olkin (KMO) test and Bartlett sphericity test were conducted to ensure that the
assumptions in this validity are met. The KMO value was 0.500. The (Measures of Sampling Adequacy)
MSA value should equal or exceed 0.500 for consideration for further analysis [21]. The Bartlett
sphericity test yielded the Chi-Square test with a p-value of &lt;0.001, lower than the significance value of
0.05. Both tests concluded that a factor analysis could be conducted for factor loading analysis.</p>
        <p>Ten factors were derived from the EFA. Components with eigenvalues greater than 1.0 are separated
into distinct components [23]. These components and their eigenvalues are found in table 8, corroborated
by the scree plot in figure 1. Therefore, this research obtained ten components from the instrument’s
EFA.</p>
        <p>The table shows that component 1 has the greatest eigenvalue in both the unrotated and rotated
solutions, which implies that it explains the highest amount of variance in the data. The portion of the
overall variation in the data that each component can account for can help comprehend the significance
of each component in capturing the fundamental structure of the data. When a component has a more
significant proportion of variance, it indicates that it has a greater impact in explaining the variability
observed in the dataset.</p>
        <p>The ten identified components account for 87.6% of the total variance. This statement suggests that
the integrated assessment tool evaluates a primary element or aspect known as the ability, as stated
in the work of Hambleton, Swaminathan, and Rogers [24]. In this study, the ability is known as the
PC2</p>
        <p>PC3</p>
        <p>PC4</p>
        <p>PC5</p>
        <p>PC6</p>
        <p>PC7</p>
        <p>PC8</p>
        <p>PC9</p>
        <p>PC10
combination of SPS and conceptual understanding.</p>
        <p>Factor loadings indicate how strongly individual items are related to the underlying components and
in which direction [25]. This analysis used a varimax rotation to make the factor structure easier to
understand. The displayed loadings in table 9 show the loads beyond 0.45. The missing loads are below
0.45 and were not shown.</p>
        <p>The factor loadings show how the individual items are associated with each principal component.
Items such as Q6 and Q25 have high loadings on PC1, which suggests they are strongly connected to this
component. On the other hand, Q24 and Q35 have high loadings on PC3, indicating their association
with a diferent underlying dimension. In contrast, Q33 and Q37 have low loadings across all principal
components, which suggests weaker associations with the identified factors.</p>
        <p>Interpreting the factor loadings can help understand the latent with. High loadings on the same
principal component likely measurement principal component are likely measuring a common
underlying construct. For instance, Q6, Q25, Q26, and Q7 all have strong loadings on PC1, suggesting they
contribute to measuring a specific dimension or trait together. Conversely, items with high loadings on
diferent principal components may represent distinct constructs or dimensions.</p>
        <p>The factors were grouped according to their level of cognition. For the researcher, the EFA factor
loads were categorized according to the level of cognition. Since the participants were answering the
instrument with increasing levels of cognitive ability based on the revised Bloom’s taxonomy, the
degree of item connection was seen. These results were similar to the findings of Sadhu and Laksono
[25], who found that their integral assessment instrument in chemistry can measure nine factors. They
added that quantitative skills are the most dominant.</p>
        <p>It was observed that items grouped in one principal component (PC) tend to come from similar items.
For instance, questions 15 and 29 were identified as evaluation questions based on the EFA, but even
before the analysis, these two questions were also classified as evaluation questions. The most dominant
principal component is the PC1, which contributes to eight items, followed by the PC2 with seven items.
This means that the participants’ most dominant scores were from remembering questions.</p>
        <p>Furthermore, these findings show that the instrument can measure the teachers’ combined conceptual
understanding and SPS. Thus, it is worth noting that this instrument is an integrated assessment for
physics. This paper has found that integrating SPS and conceptual understanding in an instrument is
feasible.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Limitations of the test</title>
        <p>Administering a three-tier test in a classroom environment may pose time constraints compared to
traditional assessments, which could be challenging due to the packed curriculum. Interpreting the
results of a three-tier test can be complex and may require advanced statistical methods to diferentiate
between genuine comprehension and surface-level knowledge. Teachers may require additional training
to efectively utilize and understand the outcomes of these tests, which could be a hurdle in some
educational settings.</p>
        <p>
          The second tier of the test depends on students’ self-evaluation of their confidence, but students’
self-perception may not always align with their actual understanding, leading to biased outcomes.
High confidence does not necessarily indicate accuracy, especially when students are unaware of their
misunderstandings ([
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]).
        </p>
        <p>Additionally, the integration of smartphone-based experiments assumes that all students have access
to smartphones or similar technology, which may not be the case. Disparities in the type of smartphones
or the availability of specific apps could impact how students engage in related experiments, potentially
afecting the fairness and reliability of the assessment.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>Based on the instrument’s analysis results, the integrated three-tier multiple-choice test this paper
developed is now considered valid and reliable. The instrument has good internal consistency based on
Cronbach’s Alpha. The revised and retained items were relevant, clear, and simple, and their meanings
were clear.</p>
      <p>The instrument passed the KMO and Barlett’s sphericity test for EFA. The EFA results found that
the instrument measured the teachers’ combined conceptual understanding and science process skills.
However, the most dominant level is the remembering level, while the evaluating level is the lowest.
Therefore, this paper suggests using this instrument to measure the level of integrated physics SPS and
conceptual understanding during a smartphone-based physics experiment.</p>
      <p>This paper recommends evaluating the level of teachers’ SPS and conceptual understanding using
a three-tier test to determine the alternative conceptions the teachers have and devise interventions.
This intervention can be considered during a continuing professional development program. The use of
this instrument with students is also feasible. Moreover, this instrument can be enhanced if the item’s
dificulty index and face validity are measured.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Acknowledgments</title>
      <p>I would like to thank the teachers who participated in the instrument’s pilot testing and final
administration. I also appreciate the experts who thoroughly scrutinized the instrument. Thanks also to Dr.
Reston of USC for the comments and suggestions for the instrument and this paper.
[12] B. S. Bloom, A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of
educational objectives, Longman, 2010.
[13] M. Napal, A. M. Mendióroz-Lacambra, A. Peñalva, Sustainability teaching tools in the digital age,</p>
      <p>Sustainability 12 (2020) 3366. doi:10.3390/su12083366.
[14] I. Mullis, M. Martin, K. Cotter, V. Centurino, TIMSS 2019 Item Writing Guidelines, 2019. URL:
https://timssandpirls.bc.edu/timss2019/methods/pdf/T19-item-writing-guidelines.pdf.
[15] M. O. Martin, I. V. Mullis, M. Hooper, Methods and procedures in TIMSS Advanced 2015, Boston
College, TIMSS &amp; PIRLS International Study, 2016. URL: http://timss.bc.edu/publications/timss/
2015-a-methods.html.
[16] F. A. A. Pacala, Development and Validation of Three-Tier Multiple Choice Test for Conceptual
Understanding in Momentum and Collision, International Journal of Multidisciplinary Approach
&amp; Studies 5 (2018) 1–7. URL: http://ijmas.com/upcomingissue/01.02.2018.pdf.
[17] D. F. Polit, C. T. Beck, S. V. Owen, Is the CVI an acceptable indicator of content validity? Appraisal
and recommendations, Research in nursing &amp; health 30 (2007) 459–467. doi:10.1002/nur.20199.
[18] C. F. Waltz, B. R. Bausell, Nursing research: design statistics and computer analysis, Philadelphia :
F.A. Davis Co., 1981. URL: https://archive.org/details/nursingresearchd0000walt/page/n3/mode/
2up.
[19] H. W. Marsh, B. Muthén, T. Asparouhov, O. Lüdtke, A. Robitzsch, A. J. Morin, U. Trautwein,
Exploratory structural equation modeling, integrating CFA and EFA: Application to students’
evaluations of university teaching, Structural equation modeling: A multidisciplinary journal 16
(2009) 439–476. doi:10.1080/10705510903008220.
[20] Y. Chua, Research methods and statistics book 4: Univariate and multivariate tests, Shah Alam,</p>
      <p>Malaysia: McGraw-Hill Education, 2009.
[21] N. W. D. Ayuni, I. G. A. M. K. K. Sari, Analysis of factors that influencing the interest of Bali State
Polytechnic’s students in entrepreneurship, in: Journal of Physics: Conference Series, volume 953,
IOP Publishing, 2018, p. 012071. doi:10.1088/1742-6596/953/1/012071.
[22] V. Zamanzadeh, A. Ghahramanian, M. Rassouli, A. Abbaszadeh, H. Alavi-Majd, A.-R. Nikanfar,
Design and implementation content validity study: development of an instrument for measuring
patient-centered communication, Journal of caring sciences 4 (2015) 165. doi:10.15171/jcs.
2015.017.
[23] Z. Awang, A handbook on structural equation modelling using AMOS, Universiti Technologi</p>
      <p>MARA Press, Malaysia, 2012.
[24] R. Hambleton, H. Swaminathan, R. Jane, Fundamentals of item response theory, Sage Publications,
1991.
[25] S. Sadhu, E. W. Laksono, Development and Validation of an Integrated Assessment for Measuring
Critical Thinking and Chemical Literacy in Chemical Equilibrium, International Journal of
Instruction 11 (2018) 557–572. doi:10.12973/iji.2018.11338a.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Arista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kuswanto</surname>
          </string-name>
          ,
          <source>Virtual Physics Laboratory Application Based on the Android Smartphone to Improve Learning Independence and Conceptual Understanding</source>
          ,
          <source>International Journal of Instruction</source>
          <volume>11</volume>
          (
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          . doi:
          <volume>10</volume>
          .12973/iji.
          <year>2018</year>
          .1111a.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ismail</surname>
          </string-name>
          , I. Festiana,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hartini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yusal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Malik</surname>
          </string-name>
          ,
          <article-title>Enhancing students' conceptual understanding of electricity using learning media-based augmented reality</article-title>
          ,
          <source>in: Journal of Physics: Conference Series</source>
          , volume
          <volume>1157</volume>
          ,
          <string-name>
            <given-names>IOP</given-names>
            <surname>Publishing</surname>
          </string-name>
          ,
          <year>2019</year>
          , p.
          <fpage>032049</fpage>
          . doi:
          <volume>10</volume>
          .1088/
          <fpage>1742</fpage>
          -6596/1157/3/032049.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Cai</surname>
          </string-name>
          , C. Liu,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          , E. Liu, J.-C.
          <article-title>Liang, Efects of learning physics using Augmented Reality on students' self-eficacy and conceptions of learning</article-title>
          ,
          <source>British Journal of Educational Technology</source>
          <volume>52</volume>
          (
          <year>2021</year>
          )
          <fpage>235</fpage>
          -
          <lpage>251</lpage>
          . doi:
          <volume>10</volume>
          .1111/bjet.13020.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hochberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Becker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Louis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Klein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          ,
          <article-title>Using smartphones as experimental tools-a follow-up: cognitive efects by video analysis and reduction of cognitive load by multiple representations</article-title>
          ,
          <source>Journal of Science Education and Technology</source>
          <volume>29</volume>
          (
          <year>2020</year>
          )
          <fpage>303</fpage>
          -
          <lpage>317</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10956-020-09816-w.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Nikou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Economides</surname>
          </string-name>
          ,
          <article-title>Mobile-Based micro-Learning and Assessment: Impact on learning performance and motivation of high school students</article-title>
          ,
          <source>Journal of Computer Assisted Learning</source>
          <volume>34</volume>
          (
          <year>2018</year>
          )
          <fpage>269</fpage>
          -
          <lpage>278</lpage>
          . doi:
          <volume>10</volume>
          .1111/jcal.12240.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D. T.</given-names>
            <surname>Tiruneh</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. De Cock</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          <string-name>
            <surname>Weldeslassie</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Elen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Janssen</surname>
          </string-name>
          ,
          <article-title>Measuring critical thinking in physics: Development and validation of a critical thinking test in electricity and magnetism</article-title>
          ,
          <source>International Journal of Science and Mathematics Education</source>
          <volume>15</volume>
          (
          <year>2017</year>
          )
          <fpage>663</fpage>
          -
          <lpage>682</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s10763-016-9723-0.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Upper Secondary School Students' Conceptions of Chemical Equilibrium in Aqueous Solutions: Development and Validation of a Two-Tier Diagnostic Instrument</article-title>
          ,
          <source>Journal of Baltic Science Education</source>
          <volume>21</volume>
          (
          <year>2022</year>
          )
          <fpage>428</fpage>
          -
          <lpage>444</lpage>
          . doi:
          <volume>10</volume>
          .33225/jbse/22.21.428.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hanson</surname>
          </string-name>
          ,
          <article-title>The impact of two-tier instruments on undergraduate chemistry teacher trainees: An illuminative assessment</article-title>
          ,
          <source>International Journal for Infonomics</source>
          <volume>12</volume>
          (
          <year>2019</year>
          )
          <article-title>9</article-title>
          . doi:
          <volume>10</volume>
          .20533/iji. 1742.4712.
          <year>2019</year>
          .
          <volume>0198</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>C. M. H. Laeli</surname>
          </string-name>
          , et al.,
          <article-title>The 3 Tiers Multiple-Choice Diagnostic Test for Primary Students' Science Misconception</article-title>
          ,
          <source>Pegem Journal of Education and Instruction</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>103</fpage>
          -
          <lpage>111</lpage>
          . doi:
          <volume>10</volume>
          .47750/ pegegog.13.02.13.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Türkoguz</surname>
          </string-name>
          ,
          <article-title>Investigation of Three-Tier Diagnostic and Multiple Choice Tests on Chemistry Concepts with Response Change Behaviour</article-title>
          ,
          <source>International Education Studies</source>
          <volume>13</volume>
          (
          <year>2020</year>
          )
          <fpage>10</fpage>
          -
          <lpage>22</lpage>
          . doi:
          <volume>10</volume>
          .5539/ies.v13n9p10.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Debbye</surname>
            <given-names>Vitti"</given-names>
          </string-name>
          <source>and "Angie Torres, Practicing Science Process Skills at Home</source>
          ,
          <year>2006</year>
          . URL: https://www.studocu.com/en-za/document/stadio/teaching
          <article-title>-natural-sciences/ teaching-scientific-process-skills/88305872.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>