<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Psychometrical Modeling of Components of Composite Constructs: Recycling Data Can Be Useful1</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research University Higher School of Economics</institution>
          ,
          <addr-line>Potapovsky Lane 16, build. 10, 101000 Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>This paper describes a list of studies necessary to justify the simultaneous use of both the overall test score and the subscale scores when measuring complex constructs. We investigate in detail one of the strategies for modeling composite constructs, which is popular within the international comparative studies of education. This strategy is based on repetitive recalibrations of the same data using unidimensional models for reporting overall test score and multidimensional models for reporting its components. We use Monte-Carlo simulations to illustrate that repetitive recalibrations of the data using unidimensional and multidimensional models yield, basically, the same results after their transformation to the same scales. However, we also illustrate that the fit of the unidimensional models to the data may be confounded if the components of the composite vary in terms of their relations with each other and their variance. We illustrate the studied strategy for modeling composite constructs using the computer adaptive test PROGRESS-ML, which measures basic math literacy in the third grade.</p>
      </abstract>
      <kwd-group>
        <kwd>Composite Constructs</kwd>
        <kwd>Composite Tests</kwd>
        <kwd>Multidimensional Rasch Models</kwd>
        <kwd>Unidimensional Rasch Models</kwd>
        <kwd>PROGRESS-ML</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Within contemporary educational sciences and broadly, in the social sciences, there is
a growing need for composite measurement instruments – instruments that have a
complex structure, for example, those which consist of subscales that invest in some
way in the overall test score. This may be a consequence of the trend for measuring
complex constructs - such as 21st century skills or new literacies. Such constructs
consist of multiple components, and it is not easy to portrait them as a classic
unidimensional or single-component trait of respondents. It is widely assumed that the
information about the integral trait level is valuable for policymakers, while
information about its components is valuable for practitioners. Such information provides
important insights for improving the performance of, for example, the educational
system or psychological practice at different levels of the social system.</p>
      <p>
        The standards for educational and psychological testing [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] clearly state that (1)
test scores should not be reported to users until their validity, fairness, and reliability
have been studied, and (2) if the test produces more than one test score, the
psychometric quality of all reported scores must be confirmed.
      </p>
      <p>
        This is important because inaccurate information about the overall test score can
lead to decisions with undesirable social consequences, while erroneous information
about subscores can lead to incorrect decisions to correct or improve the situation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
In an academic environment, low-quality subscores can lead to false conclusions
about the nature of the phenomenon being studied.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Psychometrics of composite instruments</title>
      <p>
        In psychometric terms, composite tests are multidimensional. Therefore, the task is to
evaluate, if possible, both the overall ability and its components. Psychometric
modeling of such tests consists of several stages. First of all, a researcher needs to check
whether the test is essentially unidimensional. It is possible to do so by utilizing the
weak definition of local item independence stating that item residual correlations are
zero after extracting a single factor estimated by the unidimensional model (figure 1a)
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. If so, a researcher can report the overall test score - of course, given that it is
proven to be valid and psychometrically consistent. If the test is not unidimensional, it
is necessary to use multidimensional models, and then the overall test score requires
additional research using hierarchical models [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Two types of hierarchical models
are particularly popular – models with higher-order factors (figure 1c) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and bifactor
models [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] (figure 1d). Despite the algebraic similarities [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ] and the fact that
both groups of models assume the use of the overall test score (called the general
factor in factor-analysis terminology), their interpretation is different [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. While
models with higher-order factors estimate the general factor that manifests in items
through subscores, bifactor models assume a complete separation of the general factor
and specific factors.
      </p>
      <p>
        Second, if a researcher intends to report subscores (for example, cognitive
operations or content areas), several approaches are available. The first is to apply the
unidimensional model to each subscale separately [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This approach is called the
"Consecutive Approach". The consecutive approach is the least attractive since the
number of items in each subscale is usually small. Therefore, the measurement
reliability will not be high enough, and the measurement error will be too large. This leads
to the impossibility of reporting subscores [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        The second approach involves the use of bifactor models. These models,
hypothetically, allow simultaneous reports of the overall score and subscores as additional
independent information. However, studies show that subscores estimated in bifactor
models rarely have satisfactory reliability because they describe information not
extracted by the overall score. Therefore, valuable information is often suppressed by
random noise [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Moreover, their interpretation is difficult due to model
assumptions.
      </p>
      <p>Fig. 1. Structural models for modeling composite constructs. Latent variables are drawn using
circles, while observed variables are drawn using squares. One-headed arrows represent
regression dependencies, while two-headed arrows represent correlations</p>
      <p>
        The third approach involves the use of non-compensatory multidimensional
models [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] (correlated traits models or models for between-item multidimensionality
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], figure 1b). Such models represent, essentially, several unidimensional models
combined in a single likelihood equation. This approach is under investigation in this
paper. From a modeling perspective, it is crucial to distinguish this analysis strategy
from the bifactor modeling and consecutive approach. The described approach breaks
the general factor into its parts, proportional to the number of items dedicated to a
particular dimension. Each latent trait is calculated based on respondents' responses to
the corresponding items and considering the latent variables' estimated correlations.
Thus, multidimensional models use information about each dimension and compute
the probability of completing or endorsing an item as a function of several latent
variables, taking into account the relationships between them. As a result, such
measurements' reliability will be greater compared to the consecutive approach. Therefore, it
is more likely that it will be possible to report subscores. At the same time, bifactor
modeling suggests modeling additional subscale-specific components, which add up
to the general factor to produce the observed item scores. Consequently, the
interpretation of the subscale-specific scores from bifactor models is too convoluted for the
most practical tasks. As a result, the application of the bifactor models is mostly
limited to modeling testlet-based assessments and local item dependence conditional on
person parameters.
      </p>
      <p>
        The third analysis strategy illustrates the use of collateral information. Collateral
information is any information about items, respondents, or their interaction, which,
being introduced in the measurement model, does not change the parameters'
interpretation. However, collateral information reduces the uncertainty in the estimates [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
In this case, for each subscale, the responses to all other subscales (together with the
correlation matrix of latent dimensions) are collateral information [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>Thus, to use the results of composite tests, regardless of the chosen strategy of data
analysis, it is necessary to conduct extensive psychometric research. It is necessary to
decide whether the overall test score and subscores are reliable and psychometrically
consistent enough to be reported to users.
2.1</p>
      <sec id="sec-2-1">
        <title>Modeling components of the composite</title>
        <p>
          Breaking the overall test score into its components is popular within cross-national
comparative studies of education. For example, PISA [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] and TIMSS [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] use
repetitive recalibrations of their testing data to decompose the overall test score into the
components, which produce it. TIMSS uses its theoretical framework to report
subscores on cognitive operations required to solve an item. From a statistical point of
view, de facto, it leads to ignoring model-fit indices and recirculation of the data.
Nevertheless, its interpretation allows researchers to describe the composition of the
overall test scores in terms of how respondents achieve those test scores. This enables
policymakers to make decisions based on the information described in terms of social
sciences.
        </p>
        <p>
          However, the difference and equivalence between multidimensional and
unidimensional models is a challenging area of psychometric research. Many studies
have already touched upon the idea of the unidimensional interpretation of
multidimensional measurements. For example, Reckase et al. [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] showed that if the test
items are selected according to specific conditions, the unidimensional model can fit
such data. However, it requires strict guiding the process of test development by the
psychometric parameters of the items. Several researchers have also tried to
conceptualize the fit of the unidimensional models to multidimensional data in terms of the
general factor's strength. For example, Drasgow and Parsons [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] demonstrated that if
the general factor is "strong" (if the factors in the multidimensional model are firmly
positively correlated), then the unidimensional model can fit the data well. Our paper
describes the same phenomenon directly in terms of the correlation matrix of latent
dimensions. Many other researchers studied how model modification can allow the
unidimensional model to fit multidimensional data. The main implication of those
findings is that it is possible to use the overall test score even if the general factor is
weak as long as the multidimensional structure of the data is explicitly modeled [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ].
        </p>
        <p>
          Nonetheless, much research found that the differences between parameter
estimates from multidimensional and unidimensional models are expressed, mainly, in
item parameters. Numerous researches have highlighted unpredictable distortion in
the item parameters estimated when the model's dimensionality is misspecified
regarding the data-generating model [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. However, another conclusion from this
stream of research concerns the stability of the person parameters. As DeMar noted
(although, in another context), "if the focus is on estimated θ's and not on the item
parameters, any of the models will perform satisfactorily" [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. Reise et al. [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]
summarized that the correlation of person parameters from different models tends to be
close to 1 regardless of the model's misspecification.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Simulation study</title>
      <p>To illustrate the possibility of fitting the unidimensional models to the
multidimensional data, we perform a small-scale Monte-Carlo simulation study. We generate the
data under the multidimensional Rasch model and calibrate both unidimensional
(misspecified) and multidimensional (correctly specified) Rasch models on the data.
We then compare the average of the multiple person abilities from the
multidimensional model and the estimated person ability from the unidimensional model. To
compare the results, we used the Pearson linear correlation.</p>
      <p>We also analyze the essential unidimensionality of the simulated data by utilizing
residual analysis. To do so, we apply principal components analysis to the
standardized response residuals under the unidimensional model. This is standard practice for
the analysis of unidimensionality under the Rasch modeling paradigm. This method
rests upon the assumption that if the data is unidimensional (and does not exhibit local
item dependence conditional on person parameters), the residuals are noise, and any
significant principal component cannot be extracted from the data [26; 27]. To
analyze local fit, we used Rasch InFit and OutFit item-wise statistics [28], particularly
their range from the maximum to minimum values. The larger range in InFit and
OutFit means that some items deviate from the model prediction and do not fit the Rasch
model, while smaller variance means that all items fit the Rasch model.</p>
      <p>
        We conduct the simulations for 2000 respondents responding to 30 dichotomous
items, separated into five subscales equally (6 items per subscale). We carry out 100
replications for randomly varying positive definite variance-covariance matrices with
positive manifold (where all latent dimensions are non-negatively correlated). Note,
however, that during random varying of the variance-covariance matrix, we also alter
the variance of latent dimensions. To control this source of the difference of the
results, we also carry out 50 replications for three fixed variance-covariance matrices of
person parameters (where all correlations were equal to 0.80, 0.50, or 0.20, and all
variances are equal). For the randomly varying variance-covariance matrices, we
calculate the difference between correlations by taking the standard deviation of the
values in the lower triangle of the correlation matrix. We do so to analyze the fit of
the unidimensional models conditional on the difference between the
variancecovariance matrix values. Both the multidimensional model and unidimensional
model can be considered as special cases of the Multidimensional Random Coefficients
Multinomial Logit Model [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The quasi-Monte-Carlo algorithm implemented in the
Tam v. 3.5-19 package [29] for the R V. 3.6.2 software was used to estimate all
models.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Results of the simulation study</title>
        <p>The average correlation between person parameters from the unidimensional model
and the average of person parameters from the multidimensional model is 0.99 (p &lt;
0.01) with a standard deviation of less than 0.01 across all simulated conditions.
These results hold for any case – whether the variance-covariance matrix was fixed or
not. This result is in agreement with other similar research, suggesting that person
parameters are more stable in the situation of model dimensionality misspecification.</p>
        <p>Further, the results of dimensionality analysis using PCA on unidimensional model
residuals do vary depending on the size of correlations of latent dimensions in the
data-generating multidimensional model. They suggest that the eigenvalue of the first
component depends on the mean correlation of those dimensions (r = -0.48, p &lt; 0.01,
figure 2) and less depends on differences in the values of correlation matrix (r = -0.25,
p &lt; 0.05, figure 3). Note, however, that the critical value for the first eigenvalue is 2
[30, 31]. Since the first component's eigenvalue is larger than the critical value, all
unidimensional models are critically misspecified for the simulated data, and,
therefore, their results are inconsistent.</p>
        <p>Fig. 2. Scatterplot of eigenvalues of the first component from PCA applied to the standardized
model residuals versus mean correlation of latent dimensions from the simulations with
randomly varying variance-covariance matrices. Each point represents a single simulation
Fig. 3. Scatterplot of eigenvalues of the first component from PCA applied to the standardized
model residuals versus standard deviation of correlations of latent dimensions from the
simulations with randomly varying variance-covariance matrices. Each point represents a single
simulation</p>
        <p>To support these findings, we additionally analyzed the eigenvalue of the first
component from PCA applied to the unidimensional model residuals when the
variance-covariance matrix was fixed. We compared the eigenvalue of the first
component across different values of the fixed correlation and the varied matrix. The results
are presented in figure 4. They also suggest that the analyzed eigenvalue depends on
the size of the fixed correlation. However, they never exceed the critical value of 2.
Therefore, the data with small (or absent) variance in the values of the correlation
matrix of underlying latent factors can be considered unidimensional.</p>
        <p>Fig. 4. Boxplot of the variance of the first component from PCA applied to the standardized
model residuals depending on conditions for simulations</p>
        <p>Next, we analyzed item fit statistics. The results are presented in figure 4. We
compared values of item fit statistics across different conditions of simulations
similarly with previous results. The results are presented in figures 5 and 6. We
discovered similar findings: the range of item fit statistics from unidimensional models in
case of randomly varied variance-covariance matrix exceeds that of the fixed
variance-covariance matrix. However, since the model used for data-generating is the
Rasch model as well as the model used for data analysis, item fit statistics do not react
to differences in item discrimination parameters. Instead, they react to the violation of
unidimensionality, which is expected [27].
Thus, we showed that the unidimensional IRT model could fit the data well even if
the data was actually generated under the multidimensional model. This is fair for the
cases where the values in the correlation matrix of latent dimensions are positive,
correlations are strong, and they do not vary much. However, regardless of that, the
average of the ability estimates from the multidimensional model is equal to the
ability estimate from the unidimensional model. For this, of course, their transformation
to the scales with the same numerical values is necessary (e.g., linear transformation
to the scale N(500,100)). This finding is in agreement with previous studies, which
found that the person parameters are not as sensitive to the model dimensionality
misspecification as the item parameters.</p>
        <p>Nonetheless, psychometric consistency of the unidimensional score can be
confounded if there is variation in the correlation matrix of "true" latent dimensions. If
this is a case, the extraction of the overall test score from multidimensional data
cannot be conducted by averaging the multidimensional model's estimates. Additional
research on "sufficient unidimensionality" of the data is crucial for overall test score
reporting.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Real data example</title>
      <p>This section demonstrates the scope of psychometric studies necessary for reporting
both overall test scores and specific scores, interpreting them as components of the
composite construct. We do so by applying them for the PROGRESS-ML basic
mathematical literacy test.</p>
      <p>The PROGRESS-ML test evaluates how well a student is oriented in mathematics
after completing two years of primary school. When developing the test, we relied on
the following definition of basic mathematical literacy [32]: "basic mathematical
literacy (including working with data) – the ability to apply mathematical tools,
reasoning, and modeling in everyday life, including in the digital environment".</p>
      <p>The PROGRESS-ML basic math literacy test consists of 30 dichotomous items.
The assessment is built as a computerized adaptive test with an automated stopping
rule.</p>
      <p>The content of the test was selected in a way that, on the one hand, it meets the
definition of basic mathematical literacy, and on the other hand, it takes into account the
content of the Russian Federal Educational Standard. As a result, we identified five
content areas: spatial representations, measurement of quantities, regularities,
modeling, and information processing. Test items are grouped into blocks according to the
content area.</p>
      <p>Additionally, the PROGRESS-ML test evaluates students' cognitive processes
required to solve the items. When developing the test items, we used the TIMSS'
theoretical framework for the 4th grade [33]. Therefore, in addition to assessing the
content area, three cognitive operations groups are measured — knowing, application,
and reasoning.</p>
      <p>Thus, the PROGRESS-ML test is a composite tool: it includes five content areas
and reflects three cognitive operations groups. It is assumed that the test results will
report the students' overall test score (in this case, the level of their basic
mathematical literacy), as well as subscores (in this case, content areas and cognitive
operations).</p>
      <p>The sample consisted of 6078 the 3rd grade students from two regions of the
Russian Federation. The samples were representative for the regions. Average age = 9.06
years (SD = 0.46), number of girls = 52.36%.
4.1</p>
      <sec id="sec-4-1">
        <title>Results of the analysis of the real data</title>
        <p>In the analysis of standardized residuals by the PCA, we found that the first
component's eigenvalue is 1.45, which corresponds to 4.2% of the residual variance. The
next four components' eigenvalues are in the interval from 1.15 to 1.2. The
distribution of the explained variance of residuals among the components is almost uniform –
about 4% per component. Therefore, we conclude that the unidimensional model
sufficiently describes the response probability distribution across persons, and the test
can be considered unidimensional.</p>
        <p>The model Expected-a-Posteriori reliability [34] of the entire test score from the
unidimensional model was 0.76. For comparison, we calculated the reliability using
the methods of Classical Test Theory (CTT): Greatest Lower Bound (GLB) [35]
reliability was 0.86, the Cronbach's α [36] was 0.81. However, it is essential to note that
the design of testing (computerized adaptive) implies that not all items are
administrated to all respondents, and the CTT parameters become unstable in the presence of
missing responses. Therefore, even though, in our example, the reliability estimated in
the CTT (both GLB and Cronbach's α) is slightly higher than the reliability of the
scores evaluated in the IRT, these indices should not be trusted.</p>
        <p>Overall, the analysis results suggest that the test can be considered unidimensional,
even though there are different ways to group items. This implies that it is possible to
report one overall test score of mathematical literacy based on the test results, which
will have good reliability and psychometric consistency.</p>
        <p>Then, we calibrated the multidimensional IRT model to estimate if they will have
good psychometric characteristics. The reliability analysis results by content areas are
shown in table 1, by cognitive operations are shown in table 2.</p>
        <p>From the tables, we can conclude that all dimensions have sufficient reliability for
the monitoring test use. Despite the small number of items per subscale, relatively
high reliability is possible due to the approach used for IRT modeling. In fact, such a
small number of items per dimension makes raw subtest scores unusable.
Additionally, we looked at correlations between the latent dimensions: both content areas and
cognitive operations correlate approximately equally – at the level of 0.8-0.9. Based
on the simulation study, we conclude that this can be seen as an additional argument
in favor of the unidimensional model, even though multidimensional models fit the
data better than the unidimensional model according to the AIC [37] and BIC [38]
indices. These indices can estimate the relative model fit to the data introducing a
penalty for extra model parameters (AIC) with respect to sample size (BIC). The
lower values of these indices indicate a better model fit. These indices are presented in
table 3.</p>
        <sec id="sec-4-1-1">
          <title>Model</title>
          <p>Unidimensional
Content areas
Cognitive
operations</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>Deviance 144255.6 143875.7 143965.4</title>
          <p>Thus, multidimensional IRT models allowed us to get reasonably reliable
subscores (for both content areas and cognitive operations) and therefore made it
possible to report them to users. Moreover, the described reliability estimates are
derived from IRT models in which no context variables were entered. Note that the
introduction of these variables into the model (using latent regression modeling) leads
to the estimation of more reliable scores for subscales due to explaining ability
variance.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>
        Contemporary psychometric literature notes the growing popularity of composite tests
designed to produce both the overall test score and the subscores. There are several
strategies for processing such test data. They include the use of raw test scores or the
application of hierarchical models. However, in most cases, raw test scores cannot be
used due to their low reliability [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and hierarchical models require extraordinary
caution in use due to their complex mathematical nature and interpretation.
      </p>
      <p>In this paper, we describe the strategy for modeling subscores as components of the
overall composite test score. This strategy is based upon repetitive recalibration of the
same data using unidimensional and multidimensional models. We demonstrate that
the average of the ability estimates from the multidimensional models is equal to the
ability estimate from the unidimensional model estimated on the same data.
Interestingly, this statement holds regardless of whether or not the unidimensional model fits
the data. However, the application of any statistical models in social sciences needs to
be backed by checking its assumptions and thinking through its theoretical
consequences. Therefore, the unidimensional model's meaningfulness must be argued in
terms of both model fit and construct definition. As we demonstrate in our simulation
study, the unidimensional model does not always fit the data despite the equivalence
of its estimates to the average of the estimates from the multidimensional models.
This means that the unidimensional model's adequacy needs to be verified either way
if a researcher intends to follow the described approach in modeling the composite
constructs.</p>
      <p>We also provide an example of the described strategy for modeling the composite
constructs using the PROGRESS-ML basic mathematical literacy test. We
demonstrate that the use of IRT models allows us to report the respondent's overall test score
and subscores with respect to test specification. For this test, the main result of testing
is the respondent's overall test score. However, repeated recalibration of data based on
content areas and cognitive operations groups required for solving items allows us to
report subscores on those dimensions. These estimates possess greater reliability and
simpler interpretation than estimates from other approaches to modeling composite
constructs. The essence of these results is the decomposition of the overall test score
into the components that make it up.
27. Smith, E.V.: Detecting and evaluating the impact of multidimensionality using item fit
statistics and principal component analysis of residuals. Journal of Applied Measurement, 3,
205–231 (2002).
28. Linacre, J.M.: What do infit and outfit, mean-square and standardized mean. Rasch
Measurement Transactions, 16(2), 878, (2002).
29. Robitzsch, A., Kiefer, T., Wu, M. Package ‘TAM’. Test Analysis Modules – Version:
3.519, (2020).
30. Raîche, G.: Critical eigenvalue sizes in standardized residual principal components
analysis. Rasch measurement transactions, 19(1), 1012 (2005).
31. Linacre, J.M.: Winsteps® Rasch measurement computer program User's Guide.
Beaverton, OR: Winsteps.com (2018).
32. Фрумин, И.Д., Добрякова, М.С., Баранников, К.А., &amp; Реморенко, И.М.
Универсальные компетентности и новая грамотность: чему учить сегодня для успеха завтра.
Предварительные выводы международного доклада о тенденциях трансформации
школьного образования (2(19); Современная Аналитика Образования) (2018).
33. Mullis, I.V., &amp; Martin, M.O.: TIMSS 2019 Assessment Frameworks. International
Association for the Evaluation of Educational Achievement, Amsterdam, The Netherlands
(2017).
34. Bock, R.D., &amp; Mislevy, R.J: Adaptive EAP estimation of ability in a microcomputer
environment. Applied psychological measurement, 6(4), 431-444 (1982).
35. Jackson, P.H., &amp; Agunwamba, C.C.: Lower bounds for the reliability of the total score on
a test composed of non-homogeneous items: I: Algebraic lower bounds. Psychometrika,
42(4), 567-578 (1977).
36. Cronbach, L. J.: Coefficient alpha and the internal structure of tests. Psychometrika, 16(3),
297-334 (1951).
37. Akaike, H.: A new look at the statistical model identification. IEEE transactions on
automatic control, 19(6), 716-723 (1974).
38. Schwarz, G.: Estimating the dimension of a model. The annals of statistics, 6(2), 461-464
(1978).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. American Educational Research Association, American Psychological Association, and
          <article-title>National Council on Measurement in Education: Standards for educational and psychological testing</article-title>
          . American Educational Research Association, Washington, DC (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Sinharay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puhan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Haberman</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          <article-title>An NCME instructional module on subscores</article-title>
          .
          <source>Educational Measurement: Issues and Practice</source>
          ,
          <volume>30</volume>
          (
          <issue>3</issue>
          ),
          <fpage>29</fpage>
          -
          <lpage>40</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hattie</surname>
          </string-name>
          , J.:
          <article-title>Methodology review: Assessing unidimensionality of tests and items</article-title>
          .
          <source>Applied Psychological Measurement</source>
          ,
          <volume>9</volume>
          ,
          <fpage>139</fpage>
          -
          <lpage>164</lpage>
          (
          <year>1985</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Yung</surname>
            ,
            <given-names>Y.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thissen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McLeod</surname>
            ,
            <given-names>L.D.</given-names>
          </string-name>
          :
          <article-title>On the relationship between the higher-order factor model and the hierarchical factor model</article-title>
          .
          <source>Psychometrika</source>
          ,
          <volume>64</volume>
          (
          <issue>2</issue>
          ),
          <fpage>113</fpage>
          -
          <lpage>128</lpage>
          (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gignac</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          :
          <article-title>Higher-order models versus direct hierarchical models: g as superordinate or breadth factor? Psychology Science</article-title>
          ,
          <volume>50</volume>
          (
          <issue>1</issue>
          ),
          <volume>21</volume>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Holzinger</surname>
            ,
            <given-names>K.J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Swineford</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>A study in factor analysis: The stability of a bi-factor solution</article-title>
          .
          <source>Supplementary educational monographs</source>
          (
          <year>1939</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Reise</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          <article-title>The rediscovery of bifactor measurement models</article-title>
          .
          <source>Multivariate behavioral research</source>
          ,
          <volume>47</volume>
          (
          <issue>5</issue>
          ),
          <fpage>667</fpage>
          -
          <lpage>696</lpage>
          , (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Leiman</surname>
            ,
            <given-names>J.M.:</given-names>
          </string-name>
          <article-title>The development of hierarchical factor solutions</article-title>
          .
          <source>Psychometrika</source>
          ,
          <volume>22</volume>
          (
          <issue>1</issue>
          ),
          <fpage>53</fpage>
          -
          <lpage>61</lpage>
          (
          <year>1957</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Rijmen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>Formal relations and an empirical comparison among the bi‐ factor, the testlet, and a second‐ order multidimensional IRT model</article-title>
          .
          <source>Journal of Educational Measurement</source>
          ,
          <volume>47</volume>
          (
          <issue>3</issue>
          ),
          <fpage>361</fpage>
          -
          <lpage>372</lpage>
          , (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Brunner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagy</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wilhelm</surname>
            ,
            <given-names>O.:</given-names>
          </string-name>
          <article-title>A tutorial on hierarchically structured constructs</article-title>
          .
          <source>Journal of personality</source>
          ,
          <volume>80</volume>
          (
          <issue>4</issue>
          ),
          <fpage>796</fpage>
          -
          <lpage>846</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mansolf</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Reise</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          :
          <article-title>When and why the second-order and bifactor models are distinguishable</article-title>
          .
          <source>Intelligence</source>
          ,
          <volume>61</volume>
          ,
          <fpage>120</fpage>
          -
          <lpage>129</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Davey</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Hirsch</surname>
          </string-name>
          , T.M.
          <article-title>: Concurrent and Consecutive estimates of examinee ability profiles</article-title>
          .
          <source>Paper presented at the Annual Meeting of the Psychometric Society</source>
          , New Brunswick, NJ. (
          <year>1991</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Haberman</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sinharay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Reporting of subscores using multidimensional item response theory</article-title>
          .
          <source>Psychometrika</source>
          ,
          <volume>75</volume>
          (
          <issue>2</issue>
          ),
          <fpage>209</fpage>
          -
          <lpage>227</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Reckase</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          <article-title>Multidimensional item response theory models</article-title>
          .
          <source>In Multidimensional item response theory</source>
          (pp.
          <fpage>79</fpage>
          -
          <lpage>112</lpage>
          ). Springer, New York, NY (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <given-names>M.</given-names>
            , &amp;
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.C.</surname>
          </string-name>
          :
          <article-title>The multidimensional random coefficients multinomial logit model</article-title>
          .
          <source>Applied psychological measurement</source>
          ,
          <volume>21</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp; Cheng, Y.:
          <article-title>Improving measurement precision of test batteries using multidimensional item response models</article-title>
          .
          <source>Psychological Methods</source>
          ,
          <volume>9</volume>
          (
          <issue>1</issue>
          ),
          <fpage>116</fpage>
          -
          <lpage>136</lpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tam</surname>
            ,
            <given-names>H.P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Jen</surname>
            ,
            <given-names>T.H.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Multidimensional IRT Models in Book: Educational measurement for applied researchers</article-title>
          .
          <source>Theory into practice.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Scaling PISA</surname>
          </string-name>
          <article-title>Data (Chapter 9)</article-title>
          .
          <source>In: PISA 2018 Technical Report. OECD</source>
          , Paris (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Foy</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>Scaling the TIMSS 2015 Achievement Data</article-title>
          . In: Martin,
          <string-name>
            <given-names>M.O.</given-names>
            ,
            <surname>Mullis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.V.</given-names>
            , &amp;
            <surname>Hooper</surname>
          </string-name>
          , M. (eds.)
          <article-title>Methods and</article-title>
          procedures in
          <source>TIMSS 2015</source>
          , pp.
          <volume>13</volume>
          .
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .62. TIMSS &amp; PIRLS International Study Center, Boston College, Chestnut Hill, MA (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Reckase</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ackerman</surname>
            ,
            <given-names>T.A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          :
          <article-title>Building a unidimensional test using multidimensional items</article-title>
          .
          <source>Journal of Educational Measurement</source>
          <volume>25</volume>
          (
          <issue>3</issue>
          ),
          <fpage>193</fpage>
          -
          <lpage>203</lpage>
          (
          <year>1988</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Drasgow</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Parsons</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Application of unidimensional item response theory models to multidimensional data</article-title>
          .
          <source>Applied Psychological Measurement</source>
          ,
          <volume>7</volume>
          ,
          <fpage>189</fpage>
          -
          <lpage>199</lpage>
          (
          <year>1983</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Ip</surname>
            ,
            <given-names>E.H.</given-names>
          </string-name>
          :
          <article-title>Empirically indistinguishable multidimensional IRT and locally dependent unidimensional item response models</article-title>
          .
          <source>British Journal of Mathematical and Statistical Psychology</source>
          ,
          <volume>63</volume>
          (
          <issue>2</issue>
          ),
          <fpage>395</fpage>
          -
          <lpage>416</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Steinberg</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Thissen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Uses of item response theory and the testlet concept in the measurement of psychopathology</article-title>
          .
          <source>Psychological Methods</source>
          ,
          <volume>1</volume>
          ,
          <fpage>81</fpage>
          -
          <lpage>97</lpage>
          (
          <year>1996</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>DeMars</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          <article-title>Application of the bi-factor multidimensional item response theory model to testlet-based tests</article-title>
          .
          <source>Journal of Educational Measurement</source>
          ,
          <volume>43</volume>
          ,
          <fpage>145</fpage>
          -
          <lpage>168</lpage>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Reise</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cook</surname>
            ,
            <given-names>K.F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>T.M. Evaluating</given-names>
          </string-name>
          <article-title>the impact of multidimensionality on unidimensional item response theory model parameters</article-title>
          . In Reise,
          <string-name>
            <given-names>S.P.</given-names>
            , &amp;
            <surname>Revicki</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.A</surname>
          </string-name>
          . (eds.)
          <article-title>Handbook of item response theory modeling</article-title>
          . Routledge, New York (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Linacre</surname>
            ,
            <given-names>J.M.:</given-names>
          </string-name>
          <article-title>Structure in Rasch residuals: why principal components analysis</article-title>
          .
          <source>Rasch measurement transactions</source>
          ,
          <volume>12</volume>
          (
          <issue>2</issue>
          ),
          <volume>636</volume>
          (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>