Valid measures of student motivation can inform the design of learning environments to engage students and maximize learning gains. This study validates a measure of student motivation, the Reduced Instructional Materials Motivation Survey (RIMMS), with a sample of Chinese middle school students using an adaptive learning system in math. Participants were 429 students from 21 provinces in China. Their ages ranged from 14 to 17 years old, and most were in 9th grade. A confirmatory factor analysis (CFA) validated the RIMMS in this context by demonstrating that RIMMS responses retained the intended four-factor structure: attention, relevance, confidence, and satisfaction. To illustrate the utility of measuring student motivation, this study identifies factors of motivation that are strongest for specific student subgroups. Students who expected to attend elite high schools rated the adaptive learning system higher on all four RIMMS motivation factors compared to students who did not expect to attend elite high schools. Lower parental education levels were associated with higher ratings on three RIMMS factors. This study contributes to the field’s understanding of student motivation in adaptive learning settings.
In today's increasingly digital world, it is critical that all students learn to think computationally from an early age. Assessments of Computational Thinking (CT) are essential for capturing information about student learning and challenges. Several existing K-12 CT assessments focus on concepts like variables, iterations and conditionals without emphasizing practices like algorithmic thinking, reusing and remixing, and debugging. In this paper, we discuss the development of and results from a validated CT Practices assessment for 4th-6th grade students. The assessment tasks are multilingual, shifting the focus to CT practices, and making the assessment useful for students using different CS curricula and different programming languages. Results from an implementation of the assessment with about 15000 upper elementary students in Hong Kong indicate challenges with algorithm comparison given constraints, deciding when code can be reused, and choosing debugging test cases. These results point to the utility of our assessment as a curricular tool and the need for emphasizing CT practices in future curricular initiatives and teacher professional development.
The K-12 CS Framework provides guidance on what concepts and practices students are expected to know and demonstrate within different grade bands. For these guidelines to be useful in CS education, a critical next step is to translate the guidelines to explicit learning targets and design aligned instructional tools and assessments. Our research and development goal in this paper is to design a playful, curriculum-neutral assessment aligned with the 'Data and Analysis' concept (grades 6-8) from the CS framework. Using Evidence Centered Design and Participatory Design, we present a set of assessment guidelines for assessing data and analysis, as well as a set of design considerations for integrating data and analysis across middle school curricula in CS and non-CS contexts. We outline these contributions, describe how they were applied to the development of a game-based formative assessment for data and analysis, and present preliminary findings on student understanding and challenges inferred from student gameplay.
Past research suggests revised parallel analysis (R-PA) tends to yield relatively accurate results in determining the number of factors in exploratory factor analysis. R-PA can be interpreted as a series of hypothesis tests. At each step in the series, a null hypothesis is tested that an additional factor accounts for zero common variance among measures in the population. Integration of an effect size statistic-the proportion of common variance (PCV)-into this testing process should allow for a more nuanced interpretation of R-PA results. In this article, we initially assessed the psychometric qualities of three PCV statistics that can be used in conjunction with principal axis factor analysis: the standard PCV statistic and two modifications of it. Based on analyses of generated data, the modification that considered only positive eigenvalues ( π ^ SMC : k ' + Λ ^ ) overall yielded the best results. Next, we examined PCV using minimum rank factor analysis, a method that avoids the extraction of negative eigenvalues. PCV with minimum rank factor analysis generally did not perform as well as π ^ SMC : k ' + Λ ^ , even with a relatively large sample size of 5,000. Finally, we investigated the use of π ^ SMC : k ' + Λ ^ in combination with R-PA and concluded that practitioners can gain additional information from π ^ SMC : k ' + Λ ^ and make more nuanced decision about the number of factors when R-PA fails to retain the correct number of factors.
As K-12 computer science (CS) education initiatives scale throughout the U.S., researchers seek to understand the context-specific relationships between CS instruction and student learning. Evaluation of instruction requires valid measures of curriculum implementation. We have developed measures for identifying conditions for successful implementation of an introductory high school computer science curriculum along two-dimensions: teaching quality and curriculum enactment. Additionally, we have defined three types of instructional strategies for teaching quality. Quantitative and qualitative data were collected from 53 teachers through surveys and interviews. Data were aggregated and integrated to derive scaled measures for the instructional strategies and curriculum adaptation, and implementation measures were correlated with student end-of-unit assessment data. We found potential factors that can enhance or impede the successful implementation of CS curriculum materials, and we have identified several broad issues associated with scaling up CS curricular implementation.
Inferences about student knowledge, skills, and attributes based on digital activity still largely come from whether students ultimately get a correct result or not. However, the ability to collect activity stream data as individuals interact with digital environments provides information about students' processes as they progress through learning activities. These data have the potential to yield information about student cognition if methods can be developed to identify and aggregate evidence from diverse data sources. This work demonstrates how data from multiple carefully designed activities aligned to a learning progression can be used to support inferences about students' levels of understanding of the geometric measurement of area. The article demonstrates evidence identification and aggregation of activity stream data from two different digital activities, responses to traditional assessment items, and ratings based on observation of in-person non-digital activity aligned to a common learning progression using a Bayesian Network approach.
Parallel analysis (PA) assesses the number of factors in exploratory factor analysis. Traditionally PA compares the eigenvalues for a sample correlation matrix with the eigenvalues for correlation matrices for 100 comparison datasets generated such that the variables are independent, but this approach uses the wrong reference distribution. The proper reference distribution of eigenvalues assesses the kth factor based on comparison datasets with k−1 underlying factors. Two methods that use the proper reference distribution are revised PA (R-PA) and the comparison data method (CDM). We compare the accuracies of these methods using Monte Carlo methods by manipulating the factor structure, factor loadings, factor correlations, and number of observations. In the 17 conditions in which CDM was more accurate than R-PA, both methods evidenced high accuracies (i.e.,>94.5%). In these conditions, CDM had slightly higher accuracies (mean difference of 1.6%). In contrast, in the remaining 25 conditions, R-PA evidenced higher accuracies (mean difference of 12.1%, and considerably higher for some conditions). We consider these findings in conjunction with previous research investigating PA methods and concluded that R-PA tends to offer somewhat stronger results. Nevertheless, further research is required. Given that both CDM and R-PA involve hypothesis testing, we argue that future research should explore effect size statistics to augment these methods.
As K-12 computer science (CS) initiatives scale throughout the U.S., educators face increasing pressure from their school systems to provide evidence about student learning on hard-to-measure CS outcomes. At the same time, researchers studying curriculum implementation and student learning want reliable measures of how students apply their CS knowledge. This paper describes a two-year validation study focused on end-of-unit and cumulative assessments for Exploring Computer Science, an introductory high school CS curriculum. To develop the assessments, we applied a principled methodology called Evidence-Centered Design (ECD) to (1) work with various stakeholders to identify the important computer science skills to measure, (2) map those skills to a model of evidence that can support inferences about those skills, and (3) develop assessment tasks that elicit that evidence. Using ECD, we created assessments that measure the practices of computational thinking, in contrast to assessments that only measure CS conceptual knowledge. We iteratively developed and piloted the assessments with 941 students over two years and collected three types of validity evidence based on contemporary psychometric standards: test content, internal structure, and student response processes. Results show that reliability was moderate to high for each of the unit assessments; the assessment tasks within each assessment are well aligned with each other and with the targeted learning goals; and average scores were in the 60 to 70 percent range. These results indicate that the assessments validly measure students' computational thinking practices covered in the introductory CS curriculum. We discuss the broader issues we faced of balancing the need to use the assessment results for evaluation and research, and demands from teachers for use in the classroom.
Two models can be nonequivalent, but fit very similarly across a wide range of data sets. These near-equivalent models, like equivalent models, should be considered rival explanations for results of a study if they represent plausible explanations for the phenomenon of interest. Prior to conducting a study, researchers should evaluate plausible models that are alternatives to those hypothesized to evaluate whether they are near-equivalent or equivalent and, in so doing, address the adequacy of the study's methodology. To assess the extent to which alternative models for a study are empirically distinguishable, we propose 5 indexes that quantify the degree of similarity in fit between 2 models across a specified universe of data sets. These indexes compare either the maximum likelihood fit function values or the residual covariance matrices of models. Illustrations are provided to support interpretations of these similarity indexes.
The objective was to offer guidelines for applied researchers on how to weigh the consequences of errors made in evaluating measurement invariance (MI) on the assessment of factor mean differences. We conducted a simulation study to supplement the MI literature by focusing on choosing among analysis models with different number of between-group constraints imposed on loadings and intercepts of indicators. Data were generated with varying proportions, patterns, and magnitudes of differences in loadings and intercepts as well as factor mean differences and sample size. Based on the findings, we concluded that researchers who conduct MI analyses should recognize that relaxing as well as imposing constraints can affect Type I error rate, power, and bias of estimates in factor mean differences. In addition, fit indexes can be misleading in making decisions about constraints of loadings and intercepts. We offer suggestions for making MI decisions under uncertainty when assessing factor mean differences.
The standardized generalized dimensionality discrepancy measure and the standardized model-based covariance are introduced as tools to critique dimensionality assumptions in multidimensional item response models. These tools are grounded in a covariance theory perspective and associated connections between dimensionality and local independence. Relative to their precursors, they allow for dimensionality assessment in a more readily interpretable metric of correlations. A simulation study demonstrates the utility of the discrepancy measures' application at multiple levels of dimensionality analysis, and compares them to factor analytic and item response theoretic approaches. An example illustrates their use in practice.