
Traditional assessments (e.g. multiple-choice quizzes and essays) fail to meet the diverse needs of students and are increasingly prone to academic integrity risks. This study compared students' perceptions of traditional assessments with two innovative assessment types: Reflective E-portfolios and Interactive Oral Assessments (IOAs). Participants were 85 undergraduate biopsychology students, who completed Assessment Experience and Perceived Academic Misconduct scales, the University Needs Instrument, and provided qualitative responses. Students with high academic and emotional support needs rated E-portfolios more positively than IOAs, multiple choice quizzes, and written assessment tasks. IOAs, followed by E-portfolios were viewed as less vulnerable to academic misconduct than traditional tasks. Students had initial polarized reactions to innovative tasks, followed by task enjoyment, and highlighted the importance of engaging teaching methods and learning support. Our findings support the need for more innovative assessment that supports diverse student learning needs, while maintaining academic integrity.
Existing formative assessments have key limitations for evaluating PST skill early in teacher preparation program coursework. This study examines the validity of performance tasks as feasible-to-use assessments providing reliable and actionable information about PSTs' instructional skills. Participants were 153 PSTs from two universities participating in a mathematics methods course. Participants completed a set of three performance tasks targeting a focal skill and filmed themselves performing this skill during classroom instruction. All tasks were double coded by trained raters on a proximal measure of the focal skill and distal measures of instructional quality as defined by general and special mathematics educators. ANOVA and Generalizability Theory (GT) analyses were conducted to investigate score accuracy, consistency, bias, and generalizability over different conditions. Regression analyses were conducted to examine how performance task scores relate to PSTs' classroom performance. Study results indicated that performance tasks were reliable and valid measures of PST skill.
This study examined the validity of a computerized adaptive testing (CAT) version of a multidimensional test of reading comprehension that produces two scores: an overall reading comprehension score and a diagnostic classification. The Multiple-choice Online Causal Comprehension Assessment, or MOCCA, uses several features to render scores across two measurement dimensions. We sought evidence that CAT increased testing efficiency by reducing test time and length without sacrificing measurement precision, while also producing better diagnostic classifications and equivalent convergent validity compared to a fixed-item test (FIT). Using a blocked randomized control trial design, we compared CAT and FIT versions on several dependent variables related to testing efficiency, precision, and validity. CAT reduced average test length and testing time while improving measurement precision for the reading comprehension score and improving diagnostic classifications. Results also revealed similar correlations between reading comprehension scores and criterion measures, and evidence of the predictive utility of both scores.
When responding to a low-stakes assessment, students may not provide effortful responses due to the lack of personal consequences which threatens the validity of interpretations of assessment results. To mitigate this threat, it is critical to identify patterns among disengaged students. Understanding these patterns provides insight into the behavioral dynamics of disengaged students and can be useful for interventions aimed at promoting student engagement. By investigating disengaged students at the test-level, we classified them into two newly proposed classifications: accelerated responders and non-accelerated responders. Accelerated responders reach an item such that they begin to proportionally rapidly guess more often for all subsequent items, whereas non-accelerated responders do not respond with this pattern. We found accelerated responders constitute a large majority of students disengaged at the test level. These findings explain the extent to which students disengage in a low-stakes assessment and may inform intervention strategies aimed at mitigating non-effortful responding.
This study complements TIMSS 2023 proficiency reporting by applying cognitive diagnostic modeling to examine Grade 8 science mastery across eight systems. Using a framework-aligned Q-matrix, we estimated country-specific cognitive diagnostic models (restricted DINA parameterization) on 16 dichotomized items (N = 18,131). Attribute-level EAP mastery and latent profiles were linked to TIMSS IRT-based domain scale scores, reported via plausible values with jackknife replication. Models converged stably, yielding 8-13 profiles per system. Environmental/Applied Reasoning was most frequently mastered. Profile-based scale means increased monotonically with the number of mastered attributes, and top - bottom gaps reached 250 points in some systems. Results show that CDM-derived mastery profiles and IRT-based scale scores are complementary: scale scores indicate how much proficiency is demonstrated, whereas CDM clarifies where and how mastery is distributed across framework-aligned attributes. The study provides a replicable, population-scaled diagnostic approach to inform curriculum sequencing, equity monitoring, and cross-system hypothesis generation.
Using log-file data from a graduate admissions test, we examined whether test-takers from different demographic backgrounds show similar test-taking behaviors and whether these behaviors relate to test performance (correct items) similarly across groups. Data from 118,657 U.S. test-takers showed generally small demographic differences in behaviors, such as the average time spent on initial visits and the number of revisited and revised answers, with the largest differences occurring in item revisitation. Behavior - performance associations varied by demographic group and section (quantitative vs. verbal), but were small-to-moderate in magnitude. In contrast, performance gaps were larger and ranged from small-to-moderate to moderate-to-large across demographic comparisons. These patterns - small behavioral differences coupled with substantial performance gaps - suggest that although test-taking behaviors matter, they do not fully explain performance disparities and other factors contribute to observed gaps. The implications of these findings for understanding performance outcomes on high-stakes assessments are discussed.
Scholars who hope to measure teachers' mathematics instructional practice at scale often face a choice between self-report surveys, which face drawbacks associated with reporting accuracy, and classroom observations, which possess stronger validity but potentially lower reliability and greater expense. This article presents an alternative. In it, we describe a method for measuring teachers' inclination to engage with three mathematics instructional practices - eliciting student thinking, using student thinking, and supporting student sensemaking - by presenting teachers with animated vignettes each featuring an embedded student idea, asking teachers to voice what they would say next in response to that idea, and then to report what instructional steps they would take following that initial response. We describe the prevalence of these three instructional moves in teacher responses to the six vignettes as well as the factor structure and reliability of scales derived from coding this data.
Gradeless assessment is often introduced with the aim of reducing pressure associated with grades. This study investigates how various factors predict differences in students' perceived grade-related pressure in the context of gradeless assessment. Using an exploratory design, we conducted a regression analysis of survey data from 1,708 students across seven Norwegian upper secondary schools. The findings indicate that: (1) consistent implementation of gradeless assessment is associated with lower perceptions of grade-related pressure, (2) students' perceptions of feedback emerge as a strong predictor of lower grade-related pressure, and (3) students' perceptions differed across educational tracks. The results highlight the importance of consistent assessment practices, feedback quality, and contextual factors in shaping students' experiences of grade-related pressure. The study contributes to ongoing discussions on gradeless assessment by underscoring the need for a nuanced understanding of grade-related pressure in gradeless contexts. Implications for practice, policy, and future research are discussed.
During COVID-19 lockdowns, teachers worldwide had to quickly adapt to emergency remote teaching, impacting their instruction and assessment practices. This transition resulted in increased perceived workloads for teachers. Through a moderated mediation model using data from 3675 Portuguese teachers, this study found that the relationship between teachers' adaptation to remote teaching and their perceived workloads was mediated by instruction and assessment. Moreover, changing assessment methodologies, as proposed in an emergency assessment policy, moderated this relationship. However, for special education teachers, these changes did not moderate the relationship, suggesting that their challenges extended beyond assessment adjustments. These findings emphasize the unique difficulties faced by special education teachers and the importance of avoiding a one-size-fits-all assessment policy in future disruptive events.
The Next Generation Science Standards (NGSS) calls for assessments that can measure students' three-dimensional (3D) science learning. In this study, we developed and collected validity evidence for 3D assessment tasks that assess students' ability to reason about energy disciplinary core ideas using science and engineering practices and crosscutting concepts across three grade bands. We field tested the tasks with 3446 elementary, middle, and high school students. Validity evidence based on test content, response processes, internal structure, and relations to other variables was collected. The evidence supports the use of these tasks as measures of students' ability to reason about energy as recommended by NGSS. The field test data revealed that while many middle and high school students demonstrated understanding of the elementary-level NGSS expectations, many have not yet met the NGSS expectations for their own grade levels. This finding likely reflects the uneven implementation of NGSS-aligned instruction.
Large-scale assessments (LSAs) primarily support system-level monitoring, but their instructional diagnostic potential remains underused. Cognitive Diagnostic Models (CDMs) offer a promising avenue, though their application in LSAs poses theoretical and practical challenges. This study explores whether cognitive models used in item generation can directly inform Q-matrix construction for CDM analyses, enhancing diagnostic value in early numeracy assessment. Using data from Luxembourg's school monitoring program (N = 35,058), we analyzed four cognitive attributes (counting, addition < 10, decomposition, addition > 10) using developmental frameworks. We compared a Single-Attribute Hierarchical Model, assuming linear progression, with a Multiple-Attribute Hierarchical Model, allowing skill interactions. Both hierarchical models reproduced expected developmental progressions, with decomposition emerging as a key threshold skill and socio-economic status showing the largest subgroup differences. Subgroup analyses revealed a smaller-than-expected impact of migration background, while math anxiety peaked at intermediate skill levels. Embedding cognitive models in LSAs can bridge system-level monitoring and instructional support.
University teachers are primarily responsible for conducting assessments in their classrooms. Their assessment skills are very important in the assessment process. This survey investigated the relationship between public university teachers' perceived formative assessment skills and assessment practices based on self-reported data. An adapted version of the Assessment Practices Inventory (API) was distributed to a random sample of 380 teachers from four public universities in Ethiopia that were randomly selected. Third-order partial correlation and canonical correlation analyses were used to examine the relationship. The study found that teachers' perceptions of their formative assessment skills were significantly correlated with their assessment practices, as perceptions influence actions. Additionally, teachers' perceived formative assessment skills in various areas of assessment were strongly correlated with their practices in those areas. Therefore, the study suggests that the formative assessment skill requirements of teachers should be addressed to enable them to effectively deal with their classroom assessment tasks.
Assessing the personal and professional skills alongside the academic ability of applicants is vital for holistic admissions practices. While these soft skills can be assessed via interviews, the time and resources required to conduct them mean that they are often impractical to offer to every applicant. Situational judgment tests (SJT), particularly open-response SJTs, offer a resource-friendly method for assessing these skills early in the admissions process prior to deciding who should interview. We explored the extent to which a new SJT format, an open-response SJT with both typed- and video-response items, could predict future interview performance using admissions data from 1,011 applicants to two US medical school programs. Correlation (r=0.37-0.48) and logistic regression results (OR = 1.668-2.595) indicate that the updated SJT format has an increased ability to predict future interview performance relative to the previous format which consisted solely of typed-response items.
The design of instructionally meaningful science assessments has improved significantly in recent years, but rubrics have not evolved in parallel. Poor alignment with interpretation and feedback systems can lead these assessments to fall short of their goals for enhancing instruction. Although there is consensus about essential qualities of three-dimensional (3D) science assessments, there is no such agreement about features of quality rubrics. By translating assessment design principles into criteria for rubric design, we ensure coherence between the interpretation model and the assessment. These criteria were used to analyze existing rubrics to identify advantageous features that formed the core of a new rubric model, the Learning-Centered Rubric (LCR). The LCR provides actionable, equitable, efficient, and asset-based information about progress with 3D learning. This work brings overdue attention to the need for transparency in rubric design and for innovation in rubric development that matches the flourishing advancement of science assessments aligned to practice.
This study explores the feasibility and potential advantages of using a computer adaptive testing (CAT) paradigm for surveys consisting of items that most respondents endorse, such as when assessing socio-emotional outcomes. Such constructs oftentimes produce item responses concentrated at one end of the Likert scale (in psychometric parlance, none of the items are "difficult"). Perhaps given these challenges, CAT versions of socio-emotional surveys are rarely if ever produced. As a result, researchers could be missing an opportunity to harness advantages of CAT, which include shorter but more reliable and more personalized measures. By simulating data using Monte Carlo methods, this study compares CAT with traditional fixed-form surveys, investigating factors such as item pool size, item difficulty, and test length on bias and precision of scores. The results indicate that CAT significantly improves score precision while reducing the number of items administered.
Extensive research connects enhanced formative assessment to the use of rubrics. However, their value in improving student performance is mixed. This study investigated the impacts of rubrics on students' academic performance on quadratic equation problems. Using a quasi-experimental design, 149 grade eleven students in six classes were randomly sampled from three secondary schools in a rural district (Zambia). Three classes were randomly assigned into experimental groups (n = 78) and used rubrics to obtain feedback that encouraged self-assessments for six weeks in the second term of 2024. The same formative assessments were implemented in three comparison groups (n = 71) without rubrics. Data were collected using a quadratic equations test and analyzed using a t-test. The students from experimental groups ($M = 47.27, SD = 24.21$M=47.27,SD=24.21) significantly outperformed those from comparison groups ($M = 24.54, SD = 19.48$M=24.54,SD=19.48) after the intervention regardless of gender. Thus, by providing clear expectations and promoting self-assessment, rubrics can significantly improve male and female students' academic performance on quadratic equation problems.
This paper presents and illustrates a framework for visualizing large-scale assessment results in a dynamic score reporting interface to support teachers in making content-referenced interpretations of student growth. The reporting interface maps student performance to locations along a research-based learning progression to facilitate interpretations of the quantitative differences along a vertical score scale. We illustrate content-referenced growth reporting in the context of a learning progression for how students understand fractions. Two key aspects of the illustration include evidence showing that discrete levels of the learning progression have a moderate to strong association with the difficulty of assessment items that were coded to the levels, and results from a small-scale pilot test of the reporting interface with practicing classroom teachers. We speculate about important aspects of implementation that would strengthen the validity of CRG score reporting interpretations and uses.
Response anchoring, the tendency for respondents to provide item responses equivalent (or in close proximity) to immediately preceding responses, can reflect disengaged responding and/or cognitive challenges in comprehending test items. Using Grade 4-12 student responses to a survey measure of four social-emotional learning constructs collected from the CORE Districts in the state of California, we examine student variables that predict anchoring, and consider potential score bias and survey design implications. Response anchoring is studied using both an index of sequential response separation as well as through an IRT-based anchoring index recently considered in the literature. Both approaches show anchoring behavior increases at higher grade levels; the IRT index, however, uniquely shows anchoring to be reduced for students higher on the SEL constructs, as might be theoretically anticipated, suggesting its greater validity. Both methodological and practical implications of the analyses are considered in discussion.