
While high school grades are widely used for university admissions, little is known about which specific high school grades best predict what type of performance at university. This study examines the predictive value of overall high school GPA (grade point average), grades for subsets of subjects, and the added value of having taken specific subjects, for university performance across different cognitive learning objectives, different programmes and over time. Using data from multiple cohorts of six undergraduate programmes at a large Dutch research university, we show that the overall high school GPA consistently outperforms subsets of discipline-related subjects, suggesting that high school grades primarily represent general learning skills and traits. However, having taken a specific related high school subject was generally associated with better university performance, although effect sizes were small. High school grades predicted performance better on assessments targeting lower-order cognitive skills than complex academic tasks. No significant differences emerged between the predictive value of high school final-year and penultimate-year grades. Finally, the predictive strength declined over the course of the three-year bachelor programmes. These findings highlight the need for careful consideration of which high school grades to use in admissions and provide practical suggestions for university admissions officers to do so.
Higher education sectors worldwide have largely adopted a rhetoric of assessment redesign, calling for fundamental 'rethinking', 'reforming' and 'reimagining' of assessment. This rhetoric is particularly evident in the so-called age of Generative Artificial Intelligence (GenAI). But what about grades? In this conceptual study, we address the elephant in the room, grading: a large looming presence that is nonetheless absent from conversation about assessment reform. By using the surge of GenAI technologies as the context for assessment reforms, we analyse recent responses in grading policy and practice to GenAI technologies by asking what has changed, if anything. We discuss how two approaches to grading in an age of GenAI - AI-proof grading and AI-powered grading - both further reinforce the status quo of grading systems. Next, we conduct a policy review and note a critical silence regarding grades. Based on these two analyses, we showcase that entire programs of assessment are being reworked with little to no thought given to grades. Our study aims to generate broader, sector-wide conversations about how higher education institutions should represent student achievement.
Student evaluations of teaching can be biased by the 'way' the instructor speaks or their accent. In the current study, we investigate whether a brief awareness-based intervention can be effective in reducing accent-based bias in students' evaluations of instructors. In a series of controlled laboratory experiments, college student participants were randomly assigned to watch a short lecture narrated by either an instructor who spoke English with a Mandarin accent or an American accent. Prior to evaluating the instructor, participants in the intervention condition were given instructions designed to reduce bias, whereas participants in the control condition proceeded directly to the evaluation form. Although there were no statistically significant differences in performance on a multiple-choice assessment of learning, participants in both experiments showed evidence of bias, rating the Mandarin-accented instructor more negatively than the American-accented instructor. In Experiment 1, the intervention reduced the disparity in participants' overall ratings of the two instructors; however, in Experiment 2, the intervention was less effective with different lecture materials and a more diverse sample of students. Awareness-based interventions may show promise in reducing real-world biases in college classes and could support efforts to improve diversity and inclusivity in higher education.
Peer review serves as the structural foundation of scientific integrity, yet the system currently faces unprecedented strain. In response, AI tools are being deployed to assist human expert review, a development that has generated considerable debate. However, the psychological impact of this transition on the scientific community remains only partially understood. Our study utilises a sequential mixed-methods approach to examine the perceptions of AI-mediated feedback among a cohort of elite researchers (Nature and Science authors). Quantitative findings from our randomised experimental survey (N = 495) demonstrated that AI-assisted reviews were viewed as deficient in fairness, usefulness, and acceptance compared to human-led evaluations. Qualitative evidence from 47 in-depth interviews further identified a dual-layered aversion to AI-assisted feedback, concerning both the technology (AI use) and the agent (AI user). Specifically, scholars perceived AI-generated critiques as devoid of the necessary disciplinary nuance required for high-stakes evaluation. Moreover, a notable 'AI user aversion' emerged: reviewers who delegated tasks to AI were perceived as lacking the diligence and empathic engagement essential to the peer-review contract. Together, these finding suggest that the integration of AI into peer review may erode trust in the research evaluation process and journals should implement robust governance that preserves human oversight.
Evaluative judgement, the ability to know what good work looks like and act on that understanding, is increasingly recognised as a pedagogical goal in higher education. However, without a consolidated evidence base, efforts to develop students' evaluative judgement remain scattered, making it difficult to tell which approaches are effective, what problems remain, and what requires further study. To address this gap, this paper maps how evaluative judgement has been conceptualised, measured, and implemented in higher education through a scoping review of pedagogical approaches that support students' judgement. Across the studies, evaluative judgement was most often framed as a competency and examined through outcome-focused measures, with evidence typically reflecting teacher or researcher intentions rather than what students do during learning activities. In practice, evaluative judgement was most supported through purposefully structured learning designs that created authentic evaluative experiences. Future research should adopt a wider range of theoretical framings and prioritise process-focused evidence that captures how students form and apply evaluative judgements in practice, including considering the contextual influences that shape their decisions.
Teacher feedback literacy plays a crucial role in effective feedback processes in higher education. Scholars have recently developed theoretical frameworks to guide research in this area. However, there remains a gap in empirical research which explores these frameworks within authentic educational contexts. In response, this study uses a prominent teacher feedback literacy framework to examine feedback practices in a Taiwanese university English for academic purposes (EAP) writing context. Specifically, this research explores critical considerations in the relationship between teacher feedback literacy and learner uptake, with findings that extend beyond current framings of this concept. Employing a comparative case study methodology of two teachers (Teacher A and Teacher B), multiple data sources were analysed to explore how differences in how teachers' approaches to feedback literacy influenced learner uptake over a single writing cycle. Findings highlight the impact and trade-offs of different approaches on uptake, learner engagement, and learning. Research implications include contributions to the ongoing dialogue on enhancing feedback literacy, with propositions to better align theoretical frameworks with diverse educational practices and authentic contexts. By exploring practical applications of teacher feedback literacy, this research also offers insights that may improve both teaching and learning in higher education EAP settings.
Project-Based Assessment (PBA) is an integrated assessment paradigm widely applied in STEM and other interdisciplinary fields. It holds significant potential for fostering students' higher-order thinking skills and enhancing their comprehensive abilities. However, research in this field remains fragmented and lacks a systematic synthesis. This study addresses this gap by conducting a scoping review of 75 articles aiming to elucidate core concepts, diagnose implementation practices, and provide strategic guidance for future development. The synthesis revealed that PBA is a comprehensive assessment paradigm grounded in constructivism, which integrates multidimensional assessment functions based on authentic problems and process-oriented assessment. Essentially, PBA is positioned as a learning-centered assessment. The findings indicated that PBA is predominantly implemented in higher education and applied disciplines, primarily targeting higher-order cognitive objectives through mixed-method assessment tools. While PBA demonstrates a positive trend in fostering learner engagement, its effectiveness is often hindered by heavy implementation burdens and concerns regarding assessment reliability. Furthermore, several critical deficiencies were identified, including a teacher-centric power structure, the absence of initial assessment phases, and insufficient iterative improvement. Consequently, to enable the effective implementation of PBA, future efforts should focus on reconstructing evidence frameworks, redefining teacher roles, and refining practice models.
This study investigates how graduate educators conceptualize and apply the AI Assessment Scale (AIAS) within their assessment design practice. Drawing on an exploratory qualitative case study design, the study analyzed eight AI-integrated lesson plans produced by in-service teachers in a technology elective course, supplemented by semi-structured interviews with four participants. Findings suggest that AIAS level selection is experienced not as a technical classification but as an ethical boundary-setting practice, a judgment about delegating evaluative responsibility between human and AI agents. Participants demonstrated variably developed AI assessment literacy: procedural ethics (integrity, authorship) and experiential ethics (learner agency) were more fully articulated than structural ethics (algorithmic bias, data governance). Constructive alignment was strongest when AI was constitutively embedded in the design; conversely, AI integration could shift assessment criteria from content-driven toward procedurally driven evaluation, a drift originating at the design-imagining stage before any tool was deployed. Notably, designing without AI tools heightened participants' awareness of habitual AI dependence and, in several cases, increased confidence in unassisted design, suggesting that AI assessment literacy may require experiential constraint as well as conceptual instruction. Implications are discussed for teacher education, AIAS professional development, and AI assessment literacy.
Authentic assessment has become increasingly central in higher education, reflecting a shift away from traditional, decontextualised testing toward assessment practices emphasising meaningful learning, integrated competence and the application of knowledge in context. Despite its growing prominence, authentic assessment remains conceptually fragmented, encompassing task-focused, competence-oriented, and more recent embedded and future-oriented interpretations. This article presents a comprehensive conceptual review tracing the evolution of authentic assessment from its early performance-oriented formulations in the late 1980s to contemporary conceptualisations emphasising sustainability, digital mediation, ethical engagement with emerging technologies and lifelong learning capabilities. The article synthesises historical and theoretical developments to examine how authenticity has been reframed in response to changing educational priorities and institutional contexts. It also highlights key debates in authentic assessment and tensions surrounding employability, equity, standardisation, digitally mediated learning environments and academic integrity. By providing an analytically structured account of the conceptual evolution of authentic assessment, the article clarifies conceptual ambiguities, situates contemporary developments within a broader historical trajectory and provides a foundation for future empirical research on the mechanisms, implementation and outcomes of authentic assessment across diverse higher education contexts.