Clinical psychology is a discipline reliant on self-reports but uniquely susceptible to specific biases associated therewith. Here we provide a prototype for objective behavioral assessment drawn from the field of alcohol science, describing an emerging class of wearable transdermal biosensor. We note the challenges of transdermal alcohol assessment and describe recent performance gains from updated devices and machine learning analytic tools. We indicate unanswered questions for transdermal technology, including device longevity and the accuracy of devices for producing fine-grained estimates of drinking quantity. We identify factors that can impede development of transdermal sensors and other new objective measures, including the tendency to judge new tools against an implicit ideal, and consider scientific findings divorced from methodology. Finally, in evaluating novel objective measurement tools, we argue for careful consideration of not only error magnitude but also error type (i.e., random versus systematic), identifying measurement diversification as a priority for clinical psychology moving forward.
College-level writing instructors are under growing pressure to determine how, or if at all, to integrate generative AI tools into their curricula. While some theoretical scholarship promotes generative AI integration as a means of fostering AI literacy, empirical research on how writing instructors are approaching these tools remains limited. This study explores how college-level writing instructors are currently addressing generative AI in their classrooms, what pedagogical rationales and concerns shape decisions to integrate or avoid these tools, and how instructors evaluate the effectiveness of their own approaches. This study combines survey data from 79 college-level writing instructors across a range of institution types with follow-up semi-structured interviews using an explanatory sequential mixed methods design. Survey results revealed a pronounced division among instructors, with respondents nearly evenly split between those who integrate generative AI in some form (43
As large language models (LLMs) are increasingly integrated into college-level writing instruction, it is important to examine whether AI-mediated learning outcomes are equitably distributed across students. This study investigates demographic fairness in Learning by Teaching (LbT) with an LLM, in a college-level writing context with 66 undergraduate students. Each student completed both an LbT condition with an LLM-simulated student and a video lecture condition designed to approximate a traditional lecture format. Learning was measured at multiple levels of Bloom’s Taxonomy using immediate post-tests, delayed post-tests, and post-lesson essays assessing rhetorical strategy application. Test-based assessments showed overall learning gains from pre- to post-test across conditions, but no statistically significant differences across demographic groups. Essay-based analyses, however, revealed selective instructional effects: female students applied significantly more rhetorical strategies on average under the LbT condition (2.88) than under the video condition (2.24), a difference not reflected in test-based measures. These findings indicate that assessment modality influences conclusions about fair learning outcomes across students.
Research exploring correlates of, precursors to, and consequences of psychological disorders has often relied on designs wherein both predictor and outcome are measured by self-reports. In this article, coauthored by a clinical psychologist (C. E. Fairbairn) and a data scientist (N. Bosch), we offer information surrounding an evolving class of machine-learning models as these inform an expanding measurement tool kit in clinical-psychological science. Specifically, we note the development of deep-learning applications for image analysis, language analysis, and the analysis of physiological time-series data, reviewing implications of these advances for measurement in behavioral research. We weigh strengths and limitations of these automated methods in comparison with self-reports, including the specific form of error likely yielded via each (random vs. systematic), with the aim of fostering a replicable, sustainable, and reputationally strong field of clinical-psychological science.
OBJECTIVE:Although mobile health-tracking technologies have burgeoned, offering objective health information to consumers on an unprecedented scale, opportunities to directly test effects of such monitoring have been limited. Low-cost mobile breathalyzers are one tool commonly employed for blood alcohol concentration (BAC) assessment. The authors explored outcomes linked with BAC-tracking technologies, examining effects on alcohol use and self-estimation of BAC levels in a large U.S. sample. METHODS:Participants (N=32,179) were individuals who voluntarily purchased a mobile breathalyzer and provided at least three ad-lib readings between 2016 and 2022. A paired smartphone application prompted users to enter a BAC self-estimate (a guess) before the measured BAC level was displayed. Analyses included observations collected during active consumption (BAC >0.00%) from breathalyzer users who opted to share anonymized data. Breathalyzer users who displayed inattentive patterns of guessing were excluded from self-estimation analyses. The final dataset comprised 787,393 BAC readings and 387,643 self-estimates. RESULTS:The accuracy of BAC guesses increased by 2.38% over the course of breathalyzer use. Associations between breathalyzer use and BAC levels varied significantly according to participants' initial drinking levels (b=-0.0062, 95% CI=-0.0065, -0.0059). Among heavy-drinking participants, BAC levels decreased on average from 0.106% to 0.096%, whereas the reverse trend was observed for lighter-drinking participants, whose levels increased from 0.058% to 0.067%. A similar interaction emerged for BAC underestimation (b=-0.0058, 95% CI=-0.0066, -0.0049), with odds of underestimation decreasing among heavy-drinking and increasing among light-drinking participants. CONCLUSIONS:The results indicate promise for mobile BAC-tracking technologies as a low-impact intervention with the potential to decrease drinking among individuals who drink heavily-a population particularly susceptible to alcohol-related problems. In contrast, inverted trends emerged for light-drinking individuals, highlighting the need for empirical research in the fast-moving landscape of digital health.
As AI tools increasingly reshape educational activities through applications such as intelligent tutoring systems, early warning systems, and automated essay scoring, ensuring fairness in these AI-assisted decision-making scenarios remains a critical challenge. While researchers have been addressing AI fairness in education, existing efforts remain largely algorithmically focused and less integrated with broader socio-political ideas of fairness. As generative AI enables significantly more diverse and dynamic applications in education, research on fairness needs to take more holistic and collaborative perspectives. The Fair4AIED 2026 workshop aims to bring together researchers, practitioners, and policymakers from the AIED, EDM, and L@S communities to foster interdisciplinary dialogues and develop a joint vision on fairness in decision-making for education. The half-day workshop features a keynote presentation from outside the education community, research talks on fairness-focused research, and a moderated panel discussion on emerging themes such as technical bias, downstream harms, unfairness mitigation, and fairness in generative AI. The workshop welcomes both experts and novices in fair AI. Thus, providing a platform for translating theoretical fairness frameworks into actionable strategies for educational systems. The workshop website is available at https://fair4aied.github.io/2026/ .
Students often misjudge their understanding of learning material, which can lead to the use of ineffective learning strategies and result in suboptimal learning outcomes. However, it remains unclear how misjudgments relate to the use of metacognitive strategies in online learning settings, which is essential context for developing effective interventions that support students in making (and using) accurate judgments of their performance. To address this, we analyze data from 210 college students using a computer-based learning environment, investigating the relationships among calibration discrepancy, judgments, and strategies, as well as the factors affecting shifts in metacognitive judgments during learning. Students who overestimated their pretest retrospective judgments engaged less in metacognitive strategies, particularly in preparatory actions before quizzes (b = -9.100, p < .001). Notably, pretest retrospective judgments—rather than actual pretest scores—significantly predicted students’ engagement in these metacognitive strategies (b = -9.841, p < .001). Furthermore, increased engagement in repeated quiz-taking was a significant negative predictor of changes in metacognitive judgments (b = -1.792, p = .036), indicating that students engaging in repeated quizzes tended to adjust their judgments more conservatively. These results highlight the role of pretest retrospective judgments in shaping engagement with metacognitive strategies, underscoring the importance of correcting early calibration discrepancies. Our findings advocate for early, proactive metacognitive support tools that go beyond merely presenting information, offering guidance on interpreting feedback, and implementing strategies to better align students’ judgments with their actual performance.
Uncovering algorithmic bias related to sensitive attributes is crucial. However, understanding the underlying causes of bias is even more important to ensure fairer outcomes. This study investigates bias associated with Attention Deficit Hyperactivity Disorder (ADHD) in a machine learning model predicting students’ test scores. While fairness metrics did not reveal significant bias, potential subtle bias indicated by variations in model performance for students with ADHD was observed. To uncover causes of this potential bias, we correlated SHapley Additive exPlanations (SHAP) values with the model’s prediction errors, identifying the features most strongly associated with increasing prediction errors. Behavioral and self-reported survey features designed to measure students’ use of effective learning strategies were identified as potential causes of the model underestimating test grades for students with ADHD. Behavioral features had a stronger correlation between absolute SHAP values and prediction errors (up to r =.354, p =.013) for students with ADHD than for those without ADHD. Students with ADHD often use unique yet effective approaches to studying in online learning environments—approaches that may not be fully captured by traditional measures of typical student behaviors. These insights suggest adjusting feature design to better account for students with ADHD and mitigate bias.
This paper studies a bidirectional human-AI collaborative student performance prediction problem to enhance equitable online education, aligning with the United Nations' Sustainable Development Goal (SDG) of ensuring inclusive and equitable quality education for all. The goal is to leverage collaborative intelligence to generate accurate and fair student outcome predictions from behavioral data, ensuring equitable estimation for underrepresented populations. Current fair AI solutions often fail to mitigate demographic bias in the absence of student demographic data, while human-AI collaborative approaches frequently overlook human cognitive biases, leading to inaccurate predictions. We develop CollabDebias, a novel bidirectional human-AI collaborative framework that utilizes the complementary strengths of AI and humans to mitigate the AI demographic bias and human cognitive bias. To address AI demographic bias, we propose an uncertainty learning-based bias identification method and a reliability-aware human-AI integration approach. To reduce human cognitive bias, we design uncertainty-aware visualization of AI decision area and attention mechanism. Experimental results on an online course demonstrate CollabDebias's effectiveness in improving student performance prediction accuracy and fairness.
Students struggle with accurately assessing their own performance, especially given little training to do so. We propose an AI-powered training tool to help students improve “metacognitive calibration,” or the ability to accurately predict their own learning, potentially enhancing learning outcomes by enabling students’ use of metacognition-informed learning behaviors. We present results from a randomized controlled trial (N = 133) assessing the effectiveness of the tool in a college-level computer-based learning environment. The AI-driven tool significantly improved learning gains compared to the control group by 8.9% (t = -2.384, p =.019), and this effect was significantly mediated by learning behaviors. Overconfident students who received the intervention showed significantly greater metacognitive calibration improvement than the control group by 4.1% (t = 2.001, p =.049). These insights highlight the value of AI-powered metacognitive calibration training and the importance of promoting specific metacognition-informed learning behaviors in computer-based learning.
Adaptive learning systems are increasingly common in U.S. classrooms, but it is not yet clear whether their positive impacts are realized equally across all students. This study explores whether nuanced identity categories from open-ended self-reported data are associated with outcomes in an adaptive learning system for secondary mathematics. As a measure of impact of these social identity data, we correlate student responses for 3 categories: race and ethnicity, gender, and learning identity-a category combining student status and orientation toward learning-and total lessons completed in an adaptive learning system over one academic year. Results show the value of emergent and novel identity categories when measuring student outcomes, as learning identity was positively correlated with mathematics outcomes across two statistical tests.