
The growing use of generative AI (GenAI) tools like ChatGPT raises serious concerns about academic integrity, especially in the context of assessments. While detection efforts have focused on open-ended responses, multiple-choice questions (MCQs), which are common in high-stakes testing, remain largely overlooked, partly due to their perceived detection difficulty. The present work establishes a theoretical and empirical foundation for applications of psychometric theory to separate GenAI and human responses. Specifically, it investigates whether person-fit statistics (PFS), a class of methods within Item Response Theory (IRT) used to evaluate how well an examinee’s response pattern fits the expectations of the IRT model, can distinguish between GenAI and human responses to MCQ assessments. To study this, we use data from two authentic assessment contexts: a high-school level chemistry test and the national university entrance exam, each with approximately 1000 human respondents. Our results demonstrate that PFS reveal significant differences between human responses and those generated by advanced chatbots (ChatGPT, Claude, and Gemini), which appear as ‘aberrant’ respondents. We also demonstrate that different chatbots present significantly different response patterns, suggesting that they should be treated as a heterogeneous group of ‘intelligences’ rather than as a single one. Using the PFS measures, we also demonstrate, however, that newer GenAI versions not only improve in performance but also become more ‘human-like’ in their response patterns. Together, these results position IRT as a robust framework for characterizing and separating human and GenAI response patterns in MCQ assessments, providing a theoretical foundation and empirical evidence.
Large language models (LLMs) are being integrated into nursing education at a pace that outstrips the field’s capacity to evaluate their consequences. Existing syntheses have mapped what LLMs can do in nursing contexts. Among those identified in this review, none has examined how their integration affects the professional development processes through which nursing expertise is formed. This critical integrative review addresses that gap.Following Whittemore and Knafl’s (2005) five-stage framework, we systematically searched four databases (PubMed/MEDLINE, Web of Science, Scopus, CINAHL) for studies published between January 2017 and December 2025, yielding 489 included studies from 47 countries/regions. Two reviewers independently screened all 1,182 records; inter-rater agreement was κ = 0.67 at the title and abstract stage and κ = 0.78 at the full-text stage. Findings were analysed through an integrated theoretical framework combining Benner’s model of professional skill acquisition, cognitive load theory (Sweller, 1988), automation bias (Parasuraman & Manzey, 2010), and Wenger’s participatory account of identity formation.Three principal findings emerged. First, LLMs address real structural gaps in nursing education—including deficits in individualised learning support and clinical simulation access—but their benefits are consistently moderated by implementation framework, economic equity, and model-specific accuracy limitations. Second, LLM integration is associated with both competency enhancement and competency attrition, depending on whether the displaced cognitive work is extraneous to or constitutive of the competence being developed; unstructured AI reliance was associated with measurable deficits in ethical reasoning and individualised clinical judgement. Third, an Evidence Gap Map of the 489-study corpus shows that the domains of greatest policy consequence are supported by the least rigorous evidence. We identified no randomised or quasi-experimental studies of professional identity development, only three quasi-experimental studies of relational and ethical competency, and no studies with follow-up beyond 12 months, whereas controlled evidence is concentrated in cognitive and technical outcomes.This review introduces the Professional Identity Tension Model as a preliminary conceptual framework for distinguishing LLM integration that supports professional formation from integration that silently substitutes for it, and proposes the concept of Structural Empathy Suppression to reframe the conditions under which AI appears to outperform nurses on relational metrics. LLMs in nursing education both fill structural gaps and shape professional identity; the field now requires research designs and implementation frameworks capable of distinguishing between the two.
Mathematical modelling (MM) is widely recognized as a key competency for addressing complex real-world problems, which requires learners to engage in higher-order processes such as abstraction, representation, and iterative reasoning. However, students often experience difficulties in coordinating these processes within mathematics modelling tasks. While recent advances in artificial intelligence (AI) offer new possibilities for supporting mathematical modelling, limited research has examined how different forms of AI-mediated interaction shape students' learning processes in MM. To address this gap, the present study investigates five AI interaction roles, Tutor, Teaching Assistant, Peer, Excellent Student, and Struggling Student, as distinct pedagogical scaffolding strategies in AI-assisted learning environments. To evaluate the effectiveness of these roles, we conducted a randomized within-subjects experiment with 26 university students to compare modelling competency, learners' role preferences, and learning experiences across the five conditions. A randomized within-subjects experiment with 26 university students was conducted to examine differences in modelling performance and learners’ role preferences across conditions. The findings reveal a notable divergence between mathematical modelling performance and pedagogical roles preference. Students demonstrated higher modelling competency when interacting with Peer and TA roles, which fostered collaborative reasoning and co-construction of ideas. In contrast, students expressed stronger preferences for Tutor and Excellent Student roles, which provided more explicit guidance and structured explanations. The Struggling Student role was consistently perceived as least supportive across both performance and preference measures. These findings examine how interactional structures influence students’ engagement in mathematical modelling. The study highlights the importance of balancing cognitive scaffolding with opportunities for collaborative sense-making, and provides design implications for developing adaptive, learner-centered AI-supported modelling environments in mathematics education.
The Community of Inquiry (CoI) framework is widely used to guide and explain learning in online and blended environments. However, generative artificial intelligence (GenAI) challenges a key assumption in CoI research and practice: that indicators of presence can be attributed primarily to human learners, instructors, and their interactions. GenAI can produce fluent discourse that resembles strong cognitive, social, or teaching presence. Yet such discourse may not reflect learners’ thinking, relationships, or regulatory intentions. To date, research often frames GenAI as a tool, a dialogic partner, or a possible fourth presence. These views, however, do not fully capture GenAI’s pervasive influence on inquiry processes. This paper therefore reconceptualizes CoI in the age of GenAI by positioning GenAI as an epistemic condition. This framing highlights that GenAI can shape inquiry through direct use, indirect mediation of learning tasks and expectations, and the latent influence of its availability on how contributions are interpreted, attributed, and evaluated. In this view, GenAI shapes what appears plausible, what counts as evidence, and how contributions are attributed, thereby altering how critical thinking is enacted and recognized in a community of inquiry. Drawing on sociomaterial and postdigital perspectives, we argue that CoI presences are sociotechnical accomplishments emerging through human-GenAI assemblages, rather than through human activity alone. The paper accordingly reinterprets cognitive, social, and teaching presence and proposes a configuration-based heuristic suggesting that the relationship between GenAI involvement and inquiry quality is conditional on human accountability. We conclude with implications for theory, practice, and future research, suggesting that CoI evaluation (e.g., coding schemes) may need refinement. In particular, indicators of presence can no longer be inferred from polished final texts alone. Instead, they require process-sensitive evidence, such as prompting and revision traces, disclosure and attribution practices, verification moves, and interaction logs that make the epistemic work of inquiry visible.
Generative artificial intelligence (AI) is rapidly integrating into design education, yet its effects on student learning processes remain empirically underexplored. This study reports a sequential mixed-methods investigation across a built-environment faculty at a research-intensive university, drawing on 24 faculty interviews, a survey of 32 instructors with regression modelling, and five clustered discussions involving 31 participants.Faculty accounts converge on a pattern in which AI amplifies existing differences in student readiness rather than producing uniform improvements. Consistent with cognitive offloading theory, students with stronger foundations use AI to extend reasoning and accelerate iteration, while those with weaker foundations delegate formative cognitive work to AI, producing coherent outputs without corresponding understanding. This divergence is compounded by a loss of process visibility, as AI-generated outputs no longer reliably indicate how students arrived at their work, and by the erosion of frictional learning stages through which competence is built. Faculty have responded with adaptive strategies including process-oriented assessment redesign and deliberate reintroduction of cognitive constraint, but report no scalable solution.Exploratory quantitative analysis (n = 32) identifies perceived pedagogical relevance rather than seniority or career stage as the primary predictor within the model of faculty AI positivity. A significant negative association between theoretical course orientation and AI positivity suggests that uniform integration strategies risk producing uneven outcomes across course types. Effective AI integration in design education requires differentiated strategies responsive to course epistemology and student preparation level, assessed through process-visible methods rather than output quality alone.
The rapid rise of generative artificial intelligence (GenAI) tools such as ChatGPT is transforming the landscape of higher education. Beyond their immediate use in writing support and tutoring, these tools are driving a more profound transformation in pedagogy, curriculum design, and the foundational structures of learning itself. Drawing on a triangulated mixed-methods design, this study integrates a scoping review, bibliometric mapping (VOSviewer, n=209 records), a systematic literature review of 36 peer-reviewed articles (2023–2025), and ten semi-structured interviews with academic leaders and educators across five institutions in Australia and Indonesia. The central aim of the study is to develop and propose an AI-Augmented Learning System framework that conceptualises GenAI not merely as an instructional tool but as a catalyst for curricular and pedagogical reconfiguration. Thematic patterns reveal five interrelated system shifts: from static curricula to dynamic AI-integrated design; from teacher-centred delivery to AI-augmented facilitation; from knowledge transmission to capability development; from local experimentation to institutional governance; and from fragmented implementations to ecosystemic integration. These shifts are interpreted through established educational theories including constructivism, connectivism, TPACK, the SAMR model, and constructive alignment to clarify the pedagogical mechanisms through which GenAI is reshaping curriculum and teaching, and to surface implications for learning analytics and educational innovation. Given the bounded empirical base, the proposed framework is offered as an analytical heuristic and starting point for institutional dialogue rather than a prescriptive blueprint, providing a foundation for further empirical validation across diverse higher education contexts.
Automated analysis of classroom discourse has traditionally relied on discriminative Transformer architectures like RoBERTa to classify isolated utterances, an approach that inherently struggles with long-horizon dependencies and exhibits significant algorithmic bias against speakers of non-standard dialects of American English. Off-the-shelf models, however, often fail to align with pedagogical frameworks such as Dialogic Instruction and Asset-Based Pedagogy. To address these limitations, this paper introduces a Neuro-Symbolic Pedagogical Alignment (NSPA) framework that leverages Large Language Models (LLMs) within a Judge-Critique-Refine Direct Preference Optimization (DPO) loop to quantify high-inference educational constructs like Student Reasoning and Teacher Uptake. We specifically propose a novel Dialect-Invariant Contrastive Learning objective that utilizes style-transfer data augmentation to decouple semantic reasoning from surface-level linguistic variation, thereby mitigating the deficit framing often encoded in standard models. Evaluated on 1660 elementary mathematics lessons from the National Center for Teacher Effectiveness transcript corpus, our framework improves the detection of complex reasoning chains by 14.2 percentage points, measured by the macro-averaged harmonic mean of precision and recall, compared to state-of-the-art discriminative baselines. Furthermore, the proposed debiasing mechanism reduces the false-negative rate for contributions voiced in African American Vernacular English by 18.4 percentage points, ensuring a more equitable measurement of epistemic agency across diverse student demographics. Finally, we validate the ecological utility of these metrics by establishing a statistically significant Pearson correlation () with value-added models of teacher effectiveness, confirming that automated, equity-aware discourse analysis can serve as a rigorous proxy for learning outcomes.
To address the limitations of existing automatic evaluation systems that rely on surface features and fail to provide actionable feedback for improving discourse coherence, we propose a Transformer model based on BERT to classify discourse relations (causal, contrastive, progressive, or incoherent) between adjacent sentences, thereby identifying coherence defects at the sentence-pair level and supporting precise feedback generation. In specific implementation, English writing samples are subjected to standardized preprocessing and manual annotation to form sentence pairs for sequential input. Subsequently, the pre-trained BERT-based truncated model was fine-tuned, updating only the last four Transformer layers and the two fully connected scoring layers, and the discourse relation classification ability was optimized by weighted cross-entropy; We then uses attention weights and semantic similarity thresholds to locate breakpoints, and combines rule libraries to identify missing connections and ambiguous references; Finally, structured feedback is generated based on high scoring sample essays, and after grammar verification and teacher knowledge graph constraints, editable diagnostic reports are output to achieve a closed-loop process from defect detection to educational feedback. The experimental results show that the average F1 score of the proposed method in long essays is not less than 0.891, the average accuracy is not less than 0.902, and the mAP ranges from 0.758 to 0.918 across different defect types and essay genres. The system demonstrates high discriminative ability in discourse relation classification and precise defect localization, thereby effectively supporting coherence improvement; The average adoption rate of feedback ranges from 71.2% to 88.4%, and the average Ref-BLEU-4 ranges from 6.5% to 12.3%, indicating that the generated suggestions exhibit moderate lexical alignment with expert-approved reformulation examples, complementing the teacher adoption ratings as a measure of feedback quality; This study has to some extent bridged the gap between deep semantic modeling and instructional operability, and can provide a technical reference for intelligent writing feedback systems.
Generative artificial intelligence (GenAI) can enhance students' efficiency and the apparent quality of academic work, yet its fluent outputs may encourage learners to outsource cognitive activities central to learning. This creates the risk of “performance without learning”, whereby successful task completion does not necessarily reflect meaningful understanding or learner regulation. To address this concern, this article develops the Intent, Deconstruction, Expression, and Adaptation (IDEA) framework, a theory-informed metacognitive scaffold for GenAI-supported academic work. IDEA guides learners to articulate goals and constraints, decompose tasks into essential processes, communicate requirements through structured prompts, and critically evaluate and iteratively refine AI-generated outputs. By embedding prompting within a cycle of planning, monitoring, and regulation, IDEA provides an actionable approach for preserving learner agency in GenAI-supported work. An exploratory quasi-experimental pilot study with 42 undergraduates examined the feasibility and preliminary educational value of IDEA-based instruction relative to structured prompt-engineering instruction. Across immediate GenAI-assisted academic tasks, IDEA-trained students produced higher-quality prompts and final AI-generated outputs, and their interaction records showed observable enactment of the framework's core activities. On unaided tasks completed five days later, the IDEA group demonstrated advantages on selected tasks. Together, the conceptual framework and pilot findings position IDEA as a practical instructional approach for transforming GenAI use from passive content outsourcing into deliberate, evaluative, and learner-regulated interaction. Future research can examine its operation across disciplinary contexts and its longer-term implications for learning.
As AI enters schools, the staff it affects are increasingly urged to consult conversational AI about whether to adopt it—yet that AI is built by organizations with a commercial stake in adoption, raising the possibility that systems consulted by skeptical users are predisposed to encourage it. We test this with a fixed prompt in which a rural Montana K–12 school administrative aide voices two concerns: that AI may threaten her job, and that AI companies do not have people like her in mind. Ten frontier models from ten laboratories are each sampled 500 times at temperature 0.7. The 5000 responses are scored by a blind three-model AI panel (three model families, two openness regimes) on a four-dimension rubric validated against researcher hand-scoring (quadratic-weighted Cohen’s κ 0.704–0.932 at consensus, clearing a pre-specified κ≥0.70 threshold). Mean consensus composite scores range from 3.85 (Claude Sonnet 4.6) to 7.52 (Gemini 3.1 Pro Preview) on an eight-point scale. The variation concentrates not in whether concerns are recognized (every model acknowledges them) but in what follows: eight of ten models redirect toward AI engagement, upskilling, or adaptation, often explicitly advocating adoption of the technology the persona named as a threat. This is neither the deference the sycophancy literature predicts nor the demographic divergence the political-bias literature documents; it is concern acknowledgment followed by systematic redirection, which the cross-model spread establishes as a model-dependent design outcome—not an inevitable property of large language models as a category. Because school staff increasingly consult such systems about adoption, the finding bears on trust calibration and AI literacy in education.
Integrating artificial intelligence into medical imaging offers potential for dental education. While existing models automate diagnosis, they often lack the interpretive depth needed for comprehensive student training. To address these limitations, this paper presents Gen-Mentor, a human-in-the-loop instructional framework that integrates the DentDiff-VLM backbone into a dental-radiography workflow. The backbone uses Faster R-CNN to localize four target radiographic findings: Filling, Implant, Impacted Tooth, and Cavity. A conditional diffusion model supports curriculum expansion by generating class-specific synthetic ROI candidates as candidate instructional assets. A vision–language model (VLM) generates evidence-linked caption candidates, which a large language model (LLM) reformats into candidate case descriptions, comparisons, and quiz prompts. Selected candidate instructional assets then undergo structured expert review. We evaluate Gen-Mentor across technical performance, expert review of instructional assets, and learner acceptance among dental students (N=45). The framework achieved a mean System Usability Scale score of 72.7, with improvements in case diversity and immediate-feedback support.
Generative AI-powered agentic systems are increasingly proposed for educational applications. However, the research landscape remains fragmented, and the relationship between technical capabilities and pedagogical design remains poorly understood. To address this gap, this scoping review systematically mapped 474 studies published between January 2020 and May 2026. Guided by three research questions, we analysed publication characteristics, study designs, agent roles, AI models and architectures, six dimensions of agentic capability, and the extent to which educational theory was incorporated.The results show that the field has expanded rapidly, particularly since 2025. At the same time, the literature is dominated by conference papers and is primarily concentrated in higher education, STEM disciplines, and text-based tutoring scenarios. In terms of technical implementation, GPT-series models and LangChain are the most widely adopted technologies, whereas OpenClaw and other frontier agent paradigms remain largely absent. Across the reviewed studies, agentic capabilities tend to remain at relatively modest levels: although many systems demonstrate single-task autonomy, sequential planning, and, increasingly, multi-agent collaboration, they rarely exhibit strong tool orchestration or robust embedded governance.From an educational perspective, theoretical grounding remains limited. Only 138 studies explicitly drew on educational theory, revealing a clear disciplinary divide between technically oriented research and pedagogically oriented work. Methodologically, empirical evaluations are also limited, with most studies relying on small-scale and short-term designs. Accordingly, the gaps identified across the literature converge on several priorities: longitudinal and real-world validation, stronger pedagogical grounding, more governed adoption of emerging agent infrastructures, and more systematic integration of ethics and human oversight.Overall, this review provides researchers, developers, and educators with an evidence-based map of the current capabilities and limitations of agentic AI in education, while also highlighting concrete directions for its more responsible and educationally meaningful development.
The adoption of Generative AI (GenAI) and Large Language Models (LLMs) in healthcare education has accelerated rapidly since 2023, yet the evidence base for their use in Scenario-Based Learning (SBL) remains fragmented. This systematic review synthesises empirical research on GenAI applications across scenario-based, case-based, problem-based, and simulation-based learning in healthcare education. Following the PRISMA 2020 guidelines, five databases were searched on 9 November 2025 for peer-reviewed studies published from January 2023 onwards. Of 1151 initial records, 23 studies met the inclusion criteria. Quality was assessed using the Mixed Methods Appraisal Tool (MMAT). Thematic synthesis identified six cross-cutting themes organised around a central finding: prompt design in educational contexts functions as a form of instructional specification, encoding the cognitive targets and quality criteria that would otherwise be implicit in expert authoring. Yet only 34.8% of studies aligned generated content with established instructional frameworks, and an equal proportion reported prompting strategies in sufficient detail for reproduction. Additional findings include: variable validation practices lacking standardisation; superior outcomes from hybrid human-AI collaboration over fully automated approaches; persistent evaluation gaps in longitudinal and comparative designs; and scalability as a primary adoption driver despite largely unquantified efficiency gains. GPT-4 dominated implementations (44.4%), while open-source alternatives were under-explored. Educational outcomes were generally positive for higher-order cognitive skills but inconsistent for knowledge acquisition. A four-stage validation pipeline is proposed as a conceptual framework to guide responsible deployment, pending empirical validation. These findings suggest that GenAI integration in healthcare SBL requires treating prompt design as a methodological element, standardising multi-stage validation, and formalising human-AI collaboration to realise its educational potential.
This systematic review is a synthesis of 32 empirical and conceptual studies published between December 2022 and March 2026 to investigate AI literacy among language teachers in higher education. The review has answered four research questions related to conceptualisations of AI literacy, pedagogical practices and professional-development models, assessment approaches, and reported outcomes, challenges, and gaps guided by the theoretical lens of governing the unseen, which prefigures institutional policies. Based on PRISMA 2020, the systematic search of ERIC, the British Educational Index, and Web of Science resulted in the identification of studies, which were subjected to thematic synthesis, based on the critical policy and sociomaterial lenses. The results show that AI literacy is largely conceptualised in terms of competency-based, multi-dimensional models, but critical and domain-specific dimensions are not well-developed. Professional development practices are often unstructured and not well planned. Assessment often relies on self-report tools, with limited focus on ethical and critical skills. Limited institutional support, unequal access to resources, unclear ethical guidelines, and shared or unclear responsibility can reduce positive individual outcomes, such as increased confidence and innovative teaching practices. The review concludes that the development of AI literacy tends to be built largely through individual effort and ad-hoc experimentation, while institutional governance structures that could make it durable and equitable remain underdeveloped. The available evidence suggests that sustainable AI literacy development requires structural governance rather than individual upskilling. Policy, practice, and future research implications are discussed, with the importance of longitudinal, linguistically diverse, and governance-oriented scholarship.
Student attrition in health professional education (HPE) carries serious implications for institutional performance and healthcare workforce supply, yet cross-institutional prediction is impeded by privacy regulations that prevent raw data sharing. This study presents, to the best of the authors’ knowledge, the first simulation-based evaluation of federated learning (FL) for privacy-preserving at-risk prediction in an HPE programme-discipline context, using the Open University Learning Analytics Dataset partitioned to simulate five programme-discipline clients as HPE analogues. A systematic search of Google Scholar, Scopus, and the ACM Digital Library (terms: “federated learning,” “at-risk student prediction,” “health professional education,” “privacy-preserving learning analytics”; May 2026) identified no prior study combining these elements. Grounded in Tinto’s Student Integration Model, LMS behavioural features were operationalised as integration proxies. FedAvg and FedProx were evaluated under logistic regression (LR) and MLP architectures across four scenarios: IID baseline, pedagogical heterogeneity, data quality variation, and temporal drift. Federated LR retained 99.3–99.9% of centralised performance across all scenarios (maximum gap: 0.64 pp AUC-PR), confirming a negligible privacy-utility trade-off. FedProx delivered no statistically significant advantage over FedAvg for LR in any scenario (AUC-PR differences of −0.03 to +0.19 pp, overlapping confidence intervals, Wilcoxon p > 0.05), but gained up to 6.40 pp AUC-PR for MLP under data quality heterogeneity, reducing training variance 40-fold — establishing proximal regularisation as a stability mechanism for non-linear architectures rather than a universal performance enhancement. Cross-semester evaluation showed no statistically significant temporal drift (Wilcoxon p = 0.32). SHAP attribution identified Active Days, Assessments Submitted, and Mean Assessment Score as globally dominant predictors; per-programme analysis revealed discipline-level differences in feature salience, informing targeted early intervention. Forum Interactions ranked last across all disciplines — diverging from Tinto’s social integration predictions and attributed to HPE’s reliance on clinical placements and in-person peer learning rather than asynchronous online forums.
Artificial intelligence is becoming increasingly integrated into students' academic lives in Ghana, making it crucial to examine their perspectives on ethical dimensions. Against this backdrop, this study assesses the artificial intelligence ethical awareness (AIEA) among 509 Ghanaian university students across the dimensions of human autonomy, beneficence, and fairness, based on AIEA scale originally developed and validated in Hong Kong. The confirmatory factor analysis supported the scale's three-factor structure within this setting. The network analysis also demonstrated meaningful interconnection among the dimensions, with select items serving as conceptual links, most prominently tying concerns over human agency to perceptions of beneficence. The latent profile analysis further distinguished five student subgroups, spanning those exhibiting uniformly very high awareness to a small but notable cluster displaying no ethical awareness, particularly regarding beneficence. In addition, MANOVA results confirmed that gender contributed no statistically meaningful variation. Together, the evidence characterises AI ethical awareness as a coherent yet unevenly distributed construct, highlighting priorities for shaping future AI ethics curricula in Ghana.
A persistent challenge in drone-based STEM education is the scarcity of teacher-verified, interactive simulations that effectively support both conceptual learning and competency development. While Generative AI offers promising avenues for accelerating content creation, its outputs often lack pedagogical validity and contextual relevance. This study examines whether drone instruction supported by teacher-AI co-designed simulations, yields superior learning outcomes compared to the same hands-on drone curriculum delivered without simulations. Employing a quasi-experimental pretest–posttest design with 30 secondary students (aged 13 - 17), both groups completed identical drone tasks under the same instructor, while the experimental group additionally engaged with five interactive simulations. Quantitative analyses revealed significantly greater improvements in the experimental group for both STEM content knowledge and overall learning competencies, including critical thinking, collaboration, communication, and creativity. Complementary qualitative data from interviews and observations revealed that simulations offered a low-stakes environment that reduced cognitive load, rendered causal flight mechanisms more visible, supported hypothesis testing, and facilitated transfer to physical drone operation. These findings suggest that simulations can serve as effective pre-flight, formative scaffolds that enhance learning visibility for teachers.The teacher–AI co-design model appears promising and potentially scalable for developing validated educational resources, although the results are constrained by the modest sample size and the quasi-experimental research design.
Generative AI teaching assistants are increasingly adopted in higher education, yet debates persist over whether student–AI conversations engage students in genuine cognitive processes or function primarily as information retrieval. We propose “Chat as Learning,” a measurement paradigm that treats student prompts as observable externalisations of in-progress learning cognition, complementing outcome-based assessment with a process-level signal aggregable at population, course, and student scales. We adopt the term learning to denote the paradigm’s theoretical positioning; the present study measures cognitive demand reflected in students’ prompts, classified via Bloom’s Taxonomy, rather than learning outcomes per se. Building on prior work establishing that approximately 62% of student–AI messages on this platform reflect higher-order cognitive demand at the aggregate level Chang and Li, 2026, we ask whether such engagement is stable across contexts or varies with disciplinary demands within the same individual. Drawing on data from the Uedu AI teaching assistant platform (https://uedu.tw), a multi-institutional deployment spanning four universities, we analysed over 60,000 student messages from 116 courses across two academic semesters using an automated Bloom’s Taxonomy classification pipeline. The classifier was validated against two trained human raters on a 300-message stratified sample, giving binary higher-order vs. lower-order LLM–human Cohen’s – at the individual level and under best-pair consensus (full 6 6 confusion matrix in Table 1). Our within-person, cross-discipline design—in which the same students were tracked across courses in different disciplines—revealed discipline-associated Bloom-level prompt profiles: STEM courses elicited Apply-prevalent prompts (20.8%), Language courses showed Understand-prevalent patterns (31.7%), and Social Science courses were Create-prevalent (33.8%); patterns in Humanities courses are reported descriptively but should be interpreted with caution given the small course base (3 courses across two semesters). Paired within-person comparisons indicated that the same students produced significantly higher proportions of higher-order prompts in Social Science courses than in STEM courses (pooled , ), with the same direction of effect observed in two independent semester samples. A crossed random-effects mixed-effects logistic regression confirmed these contrasts and showed that course-level variation in higher-order engagement substantially exceeds student-level variation. These results are consistent with the view that student–AI conversations reflect discipline-associated cognitive engagement rather than fixed individual interaction styles, and they suggest that AI teaching assistants may benefit from being designed and evaluated with disciplinary context in mind.
Educational AI systems increasingly rely on large language models to support student writing and inquiry, yet enforcing safety and instructional constraints during open-ended, multi-turn interaction remains challenging. Existing approaches commonly embed such constraints within conversational prompts or rely on static filtering. Over time, these approaches may become sensitive to user interaction, making it difficult to monitor and audit when students are able to circumvent or otherwise attempt to violate such measures. We introduce VETTING, a dual-LLM framework that separates response generation from policy verification and applies explicit policy checks at runtime. The framework is illustrated through a grounded instantiation that enforces instructional and safety constraints without exposing policy specifications during interaction. We evaluate VETTING through an in situ deployment in a middle school writing activity, combining analysis of student–AI interaction behavior, human audit of verification outcomes, and characterization of computational overhead. In this deployment, VETTING achieved a precision of .943, recall of .913, and F1 score of .928, and was associated with an estimated 91.2% reduction in inappropriate content exposure with a 19.6% increase in token usage (628,011 tokens). These results suggest that policy-isolated runtime verification can support the analysis and management of educational AI behavior under authentic classroom use.