
Background:Entering clinical training, dental students must learn to read tooth-centered electronic dental records, but limited teaching time in patient-centered clinics can leave gaps in chart-reading literacy. Multimodal large language models (LLMs) that process dental record images may offer scalable educational support, yet their performance has not been benchmarked. Objective:This study evaluated multimodal ChatGPT models on dental chart image interpretation and compared the best-performing model with dental students and residents. Methods:A retrospective, cross-sectional benchmark study used deidentified dental chart images from 15 patients at Seoul National University Dental Hospital (2017-2025). Charts containing Korean and English text were captured as sequential screenshots (154 images). For each patient, 24 Korean-language questions (360 total) spanned four categories: Type A, general factual retrieval; Type B, tooth- or procedure-specific retrieval; Type C, interpretation requiring multientry synthesis; and Type D, absent-information questions assessing abstention. Nine multimodal ChatGPT models (available August 2, 2025-August 9, 2025) were evaluated under standardized conditions. Outputs were scored against a gold standard using 7 metrics, with sentence bidirectional encoder representations from transformers (SBERT) similarity prespecified as the primary semantic measure. Human baselines included 2 third-year students and 2 first-year residents. Groups were compared with Kruskal-Wallis tests and Dunn post hoc analyses. Gold-standard reliability was assessed by independent senior-expert review and chance-corrected agreement (Gwet AC1) among clinical reference raters, and the model ranking was confirmed by content-based clinical accuracy analysis. Results:Across 360 items, GPT-5 Thinking achieved the highest median SBERT similarity (0.900, IQR 0.525-1.000), followed by GPT-5 Pro (0.861, IQR 0.501-1.000) and OpenAI o3 (0.831, IQR 0.489-1.000), with significant overall group differences (P<.001). Student 1 did not differ significantly from GPT-5 Thinking across Types A-D (all P≥.16), whereas Student 2 differed only on Type D (P=.002; Cliff δ=-0.27). Resident 1 scored higher than GPT-5 Thinking on Types A (P=.01) and C (P=.04) but lower on Type D (P<.001; δ=-0.35), whereas Resident 2 scored higher on Types A (P=.03), B (P=.02), and C (P=.04) and did not differ on Type D (P=.23). Type C tasks showed compressed SBERT distributions and low exact-match rates, indicating persistent difficulty in synthesis. The clinical reference rater agreement was high (Gwet AC1=0.99), and the content-based clinical-accuracy ranking was closely aligned with the SBERT ranking (Spearman ρ=0.90; P=.001). Conclusions:Statistically nonsignificant differences were observed between GPT-5 Thinking and dental students for most question-type contrasts, with Student 2 differing only on absent-information items. Compared with first-year residents, GPT-5 Thinking remained lower on several Type A-C contrasts, particularly interpretive Type C and one tooth- or procedure-specific Type B comparison; Type D contrasts require cautious interpretation because all groups had ceiling medians. Within this single-center benchmark, multimodal LLMs may have potential as supervised educational tools for chart-reading practice and verification, rather than as replacements for clinical expertise, pending external validation across institutions, specialties, and record systems.
Background:High-quality problem-based learning (PBL) during internship is resource-intensive and difficult to scale without consistent facilitation. Although generative AI is increasingly used in health professions education, many applications remain on-demand answer tools that may not reproduce core PBL processes. Objective:This study aimed to develop and evaluate the Multiagent PBL Environment for Clinical Reasoning (MAPLE-CR), an AI-supported environment that uses generative AI as a process-oriented scaffold rather than an answer-delivery aid. The system supports cognitive mechanisms through reasoning prompts, social-interactional mechanisms through simulated tutor and peer roles, and regulatory mechanisms through structured workflows and feedback loops. We separately evaluated implementation and repeated-use feasibility, including learner experience, short-term outcomes versus self-study, and preliminary framework-aligned process evidence. Methods:We conducted a 2-stage study. In an exploratory randomized evaluation (N=52; intervention: n=26; control: n=26), interns completed parallel pretests and posttests around a standardized case. The intervention group engaged in asynchronous, scaffolded PBL in MAPLE-CR, whereas controls completed case-matched self-study using materials derived from the same case and learning objectives. We summarized transcript-derived process dimensions and postsession learner experience in the intervention group. In a voluntary follow-up, 18 participants completed 1 MAPLE-CR case per week for 4 additional weeks to examine repeated use feasibility and patterns across novel cases. Results:Baseline pretest scores were comparable between groups (P=.56). The intervention group achieved higher posttest scores than controls (Hodges-Lehmann median difference 8.25 points, 95% CI 6.60-11.60; Cliff δ=0.506, 95% CI 0.21-0.76; P=.001) and greater score gains (median difference 6.60 points, 95% CI 1.60-14.95; P=.01). On a 0 to 5 scale, mean scores were highest for knowledge accuracy (4.4) and clinical reasoning (CR, 3.7), whereas active participation averaged 2.1, and interaction-oriented dimensions were more variable. Means for 12 positively worded questionnaire items ranged from 4.19 to 4.62, supporting high satisfaction, involvement, perceived support, and self-efficacy; open-ended responses contextualized perceived strengths and improvement needs. In the voluntary follow-up (n=18), weekly cross-case use was feasible; median within-session gains ranged from 13.4 to 20.0 points across 5 attempts, with marked increases in CR and aggregate scores from attempts 1 to 2 but statistically uncertain later trajectories. Conclusions:MAPLE-CR was feasible and showed high learner acceptability. Its use was associated with larger short-term CR knowledge gains than case-matched self-study, while transcript analyses provided preliminary framework-aligned process evidence. However, the post hoc sensitivity analysis did not establish prospective power sufficiency, and the confidence interval could not exclude smaller educationally meaningful effects. These findings suggest that AI-supported, process-oriented PBL may provide a scalable formative supplement for CR practice when facilitator capacity and small-group scheduling are constrained, but the study did not isolate specific scaffolding effects. Further research should confirm these findings using larger, prospectively powered trials with stronger active comparators and longer-term outcome measures.
Background:Large language model (LLM)-based AI teaching agents are increasingly used in medical education, yet their pedagogical quality is typically judged by platform-generated scores whose scoring criteria are undisclosed and may not reflect the teaching quality of the agent. Objective:This study aimed to develop and validate a multidimensional rubric for evaluating AI teaching agents and to examine the correspondence between platform scores and rubric-based teaching quality. Methods:Eight AI teaching agents covering an endocrinology curriculum were deployed across 4 role-play paradigms (patient, student, expert, and family). Twenty-two fourth-year medical students generated 167 dialogues, which were scored both by the platform and by an independently applied 8-dimension rubric (100 points, covering knowledge accuracy, pedagogical guidance, knowledge coverage, role-play quality, difficulty calibration, medical safety, student engagement, and feedback quality). Each dialogue was scored 4 times by a primary evaluator (Claude Opus 4.8; mean within-model SD 0.36), with 2 additional LLMs as robustness checks; 40 dialogues spanning all agents were rescored by a medical-education expert for validation. Results:Platform and rubric rankings diverged for most agents: the agent ranked third by the platform ranked last on rubric-based quality, and the platform's fourth-ranked agent ranked first. Agents differed most on knowledge-related dimensions (knowledge coverage coefficient of variation=27.3%) and least on role-play quality (coefficient of variation=5.7%), while difficulty calibration was a shared weakness. In a case-level observation, one agent revised specifically to strengthen empathy attained high role-play quality yet the lowest knowledge coverage of all agents. AI scores agreed with expert ratings at the total-score level (intraclass correlation coefficient=0.51) and on cognitive-process dimensions, but agreement was low for the more subjective dimensions. Student gender showed no detectable effect, though this analysis was underpowered. Conclusions:In this exploratory study, platform-generated scores reflected a construct different from agent teaching quality and should be used as a complement rather than as the sole quality indicator. The 8-dimension rubric provides a transparent, standardized alternative that reveals differences missed by platform scores, including a lack of association between empathy and knowledge coverage that warrants attention in future agent design.
Background:The Medical College Admission Test (MCAT) has been central to medical school admissions in North America, though its necessity in holistic selection processes remains debated. The COVID-19 pandemic's suspension of MCAT testing sessions allowed institutions to explore alternative admission criteria. Additionally, global challenges in test administration underscore the vulnerability of systems dependent on a single standardized test. Objective:This study aimed to (1) assess whether including MCAT scores in admissions decisions improves the prediction of medical school performance compared with grade point average (GPA) and interview-based selection and (2) evaluate whether machine learning (ML) can generate viable MCAT score predictions when testing is unavailable. Methods:We conducted a retrospective cohort study of 1898 applicants to the Lebanese American University School of Medicine (2009-2023). Among 349 admitted students with MCAT scores, we compared admission composites including vs excluding the MCAT in relation to academic outcomes (Med 1-4 final grades) and clinical performance (Med 1-2 objective structured clinical examination [OSCE] scores). Students were stratified into tertiles to examine tier-specific effects. To model an MCAT alternative, we trained an ensemble ML model (least absolute shrinkage and selection operator, kernel ridge, gradient boosting, elastic net, and light gradient boosting machine [LightGBM]) using data from 1583 applicants (2009-2020) to predict MCAT scores from cumulative GPA, core GPA, interview scores, merit points, and honors participation. The model was validated on 315 applicants admitted during the MCAT suspension (2021-2023), and its impact on admission rankings and performance correlations was evaluated. Results:Including the MCAT in the admissions composite did not meaningfully improve prediction of overall academic performance (r=0.63, 95% CI 0.56-0.69 with MCAT vs r=0.62, 95% CI 0.55-0.68 without MCAT; P=.63). However, it significantly weakened prediction of clinical skills (OSCE: r=0.37, 95% CI 0.27-0.46 with MCAT vs r=0.48, 95% CI 0.40-0.56 without MCAT; P<.001). In tertile analyses, the MCAT modestly improved prediction among top-performing students but eliminated predictive validity in the bottom tertile (r=0.12, 95% CI -0.07 to 0.30; P=.21 vs r=0.28, 95% CI 0.10-0.44 without MCAT; P=.003). The ML model explained 36% of MCAT variance (R²=0.36 with 95% CI 0.32-0.41; root mean squared error ≈10% of score range). Incorporating predicted MCAT scores reduced correlations with subsequent performance across all outcomes. Although rankings based on predicted MCAT scores strongly correlated with original rankings (r=0.80, 95% CI 0.77-0.83), 6.3% (4/64) to 9.4% (6/64) of admission decisions would have changed. Conclusions:Within this single-institution study, MCAT scores provided limited incremental validity for predicting academic performance and reduced prediction of clinical skills. ML-based predictions of MCAT scores introduced sufficient error to affect admissions decisions, supporting a robust holistic admissions process during temporary MCAT disruptions while highlighting the need for validation across diverse institutional settings before broader generalization.
Abstract This paper critically assesses the role of generative AI in anatomical illustration, identifying fundamental barriers that currently preclude AI from replacing human medical illustrators. Despite the promise of unprecedented efficiency, contemporary models exhibit persistent anatomical inaccuracies and “hallucinations” of nonexistent structures—flaws stemming from statistical pattern-matching rather than genuine anatomical understanding. These systems further lack pedagogical intent, clinical context, and the capacity for deliberate visual judgment, while raising unresolved ethical and copyright concerns regarding training data. Although a specialized AI for this purpose is theoretically feasible, its development as a standalone goal remains economically nonviable given the niche nature of the profession. Rather than replacing human illustrators, AI’s future role will be augmentative, with the requisite anatomical intelligence likely emerging as a byproduct of broader advances in clinical applications such as surgical planning and personalized medicine. For AI-generated imagery to become educationally and clinically reliable, it will require rigorous human supervision, curated gold standard datasets, and a foundation of genuine anatomical comprehension.
Background:AI is increasingly encountered in clinical care and medical education, but medical students' attitudes, perceptions, and self-reported familiarity have been assessed using heterogeneous survey instruments, AI referents, and response scales. Prior reviews often combined mixed health profession populations or summarized central estimates without fully showing variation across settings. Objective:This study aimed to synthesize quantitative evidence on medical students' AI-related attitudes, perceptions, and self-reported familiarity while examining construct harmonization, participant independence, heterogeneity, prediction intervals, risk of bias, and certainty of evidence. Methods:We searched PubMed (MEDLINE), Embase, Web of Science, Scopus, PsycINFO, and Cochrane CENTRAL from inception to April 1, 2026; supplementary searches are described in the appendices. Eligible studies enrolled students in MD, MBBS, MBChB, or DO-equivalent medical programs, or reported separable medical student data from mixed samples. Proportion outcomes were harmonized into 9 domains and synthesized using random-effects meta-analysis with Freeman-Tukey double-arcsine transformation, Hartung-Knapp-Sidik-Jonkman-adjusted CIs, and prediction intervals. Subgroup analyses and meta-regressions were exploratory because of multiple testing, ecological confounding, and construct heterogeneity. Risk of bias and certainty were assessed using the Joanna Briggs Institute analytical cross-sectional checklist and the GRADE (Grading of Recommendations, Assessment, Development and Evaluation) framework, respectively. Results:Ninety-six cross-sectional studies from 37 countries (>45,000 medical students) were included. Summary estimates suggested favorable attitudes but wide between-setting dispersion. Positive attitude toward AI was 76.9% (95% CI 72.2%-81.4%; prediction interval 42.2%-98.3%; I²=98.3%; 44 studies; N=20,806), perceived career benefit was 78.4% (95% CI 69.5%-86.2%; prediction interval 45.3%-98.3%; I²=98.0%; 16 studies; N=9799), and support for curricular integration was 76.6% (95% CI 71.8%-81.1%; prediction interval 47.8%-96.1%; I²=97.2%; 38 studies; N=16,308). Concern about physician replacement was 39.9% (95% CI 33.6%-46.5%; prediction interval 6.6%-80.1%; I²=98.8%; 32 studies; N=16,642), willingness to learn about or adopt AI was 71.5% (95% CI 64.8%-77.8%; prediction interval 37.9%-95.4%; I²=97.7%; 22 studies; N=9199), and ethical concerns were endorsed by 62.8% (95% CI 53.9%-71.3%; prediction interval 21.9%-95.0%; I²=98.8%; 28 studies; N=14,571). Self-reported familiarity or knowledge was 63.3% (95% CI 55.9%-70.3%; prediction interval 7.8%-100.0%; I²=99.5%; 52 studies; N=27,817), and trust in AI-assisted decisions was 50.6% (95% CI 28.5%-72.6%; prediction interval 7.3%-93.3%; I²=98.1%; 8 studies; N=3007). All domains had very low certainty because of cross-sectional self-report designs, frequent use of nonvalidated or adapted instruments, wide prediction intervals, and small study effects in several domains. Conclusions:Medical students' AI-related attitudes and curricular interest appear broadly favorable, but these estimates should not be interpreted as stable global prevalences. This review adds value by restricting the population to medical students, transparently harmonizing nonequivalent constructs, auditing mixed populations and participant independence, and reporting prediction intervals and certainty. Given very low certainty, the findings support locally adapted, exploratory AI-literacy planning and standardized measurement in future studies rather than strong claims about curriculum effectiveness.
Background:AI is transforming health care, creating an imperative to integrate AI into medical education. While student perspectives are well-studied, faculty views, particularly in non-Western contexts, remain underexplored. Objective:This qualitative study explores the perspectives of 10 medical faculty members from 2 institutions in the United Arab Emirates on integrating AI into undergraduate medical education. Methods:This multi-institutional qualitative study used purposive and reflexive sampling to recruit faculty from both public and private medical universities in the United Arab Emirates. Semistructured interviews were conducted with 10 faculty members involved in curriculum design, teaching, or assessment. Data collection followed COREQ (Consolidated Criteria for Reporting Qualitative Research) guidelines. Data were analyzed using a mixed inductive-deductive approach guided by the thematic analysis framework of Braun and Clarke. Findings were interpreted using the FACETS (Form, AI Use Case, Context, Education, Technology, and SAMR: Substitution, Augmentation, Modification, Redefinition) framework, which supported a structured examination of AI integration across different dimensions of teaching and learning. Results:Faculty primarily used generative AI tools, such as ChatGPT, for content creation, assessment development, and teaching support, reflecting a preference for accessible and general-purpose technologies. AI was mainly used to enhance teaching efficiency and support student learning, including personalized study planning and practice activities. Its application extended across preclinical and clinical contexts, with strong emphasis on adapting content to local cultural and ethical norms. While AI was perceived to improve efficiency and alignment between teaching and assessment, concerns were raised regarding equity, overreliance, and variability in student use. Overall, adoption remained focused on enhancing existing practices, with limited transformative use but recognition of future potential for more advanced applications. Conclusions:The UAE medical faculty demonstrate cautious optimism toward AI integration, recognizing its potential to enhance educational efficiency and personalization while emphasizing the critical importance of cultural contextualization. Current implementation remains at early adoption stages, focused on enhancement rather than transformation. Successful integration requires faculty development, context-sensitive policies, and equitable implementation strategies that address both technological and sociocultural dimensions of AI adoption in medical education.
Background:Virtual reality (VR) offers immersive learning opportunities in health professions education, yet its effects on learner experience and short-term learning outcomes in education on execution of musculoskeletal special tests remain underexplored. Objective:This randomized controlled trial evaluated the effects of an immersive 180° video-based VR instruction, compared with conventional 2D video-based instruction, on undergraduate physical therapy students' learning satisfaction, technology acceptance, learning motivation, and quiz-based learning achievement, reflecting short-term procedural knowledge. Methods:Undergraduate physical therapy students were randomly assigned to either the immersive 180° video-based VR instruction group (VR group) or the conventional 2D video-based instruction group (control group). Both groups received identical instructional content on 4 musculoskeletal special tests (pronator teres, Hawkins-Kennedy, Yergason, and Neer tests). The VR group participated in an immersive 180° video-based instructional module delivered via a head-mounted display, whereas the control group viewed the same content on a standard monitor. Participants completed 4 instructional sessions within a single day, each lasting approximately 3 to 4 minutes. The primary outcome was learning satisfaction. Secondary outcomes were technology acceptance, learning motivation, and quiz-based learning achievement. These outcomes were assessed before and after the intervention using validated questionnaires and a 12-item quiz. Between-group posttest differences were primarily analyzed using analysis of covariance (ANCOVA) with pretest scores as covariates. Supplementary within-group and unadjusted between-group analyses were also performed. Results:Fifty-two students were randomly assigned to receive either immersive 180° video-based VR instruction (VR group; n=28) or conventional 2D video-based instruction (control group; n=24) on 4 musculoskeletal special tests, with all 4 instructional sessions completed within a single day. After adjustment for baseline scores, the VR group showed significantly higher learning satisfaction (adjusted mean 107.04 vs 93.28; F1,49=13.710; P=.001; ηp2=0.219) and technology acceptance (adjusted mean 74.44 vs 68.03; F1,49=4.402; P=.04; ηp2=0.082) than the control group. Adjusted between-group differences were not significant for quiz-based learning achievement (F1,49=1.175; P=.28). For learning motivation, the homogeneity of regression slopes assumption was violated; therefore, the adjusted estimate was interpreted descriptively and did not indicate a clear between-group advantage (F1,48=1.729; P=.20). Conclusions:Compared with conventional 2D video-based instruction, immersive 180° video-based VR instruction showed a clear advantage in learning satisfaction, while technology acceptance also favored the VR group in exploratory secondary analysis, among undergraduate physical therapy students. Findings for learning motivation and quiz-based learning achievement were less conclusive. These results support immersive video as a supplementary educational approach rather than as evidence of superior psychomotor skills training.
Unlabelled:AI is no longer confined to optional decision support; it is becoming a routine presence in clinical workflows, shaping diagnostic hypotheses, triage priorities, risk estimates, and documentation. Yet most educational responses still treat AI as a tool operated by an individual clinician. This framing underestimates how AI reshapes the actual unit of practice: the interprofessional team. We propose the concept of AI-expanded interprofessional collaboration (AI-IPC), in which AI systems function as consequential participants in team cognition-not as moral agents, but as sources of recommendations, uncertainty, and constraints that reorganize communication, authority gradients, and accountability. Building on interprofessional education (IPE) theory, situated learning, and distributed cognition, we argue that "AI literacy" alone is insufficient; learners must be trained to coordinate human-AI-human collaboration in realistic clinical settings. We outline a pragmatic, theory-aligned approach for AI-expanded interprofessional education (AI-IPE): clarifying boundary conditions for AI participation, mapping AI-specific subcompetencies onto established IPE frameworks, and evaluating performance at the level of team behaviors rather than knowledge recall. We present the Interprofessional Education Collaborative (IPEC)-aligned evaluation scaffold with concrete learning activities and assessment approaches, and we outline how the approach can be tailored across undergraduate, postgraduate, and continuing education settings. We further emphasize that implementation requires digital infrastructure, institutional governance, and educator capacity that bridges AI, clinical workflow design, and IPE facilitation. Finally, we address the deliberate use of the "AI colleague" metaphor-not to anthropomorphize AI, but to make the relational and coordinative demands of AI integration visible and teachable. The pedagogical aim is calibrated, critical engagement: teams that verify, question, and, when warranted, override AI contributions rather than defer to them. AI will not replace interprofessional collaboration; it will change what collaboration requires. Education should make that change explicit, rehearsable, and assessable.
Unlabelled:AI is entering clinical practice faster than health professions curricula can teach it, leaving many educators eager to use AI-based teaching tools but unsure of how to build them. Generative AI chatbots-configured as simulated patients, clinical coaches, or formative assessment partners-offer scalable, interactive practice without any programming, yet most educators lack a structured method for designing and deploying them well. This tutorial provides that method: a practical, platform-agnostic workflow for building no-code AI chatbots using widely available large language model platforms. The workflow is organized in five sequential sections that follow the arc of a design project: (1) defining the educational purpose, learner group, and persona; (2) configuring behavior through the system prompt, graduated information disclosure, and structured feedback; (3) refining the learner experience through communication-style calibration, voice interaction, and curated knowledge documents; (4) adding realism and testing, including AI avatar generation and rigorous pilot-testing; and (5) embedding the tool within the curriculum and governing its use ethically. Each section pairs concrete, copy-ready design steps with the reasoning behind them, drawing on the technological pedagogical content knowledge framework and on established learning mechanisms-deliberate practice, self-regulated learning, formative feedback, and simulation-based learning-so that design choices are pedagogically grounded rather than merely technical. Throughout, 2 locally developed initiatives, the Virtual Integrated Patient and the Depression Avatars project, serve as illustrative implementation examples that motivated specific design decisions. These are presented as feasibility and acceptability experiences, not as evidence of educational effectiveness. The tutorial also addresses when a chatbot is not the right tool, the principal risks (hallucination, automation bias, data privacy exposure, and bias in generated personas), and a practical governance checklist for safe deployment. Although the examples are clinical, the workflow is discipline-agnostic and transferable across higher education. No-code AI chatbots are a feasible, accessible way for educators to build interactive learning tools; rigorous, multi-institutional evaluation using validated instruments remains the essential next step.
Background:Generative AI (GenAI) is increasingly integrated into clinical learning and practice. However, medical students often lack the competencies required for safe and critical use, including prompt design, output verification, and recognition of limitations. Educational interventions that integrate GenAI with clinical reasoning frameworks remain limited. Objective:This study evaluated a structured, theory-informed workshop integrating GenAI, prompt engineering, and clinical reasoning education to enhance medical students' self-perceived AI literacy and collaborative learning attitudes, and to assess whether patient-centered orientation changed following intensive AI exposure. Methods:We conducted a single-group, pre-post, explanatory sequential mixed methods study with fifth-year medical students enrolled at an academic medical center between April 2024 and May 2025. The 3-hour workshop comprised 6 modules integrating clinical reasoning, cognitive-bias awareness, verification-oriented GenAI use, and hands-on prompt engineering and centered on ChatGPT (OpenAI). Quantitative outcomes were self-reported measures from a 20-item self-report AI-literacy questionnaire adapted from the Meta AI Literacy Scale and were examined with exploratory and confirmatory factor analysis, the 6-item Patient-Practitioner Orientation Scale-Short, and a modified Collaborative Learning Attitude Scale (CLAS). Pre-post change was assessed using 2-tailed paired t tests with Benjamini-Hochberg correction and Cohen d; the Patient-Practitioner Orientation Scale-Short and CLAS were available for a subsample (n=46). Qualitative data from 6 interviews and 17 reflective narratives were analyzed using reflexive thematic analysis and integrated with the quantitative findings. Results:Among 150 eligible students, 139 (92.7%) completed paired AI-literacy assessments. Self-perceived AI literacy improved across all domains (Cohen d=0.69-0.93, all P<.001; all remaining significant after false discovery rate correction), and collaborative learning attitudes increased substantially (d=0.94, P<.001). Patient-centered orientation showed no significant change (d=0.03); however, baseline scores were concentrated at the favorable end of the scale (a floor/restricted-range effect), and the subsample analysis was underpowered (minimum detectable dz=0.42), so this null result is inconclusive rather than evidence of unchanged orientation. Gains did not differ by sex or across the sequential cohorts. Reflexive thematic analysis (interviews: n=6; reflections: n=17) identified five themes and one emergent theme describing a shift toward verification-oriented GenAI use: (1) understanding GenAI capabilities and limitations, (2) prompt-engineering skill development, (3) calibrated trust through verification, (4) GenAI-supported communication and collaboration, and (5) ethical considerations, with emerging reconceptualization of professional identity. Conclusions:A brief, theory-informed educational intervention integrating GenAI with clinical reasoning was associated with medium-to-large improvements in self-perceived AI literacy and collaborative attitudes. No detectable change in patient-centered orientation was observed; however, this finding should be interpreted as inconclusive, given measurement and power constraints. Embedding verification practices within clinical reasoning frameworks may offer a scalable approach for preparing physicians for responsible human-AI collaboration. Future studies should incorporate comparative designs, performance-based assessments, and longitudinal follow-up.
Background:Generative artificial intelligence (GenAI) is increasingly used to draft multiple-choice questions (MCQs) for health professions education, but much evidence concerns raw model outputs, expert ratings, or item difficulty alone. Educators edit GenAI drafts before use, and whether such items are psychometrically ready for postgraduate assessment remains unclear. Objective:This study aimed to compare human-edited GenAI-assisted and educator-crafted MCQs for postgraduate Family Medicine Applied Knowledge Test-level assessment, examining difficulty, discrimination, reliability, distractor functioning, and participant perceptions. Methods:We conducted a blinded cross-sectional, within-participant comparative psychometric evaluation in Singapore. Sixty best-of-five single-best-answer MCQs were evaluated, 30 human-edited GenAI-assisted items and 30 educator-crafted items, topic-matched across postgraduate FM domains and randomized across 2 assessment sets. Eligible participants were postgraduate doctors enrolled in FM residency or postgraduate family medicine programs, preparing for the Applied Knowledge Test, and blinded to item origin; incomplete paired responses were excluded. Outcomes included paired total scores, score correlation and agreement, Kuder-Richardson Formula 20 reliability, item difficulty index, corrected point-biserial discrimination, distractor functioning, and perceived difficulty, clarity, and relevance. Analyses used paired-sample tests, Pearson correlation, Fisher exact tests, and item-level psychometric statistics, with α=.05 and Bonferroni correction within comparison families. Results:Of 74 participants, 73 completed both item sets and were included in the analysis. The final sample comprised 36 graduate diploma in FM trainees, 5 MMed FM trainees, and 32 FM residents. Paired-sample testing showed lower scores on GenAI-assisted than educator-crafted items (mean 19.12, SD 2.83 vs mean 21.10, SD 3.42 out of 30; mean difference -1.97, 95% CI -2.72 to -1.23; P<.001; Cohen d=0.62), indicating that GenAI-assisted items were not easier. Scores were positively correlated (r=0.49, 95% CI 0.30-0.64; P<.001), but Bland-Altman analysis indicated limited agreement. Kuder-Richardson Formula 20 reliability was lower for GenAI-assisted items (0.38 vs 0.60). Mean difficulty index did not differ significantly (0.64 vs 0.70; mean difference -0.07, 95% CI -0.19 to 0.06; P=.29), and more GenAI-assisted items fell within the acceptable difficulty range (18/30, 60.0% vs 13/30, 43.3%). However, mean corrected point-biserial discrimination was lower for GenAI-assisted items (0.09 vs 0.18; mean difference -0.08, 95% CI -0.16 to -0.01; P=.04), and negative discrimination was more common (6/30, 20% vs 3/30, 10%). GenAI-assisted items also had more nonfunctioning and negatively discriminating distractors, although these differences were not statistically significant. Participant ratings of perceived difficulty, clarity, and practice relevance did not differ by origin. Conclusions:Human-edited GenAI-assisted MCQs can achieve plausible difficulty, but difficulty and surface acceptability did not ensure assessment readiness. Using trainee response data, this study extends work on raw outputs or expert opinion. GenAI should be used as a drafting adjunct within educator-led workflows prioritizing key verification, distractor engineering, pilot testing, empirical item analysis, and repair before item-bank or summative use.
Background:AI is increasingly discussed and deployed in health care, yet safe and effective implementation depends on the preparedness, trust, and training of the professionals who are expected to use these tools. Objective:This study aimed to assess current AI use, perceived benefits and concerns, confidence, and training needs among French health care professionals and students. Methods:We conducted a national web-based cross-sectional survey distributed through the PulseLife professional community between December 4, 2024, and March 5, 2025. The survey instrument was administered in French and included respondent characteristic items together with 12 substantive closed-ended questions covering current AI use, confidence, perceived benefits, and concerns, and interest in AI-related training. Access was restricted to authenticated individual PulseLife accounts, and multiple submissions from the same account were not allowed. Questions were not mandatory; incomplete questionnaires were retained for item-level analyses, and percentages were calculated using item-specific denominators. Because the exact invitation denominator was not retained by the platform, view, participation, and completion rates could not be calculated. Descriptive statistics and Pearson chi-square tests were performed using R. Internal consistency and exploratory psychometric properties were assessed using the Cronbach α, exploratory factor analysis, and confirmatory factor analysis. Results:A total of 1625 respondents participated, including 1212 (74.6%) health professionals and 413 (25.4%) students. Among professionals, physicians represented the largest group (642/1212, 53%), followed by nurses (232/1212, 19.1%) and pharmacists (92/1212, 7.6%). Only 6.6% (90/1366) of the respondents reported prior AI-specific training, whereas 78.3% (920/1175) wished to receive such training. Confidence in AI for diagnosis and patient management remained limited: only 9.2% (120/1301) of the respondents reported being very confident. Nearly half (673/1455, 46.3%) of the respondents who answered this item reported no current AI use in professional activity, whereas 10.5% (153/1455) reported frequent use. Physicians and younger respondents reported more frequent AI use, and prior AI training was associated with greater confidence (P<.001 in all cases). Commonly perceived benefits included improved diagnosis (774/1625, 47.6%), time savings (685/1625, 42.2%), reduced medical errors (634/1625, 39%), and improved patient follow-up (593/1625, 36.5%). Frequently reported concerns included algorithmic bias (785/1625, 48.3%), limited transparency (666/1625, 41%), deterioration of the patient-health care professional relationship (628/1625, 38.6%), and data confidentiality (557/1625, 34.3%). Conclusions:In this national French sample, formal AI training was uncommon despite high interest in receiving it. These findings support the need for more structured educational initiatives in AI literacy across undergraduate, postgraduate, and continuing professional education. Because this study relied on a convenience sample recruited through a digital platform, the findings should be interpreted as descriptive and exploratory rather than nationally representative.
Background:The focus on sustainability in university teaching and education about the climate catastrophe is constantly increasing and is essential for creating change. The chances offered by digital technological innovations are decisive in spotlighting planetary health education. Alongside welfare economies and social movements, they are among the areas with great potential to drive sustainable development. Objective:This study aimed to evaluate attitudes toward digital planetary health before and after a fully digital planetary health lecture series. This included assessing environmental health awareness, climate-related health risks, and perceptions of the need to integrate digital and planetary health education into university curricula. Methods:A repeated, pseudonymous, nonexperimental, cross-sectional online survey was conducted before (T1) and after (T2) a fully digital lecture series on digital planetary health during the 2022-2023 winter term. Because participant matching across time points was unsuccessful, descriptive analyses were conducted at the group level. Results:A total of 180 participants completed the T1 survey, and 79 participants completed the T2 survey. The mean age in the T1 survey was 27.92 (SD 10.95) years, with a minimum age of 18 years and a maximum age of 77 years. The mean age in T2 was 29.22 (SD 12.78) years, with a minimum age of 18 years and a maximum age of 78 years. In the T2 survey, 70.9% (56/79) of the participants rated it as "(very) important" (mean 5.38, SD 1.15) to incorporate climate-related health impacts and digital health issues into their studies. In the T2 survey, 51.9% (41/79) of the participants stated that the range of educational events on digital health at universities was inadequate. Qualitative responses highlighted perceived benefits of digitalization for collaboration, knowledge dissemination, and health care access, while concerns regarding resource consumption and ethical challenges were also described. Conclusions:Participants who attended the lecture series perceived digital planetary health as highly relevant to university education. The findings provide exploratory insights into attitudes toward the integration of planetary health and digital health topics into higher education curricula.
Background:Emergency nurses must be proficient in operating the Level-1 rapid infusion system to manage hypovolemic shock effectively. However, training opportunities for this infrequently used but life-critical device remain scarce, owing to resource constraints and limited access to equipment. Augmented reality (AR) has emerged as a promising educational technology that provides immersive, hands-on learning experiences without compromising patient safety; yet its application to specialized medical device training in nursing has not been rigorously evaluated. Objective:This study aimed to evaluate the effects of an AR-based training program using Microsoft HoloLens 2 on emergency nurses' clinical competency, self-efficacy, and educational satisfaction in operating the Level-1 rapid infusion system, compared with traditional guideline-based self-directed learning. Methods:A posttest-only randomized controlled trial was conducted at Samsung Medical Center in Seoul, Republic of Korea. Between July 17 and July 20, 2023, 42 registered nurses with no prior Level-1 experience were enrolled and randomly assigned in a 1:1 ratio to an experimental group receiving AR-based training on a single HoloLens 2 device (n=21) or a control group performing self-directed learning from a printed manual (n=21). Clinical competency was assessed by time (learning and performance), accuracy (a manufacturer-aligned checklist scored out of 100, and an expert-validated 22-step pass or fail evaluation), and the number of assistance requests. Self-efficacy (6-item scale; Cronbach α=0.80) and educational satisfaction (4-item scale; Cronbach α=0.87) were measured by questionnaire. Because most outcomes were non-normally distributed, groups were compared using the Mann-Whitney U test, with data reported as medians and IQRs. Results:Learning time was longer in the experimental group (median 18.20, IQR 15.48-21.67 vs 8.98, IQR 5.85-11.68 min; P<.001), but device setup time was markedly shorter (3.67, IQR 2.90-4.63 vs 9.85, IQR 8.03-11.35 min; P<.001). The experimental group achieved higher median device operation competency scores (90.00, IQR 80.00-100.00 vs 70.00, IQR 50.00-75.00 of 100; P<.001), passed more of the 22 evaluation steps (20.00, IQR 20.00-22.00 vs 16.00, IQR 13.00-18.00; P<.001), and required fewer assistance requests (0.00, IQR 0.00-1.00 vs 2.00, IQR 2.00-3.00; P<.001). Self-efficacy (20.00, IQR 18.00-25.00 vs 16.50, IQR 13.00-20.00; P=.003) and educational satisfaction (18.00, IQR 16.00-18.00 vs 12.00, IQR 9.75-15.00; P<.001) were also significantly higher in the experimental group. Effect sizes for the principal competency outcomes were large (Cohen d=1.2-3.1). Conclusions:AR-based training significantly improved emergency nurses' clinical competency, self-efficacy, and educational satisfaction in operating the Level-1 rapid infusion system compared with traditional self-directed learning. Despite requiring longer initial learning time, AR training produced faster device setup, greater accuracy, and enhanced learner independence. These findings suggest that AR technology can serve as an effective and scalable training solution for infrequently used but critically important medical devices in emergency care settings.
Early discussions of generative language models in medical education emphasized their promise for simulation, digital patients, individualized feedback, learner assessment, health information dissemination, research support, and translation, while also warning about bias, privacy, academic integrity, misinformation, legal ambiguity, and unequal access. Since that first wave, generative artificial intelligence has moved from novelty to routine exposure for learners, educators, researchers, and institutions. Medical education therefore needs a more mature framework than a catalog of opportunities and risks. This Viewpoint argues that the next phase should be organized around educational entrustment: determining which functions can be delegated to AI systems, under what conditions, with what human supervision, and with what evidence of benefit. Building on recent proposals to apply entrustment to AI in health professions education, we operationalize the concept into a graduated, function-level model that specifies which educational functions may be delegated, at what stakes, with what oversight, assessment, and governance. We classify use cases by educational stakes and AI autonomy, and outline implications for assessment redesign, curriculum development, faculty capability, cognitive autonomy, equity, and institutional governance. The central challenge is whether medical schools can integrate these tools in ways that preserve clinical reasoning, professional identity, accountability, and fairness. The next generation of research should move beyond model performance on examinations and evaluate how AI changes learning, judgment, behavior, and patient care.
Background Generative AI tools became widely available to the public in November 2022. The extent to which these tools have been used by medical school applicants during the admissions process is unknown. Objective We aimed to estimate the extent of generative AI use among cohorts of applicants spanning the rollout of these tools. Methods We retrospectively analyzed 6000 essays from 2364 applicants submitted to a US medical school in 2021 to 2022 (baseline, before the wide availability of AI) and 2023 to 2024 (test year) to estimate the prevalence of AI use and its relation to other application data. We used GPTZero, a commercially available detection tool, to generate a metric (Phuman) reflecting the predicted probability that each essay was completely human generated, ranging from 0 (the essay appears to be entirely AI generated) to 1 (the essay appears to be entirely human generated). Results Fully human-generated negative controls demonstrated a median Phuman of 0.93 (range 0.89-0.97), while fully AI-generated positive controls demonstrated a median Phuman of 0.01 (range 0.00-0.01). The “Personal Comments” essays submitted in the 2023 to 2024 application cycle had a median Phuman of 0.77 (95% CI 0.76-0.78) compared with 0.83 (95% CI 0.82-0.85) during the 2021 to 2022 cycle. Approximately 12.3% and 2.7% of essays were evaluated as having Phuman <0.5 in the test and baseline years, respectively. Essays submitted as part of the secondary application demonstrated lower Phuman values than those of the American Medical College Application Service (AMCAS) “Personal Comments” essays. In applicant-clustered, multivariable generalized estimating equation analyses, supplementary essay type and younger age were significantly associated with lower Phuman. Application completion date, self-reported gender, program type (MD vs MD-PhD), grade point average (GPA), Medical College Admission Test (MCAT) score, socioeconomic status, and undergraduate major were not significant predictors after false discovery rate correction. Phuman was not predictive of interview invitation or acceptance in adjusted applicant-level logistic regression analyses. Conclusions An AI detection algorithm identified signs of increased use of generative AI in 2023 to 2024 medical school admission applications compared to those in the 2021 to 2022 baseline period, before AI was widely available. AI use did not appear to confer an admissions advantage. Although these results provide information about the applicant pool as a whole, AI detection is imperfect. We do not recommend deploying AI detection for individual applications in live admissions cycles.
Background:Artificial intelligence (AI) is rapidly integrating into health professions education and clinical practice, creating significant opportunities alongside new ethical challenges. Although current international and professional guidance establishes essential values, it offers limited direction for how clinicians, educators, learners, and institutions should act in routine educational, research, and clinical contexts. The CARE-AI (Contextual, Accountable, Responsible, and Equitable Artificial Intelligence) project responds to this practice-level gap by articulating guidance that moves beyond values toward professional accountability and equity, with explicit attention to educational, research, and clinical practice contexts. Objective:The study objective was to develop and validate a consensus-based, actionable framework of principles to guide responsible AI use across health professions education, research, and clinical care. Methods:We conducted a 3-phase modified Delphi consensus study, reported in accordance with the Accurate Consensus Reporting Document. Phase 1 involved 2 international professional meetings and 3 purposively sampled focus groups (AI or technology, health professions education, and ethics or professionalism) to adapt and refine draft principles using an exploratory qualitative approach. Phase 2 used an online survey with a 5-point importance scale and prespecified consensus criteria (inclusion ≥70%: high ratings; exclusion ≥70%: low ratings). Phase 3 used include, exclude, or undecided voting on revised principles. Quantitative thresholds determined consensus. Qualitative free-text comments informed iterative refinement. Results:Participants represented diverse communities of practice across health professions education, clinical care, patient partners, ethics, and digital health, spanning multiple professional roles and training levels. Across all phases, 303 unique participants contributed to the study. Phase 1 focus groups (n=61) provided early insight and direction. In phase 2, the first Delphi survey round, 242 participants initiated the survey, with 120 (49.6%) participants completing it. In phase 3, the second Delphi survey round, 103 participants were invited based on expressed interest at the end of the first round; 78 participants initiated the survey and 75 completed it (75/78, 96.2% of starters). In phase 2, of the 61 statements, 58 (95%) met the inclusion criteria, and participants submitted 1887 comments (697 were content rich), prompting clearer accountability language, stronger equity commitments, and more usable wording. In phase 3, all 10 principles and their statements met the inclusion criteria. Participants contributed 224 comments (179 were content rich) that informed final refinements. Endorsement was near unanimous: 96% (72/75) agreed or strongly agreed that the framework clearly defined professionalism expectations for AI to meet educational, technological, and ethical needs in the health professions. Conclusions:The Health CARE-AI Framework, with its preamble and 10 principles, articulates actionable, consensus-validated guidance that moves from values to competence, into professional accountability, and toward structural commitments to equity. Paired with a companion implementation guide and toolkit, the framework is intended to support use across education, research, and clinical settings.
Abstract Background Canadian psychiatry residents must demonstrate consultation competency, assessed using the standardized assessment of a clinical encounter report (STACER). However, opportunities to practice these skills and receive constructive assessment remain limited in clinical settings. Objective This study aimed to evaluate the technical feasibility of an agentic AI system designed to support psychiatry residents’ consultation competence through simulated patient encounters with a patient agent and structured feedback from a rater agent. Methods We conducted a two-phase technical feasibility prospective single-arm cohort study of the STACER Agentic System, a large language model–based platform integrating a patient agent and a rater agent. Phase 1 involved automated evaluation of the patient agent using a psychiatrist agent across 227 synthetic major depressive disorder cases. Performance was assessed using DeepEval metrics (correctness, clarity, medical faithfulness, turn relevance, and role adherence) with descriptive statistics and 95% CIs. Phase 2 involved a preliminary user study with 14 convenience-sampled participants: a total of 5 members of the clinical research team and 9 psychiatry residents from the University of Alberta. Participants completed simulated diagnostic interviews and case presentations. Performance was evaluated using STACER-based scoring by the rater agent and 2 psychiatrists. Interrater reliability was assessed using intraclass correlation coefficients (α=.05). Participants rated realism, behavioral consistency, psychiatric nuance, and feedback utility using Likert scales and free-text answers. Results The patient agent demonstrated high behavioral (51/56, 91.07%) and symptom fidelity (105/110, 95.45%), with strong automated performance (medical faithfulness mean 0.99, 95% CI 0.99‐1.00; turn relevance 0.99, 95% CI 0.986‐0.992). Participants rated simulations as psychiatrically plausible and diagnostically useful, particularly for depressive symptom representation, although rapport building was moderate (mean 2.78, SD 1.56 to mean 3.00, SD 1.41, out of 5.00) due to limited nonverbal cues. The rater agent generated structured STACER-aligned feedback with high intrarater consistency, especially at the section subtotal level. Interrater reliability with psychiatrists was poor at the item level (intraclass correlation coefficient range=0.25‐0.49) but improved to good-to-excellent agreement at the section level for psychiatry resident sessions (intraclass correlation coefficient range=0.89‐0.93). The rater agent’s scores fell between those of the 2 psychiatrists for the clinical research team and were lower than both human raters for psychiatry residents. Conclusions The STACER Agentic System demonstrates the technical feasibility of using agentic AI to simulate psychiatric consultations and deliver STACER-aligned formative feedback. By combining adaptive multiturn psychiatric simulation with competency-based evaluation, it shows promise in supporting cognitive aspects of consultation, though it remains limited in facilitating relational skills such as rapport building. These findings suggest agentic AI could expand scalable, low-risk opportunities for deliberate practice and formative feedback in competency-based psychiatric education. Further controlled studies are needed to evaluate educational effectiveness and integration into residency training.