
PURPOSE:Conventional single-run accuracy may be insufficient for evaluating frontier generative artificial intelligence (AI) models when safety-related refusals occur. This 4-model benchmark examined the need for repeated, refusal-aware evaluation using the Japanese National License Examination for Pharmacists (JNLEP). METHODS:ChatGPT GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Claude Fable 5 were evaluated using all 345 questions from the 107th JNLEP. The original Japanese questions, including image-containing items, were submitted through application programming interfaces (APIs) in 3 independent runs. Refusals were treated as incorrect when overall accuracy was calculated. For Fable 5, accuracy excluding refusals, refusal rate, refusal consistency across runs, the subject-wise distribution of refusals, and system-assigned refusal categories were also evaluated. RESULTS:The mean overall accuracies were 98.7% for GPT-5.5, 98.3% for Gemini 3.5 Flash, 96.1% for Claude Opus 4.8, and 70.3% for Claude Fable 5. Fable 5 had a mean refusal rate of 29.0%, whereas its mean accuracy excluding refusals was 99.0%. Among the 345 items, 93 were refused in all 3 runs, 14 were refused inconsistently across runs, and 238 were never refused. All 300 refusal responses were assigned to the bio category. Refusals were most frequent in Biology (90.0%) and Pharmacology (69.2%) but uncommon in Practice (2.5%). CONCLUSION:Near-saturation benchmark performance coexisted with frequent and partly run-dependent refusals in a safeguard-equipped model. Overall accuracy, accuracy excluding refusals, refusal rate, and refusal consistency describe complementary aspects of performance; therefore, repeated, refusal-aware evaluation is needed to interpret frontier AI models in pharmacy education.
PURPOSE:This study aimed to develop an instrument to assess the perceived importance of core competencies for clinical nurse educators (CNEs) in Korea, defined here as hospital-based nurse educators under the Korean Nursing Act, and to provide initial evidence of its content validity, internal structure, and reliability. METHODS:A preliminary 44-item pool, developed from international nurse educator frameworks and the literature, was refined by 8 experts through a 2-round modified Delphi process. A nationwide sample of 263 CNEs then rated the perceived importance of each item. Dimensionality was examined using parallel analysis and exploratory factor analysis based on polychoric correlations; the refined model was tested using confirmatory factor analysis (CFA) with the weighted least squares mean- and variance-adjusted estimator, composite reliability (CR), average variance extracted (AVE), Fornell-Larcker and heterotrait-monotrait discriminant validity, and alternative models. RESULTS:Content validity was high (scale-level content validity index=0.986). Parallel analysis supported a 5-factor structure; 9 weak or redundant items were removed, consolidating the 8 domains into 5 factors comprising 35 items. The CFA model showed good fit (comparative fit index=0.977, Tucker-Lewis index=0.975, root mean square error of approximation=0.048, standardized root mean square residual=0.053; standardized loadings, 0.669-0.934). Internal consistency was high (Cronbach's α=0.873-0.913; categorical McDonald's ω=0.871-0.916), and convergent validity was supported (CR=0.91-0.95; AVE>0.50). Three factors were highly correlated; a second-order model fit well, and the general factor explained most of the reliable variance (ω_h=0.83). CONCLUSION:The 35-item instrument provided initial evidence of content validity, internal structure, and reliability for assessing the perceived importance of CNE competencies. These findings represent initial, rather than definitive, validity evidence and require cross-validation in an independent sample.
PURPOSE:To evaluate students' acceptance of training via artificial intelligence (AI)-based virtual patients (VPs) in pediatrics; to analyze the main factors that affect student satisfaction; and to compare student satisfaction with that reported in the literature. METHODS:Two observational studies were carried out. Study S1 analyzed students' interactions with the platform and study S2 analyzed their answers to a final questionnaire. All students enrolled in the "Pediatrics II and Pediatric Surgery" course were invited to participate during a continuous session on May 7, 2025. Study S1 analyzed whether the case/session order, the number of interactions, the time spent, and student gender were associated with student satisfaction (by fitting a linear mixed-effects model); and also performed a qualitative analysis of students' open-ended comments. Study S2 analyzed the influence of demographic data on students' feedback (by fitting proportional odds ordinal logistic regression models and linear models). All experimental data, code and results are available. RESULTS:Concerning study S1, 70 students participated. Platform rating was high (mean=9.04, standard deviation=1.09) and was positively associated with the number of interactions with the VPs (P=0.038, β=0.022 [0.001-0.043]). Concerning study S2, 60 students participated. No statistically significant influence of demographics was found either for answers to individual questionnaire items or for grouped answers. CONCLUSION:Student acceptance of training with AI-based VPs was supported by the satisfaction outcomes measured and by comparison with previous studies. However, further research is needed to confirm these findings.
PURPOSE:When instructors address students by name, this may be associated with students' engagement and sense of inclusion in higher education. However, name-based instructor-student interaction has rarely been conceptualized as a multidimensional construct, and no psychometrically evaluated instrument is available for medical education. This study developed and psychometrically evaluated a pilot questionnaire assessing students' perceived experiences of being addressed by name in medical teaching contexts. METHODS:Questionnaire development followed a multistep exploratory approach that included qualitative input, iterative item refinement, and expert content review. An initial pool of 59 items was administered to medical students (n=270) from semesters 2-10 at a German medical faculty. Factor structure was examined using exploratory factor analysis based on Pearson correlations, oblique rotation, and predefined item-reduction criteria. Internal consistency was assessed using McDonald's ω and Cronbach's α, discriminant validity using the heterotrait-monotrait ratio, and generalized partial credit models were estimated separately for each factor. RESULTS:Exploratory factor analysis yielded a parsimonious 3-factor solution comprising 15 items: cognitive activation, appreciation & social belonging, and social evaluative anxiety. The final model showed good fit, and all factors demonstrated strong internal consistency (ω=0.84-0.88), good measurement precision (expected a posteriori reliability=0.83-0.88), adequate discriminant validity (heterotrait-monotrait ratio <0.85), and acceptable item fit. CONCLUSION:The pilot questionnaire captures 3 distinct dimensions of name-based instructor-student interaction and provides promising initial psychometric evidence for medical education research. It may support confirmatory validation and further study of relational teaching practices.
PURPOSE:This review examined the impact of learner role-active participant versus observer-on learning outcomes in healthcare simulation-based education using Kirkpatrick's evaluation model. METHODS:Five databases-PubMed, Scopus, Web of Science, ScienceDirect, and EBSCOhost-were searched for studies published from January 2018 to November 2024. Eligibility criteria, defined using the PICOS framework, targeted studies comparing learning outcomes between active participant and observer roles among health professions students and practitioners. Methodological quality was appraised using the Medical Education Research Study Quality Instrument, and findings were synthesized narratively. This review was registered in PROSPERO under registration number CRD42024611988. RESULTS:Of 648 records, 15 studies involving 1,799 participants met the inclusion criteria; role allocation was specifically reported for 644 active participants and 709 observers. Most studies were conducted in high-income countries (13/15) and nursing programs (11/15). At Kirkpatrick Level 1, most studies reported no statistically significant between-role differences, although an active participant advantage emerged for emotional arousal. At Level 2, knowledge, skills, and most attitudinal outcomes did not differ significantly between roles; active participant advantages were observed for technical skills at delayed follow-up, affective competencies in palliative care, and perceived learning transfer. No study assessed Levels 3 or 4. CONCLUSION:Most studies found no significant differences between observer and active participant roles at Kirkpatrick Levels 1 and 2. Generalizability across all health professions and settings is limited. Pedagogically appropriate use of the observer role supports learning outcomes and may expand access to simulation-based education.
PURPOSE:This study aimed to evaluate a structured basic radiology course for Family Medicine residents delivered in a 3-dimensional (3D) virtual environment (Second Life). We hypothesized that participation would be associated with higher post-course knowledge test scores and high learner satisfaction. METHODS:A quasi-experimental multi-cohort pretest-posttest study was conducted across 3 consecutive cohorts of Family Medicine residents in Spain in 2019. Ninety-six participants engaged in a 15-day course combining synchronous and asynchronous activities. Sixty-five participants provided paired pre- and post-intervention knowledge assessments, which were analyzed using paired t-tests. Learner satisfaction was evaluated using a structured questionnaire with Likert-scale items, numerical ratings (0-10), and open-ended responses. Pre- and post-test scores and ratings were compared using paired t-tests, and Mann-Whitney U tests were used for Likert-scale data. Qualitative data from open-ended responses were analyzed using thematic coding. RESULTS:Participants had significantly higher post-course knowledge test scores, increasing from 47.4±11.6 to 58.2±11.9 (mean difference, 10.8; P<0.001; paired-sample Cohen's d [dz]=0.67). High agreement was observed across most satisfaction items, with median scores ranging from 4 to 5. Content relevance, usefulness for clinical practice, and instructor performance received the highest ratings. Lower scores were observed for peer interaction and platform usability. CONCLUSION:A structured radiology course delivered in a 3D virtual environment was associated with higher immediate post-course knowledge test scores and high learner satisfaction among Family Medicine residents. The findings support the feasibility of repeated course delivery, although controlled studies are warranted.
PURPOSE:This study aimed to develop and validate the Thai AI Literacy Scale for Nursing Students (TAILS-NS). METHODS:A cross-sectional study was conducted across multiple nursing institutions in Thailand from March to May 2026 to address the absence of a validated artificial intelligence (AI) literacy instrument for nursing students in Thai or Southeast Asian contexts. A total of 410 nursing students participated, yielding a response rate of 94.5%. The TAILS-NS was developed through item generation, expert content-validity assessment using the item-objective congruence index, and a 2-phase pilot study, resulting in a 40-item instrument comprising a 15-item knowledge test and 25 Likert-scale items across 6 domains. Exploratory factor analysis using maximum likelihood estimation and confirmatory factor analysis using the weighted least squares mean and variance-adjusted estimator, as well as internal consistency, convergent validity, and discriminant validity, were assessed. An independent validation sample (n=157) was recruited for cross-validation confirmatory factor analysis. Raw data are available as a supplement. RESULTS:Exploratory factor analysis supported a 6-factor structure: AI awareness, skills, ethics and professionalism, positive attitude, AI anxiety, and readiness. Confirmatory factor analysis showed acceptable model fit, although the root mean square error of approximation (RMSEA) indicated marginal fit (comparative fit index [CFI]=0.919, Tucker-Lewis index [TLI]=0.906, RMSEA=0.090). Reliability of the knowledge subscale was acceptable (Kuder-Richardson Formula 20=0.758). Internal consistency was excellent (α=0.810-0.934; total α=0.974; ω=0.989). Average variance extracted exceeded 0.50 for all factors, supporting convergent validity. Most heterotrait-monotrait ratios were below 0.90, supporting discriminant validity. Cross-validation confirmed factorial replicability (CFI=0.917, TLI=0.904). CONCLUSION:The TAILS-NS provides evidence of validity and reliability for measuring AI literacy among Thai nursing students and may serve as a standardized tool for curriculum evaluation. Future studies should expand the AI anxiety subscale and explore cross-cultural applicability in Southeast Asian nursing contexts.
PURPOSE:This study characterized the psychometric properties of Korean Medical Licensing Examination items administered from 2012 to 2022 using the Rasch and 2-parameter logistic models and descriptively examined items classified as highly difficult. METHODS:Item parameters were estimated separately for each examination year using the Rasch and 2-parameter logistic models implemented in the irtQ package in R. Descriptive statistics and correlations between model-based difficulty estimates were calculated. Items with a 2-parameter logistic difficulty estimate of b ≥2.0 were subjected to content analysis. RESULTS:Correlations between Rasch and 2-parameter logistic item-difficulty estimates ranged from 0.70 to 0.80 across examination years. Under the 2-parameter logistic model, the proportion of items with b ≥2.0 ranged from 4.7% to 8.9% and showed no monotonic temporal trend; the proportions were 5.0% in 2021 and 5.6% in 2022. The Rasch model generally produced smaller estimated standard errors than the 2-parameter logistic model. CONCLUSION:Rasch and 2-parameter logistic difficulty estimates showed strongly correlated rankings, although correlation alone did not establish agreement between their absolute estimates. The Rasch model demonstrated greater numerical stability under the present calibration conditions, but model selection should also consider model fit, test information, and the intended use of scores. Because examination years were calibrated separately and no pre-disclosure or control condition was available, the observed annual differences cannot be interpreted as evidence that item disclosure caused changes in item difficulty or discrimination.
PURPOSE:This study aimed to evaluate the validity of multimedia-assisted items (MAIs) in a mock examination by examining their psychometric properties and nursing students' perceptions. METHODS:This methodological study used a within-group counterbalanced design in which fourth-year nursing students in South Korea completed an online mock examination consisting of paired text-based items and MAIs. Psychometric properties were examined using classical test theory and item response theory, and learner perceptions and response processes were explored through post-examination surveys and focus group interviews. RESULTS:Among 515 participants, MAIs demonstrated overall difficulty and discrimination comparable to those of text-based items in both classical test theory and item response theory analyses, with similar test characteristic curves across counterbalanced groups, supporting measurement equivalence across item formats. Post-test surveys indicated that participants perceived MAIs as realistic reflections of clinical practice and as more helpful for problem solving and ability assessment than text-based items. Qualitative findings from focus group interviews (n=15) further supported the educational value of MAIs for assessing clinical judgment while emphasizing the need for appropriate multimedia use. An order effect was also observed: presenting MAIs first led to higher item response theory discrimination parameters in subsequent text-based items (P=0.020). CONCLUSION:MAIs demonstrated psychometric properties comparable to those of text-based items and were positively perceived as valid tools for assessing clinical judgment and practical competence. These findings support the feasibility of incorporating multimedia-based assessment into the Korean Nursing Licensing Examination as part of a rigorous computer-based testing framework.
Purpose: This study aimed to describe inclusive mechanisms supporting dental students with hearing disabilities, focusing on facilitators of and barriers to participation in clinical education.Methods: A qualitative case study was conducted at a Chilean university. Data were collected through semi-structured interviews with 2 students with hearing disabilities, 3 clinical instructors, and 1 Chilean Sign Language interpreter who works at the university and supports dental students with hearing disabilities, as well as through direct observation. Data were analyzed using inductive thematic analysis. Participant validation was used to enhance credibility.Results: Facilitators were identified across pedagogical, technological, communicative, and relational domains. Pedagogical adaptations included early access to materials, visual resources, and flexible assessments. Technological tools such as real-time transcription supported communication and access to information. Alternative communication strategies included lip reading, written support, visual cues, and Chilean Sign Language. The presence of an interpreter was a key enabling factor. Relational support from faculty, peers, and family members also contributed substantially. Faculty flexibility and willingness to adapt were particularly important. Barriers were identified at structural, institutional, communicative, and attitudinal levels. Clinical environments presented physical and acoustic challenges, including noise, limited space, and complex interactions. Institutional responses were often reactive and involved limited planning. Communication barriers affected interactions with patients and peers, as well as academic literacy. Faculty reported limited training in inclusive education. Attitudinal barriers included misconceptions about disability and limited experience with disability.Conclusion: Inclusive mechanisms are present but insufficiently systematized. Strengthening institutional planning and structural adaptations is essential to ensuring equitable clinical education.
Purpose: This study evaluated the implementation of a patient- and family-centered rounds (PFCR) educational intervention and a standardized assessment form developed to assess changes in medical students’ PFCR performance over time.Methods: From October 2023 to October 2024, medical students at Johns Hopkins School of Medicine attended a 1-hour PFCR simulation workshop during the Core Clerkship in Pediatrics. Students rotating at the main campus were assessed during rounds with a formative standardized form; all students received summative oral presentation and family rapport scores regardless of site. Performance was compared between students rotating in the first half (H1) and second half (H2) of the clerkship using Wilcoxon rank-sum tests with Holm correction for multiple comparisons. Linear mixed-effects models with student-level random intercepts were used to estimate changes across serial assessments.Results: Among the 74 students rotating at the main campus, assessment forms were completed for 61% of students, with a median of 3 completed forms per student. H1 and H2 students had similar scores on both the formative assessment form and summative evaluations. Both groups improved significantly in inviting patient and family concerns across serial assessments. Summative scores did not differ between students evaluated before and after the educational intervention or between students at the main campus and those at community hospitals.Conclusion: A structured PFCR educational session paired with a standardized assessment form was feasible to implement in a pediatric clerkship. Observed student-level median scores were at or above “meets expectations” across 6 PFCR domains, and no significant performance decline was observed among students rotating later in the clerkship. Further controlled studies are needed to determine whether this intervention improves PFCR skills compared with usual training.
Purpose: This study aimed to analyze recent trends in occupational therapy, derive updated job competencies for Korean occupational therapists that reflect recent clinical changes, and confirm their validity.Methods: This 5-month expert consensus study, conducted from April to August 2024, used a 3-round online Delphi survey with 20 occupational therapy experts, followed by a focus group meeting with 10 experts, to refine and validate updated job competencies. Considering recent trends in occupational therapy, we adapted the U.S. National Board for Certification in Occupational Therapy practice analysis framework to the Korean context and collected expert feedback.Results: The Delphi panel identified several competency items that required contextualization within the Korean legal scope of practice, particularly items related to physical agent modalities. After the concluding focus group discussion, the superficial thermal agents item (D3-T1-K4) was excluded, and the electrotherapeutic modality item for swallowing disorders (D3-T1-K5) was revised; the remaining competencies were finalized.Conclusion: Updated job competencies for new occupational therapists were derived across 4 domains, 16 tasks, and 62 knowledge items. These competencies may support greater flexibility in international research and strengthen responses to the expansion of occupational therapists’ scope of practice in community settings. Therefore, the findings of this study are expected to be actively used in research, education, and practice.
PURPOSE:This study explored non-surgical residents' and their domestic partners' perceptions of how residency affected their lives and relationships through semi-structured interviews. METHODS:This qualitative study consisted of 10 semi-structured interviews with non-surgical residents at a single institution and 3 interviews with domestic partners. Residents from Child Neurology, Family Medicine, Internal Medicine, and Psychiatry were represented. The interviews were recorded, transcribed, and collaboratively coded using Atlas.ti ver. 22.0. Themes were identified using team-based thematic analysis. RESULTS:Analysis of the interviews yielded 4 themes: (1) residency training results in noticeable changes in personality traits; (2) residency affects life beyond the workplace; (3) residents and their domestic partners employ strategies to support relationship success during residency; and (4) residency program interventions may support well-being. CONCLUSION:This study found that residency training was perceived to affect residents, domestic partners, and their relationships in both beneficial and challenging ways. Residents' domestic partners remain a valuable yet underutilized resource for learning more about the residency experience.
PURPOSE:This study aimed to evaluate the implementation and evolution of a peer-assisted, simulation-based clinical skills program at Miguel Hernández University, Spain, focusing on medical student participation, perceptions, and perceived educational value. METHODS:A prospective quasi-experimental pre-post study without a control group was conducted. Sessions were organized in small groups and led by senior student tutors under faculty supervision. Six workshops were offered during the first academic year and 8 during the second, with the addition of lumbar puncture and introductory clinical ultrasound. Knowledge acquisition was assessed using pre- and post-workshop questionnaires, and satisfaction was evaluated with Likert-scale surveys. RESULTS:A total of 154 students participated, including 77 in 2023-2024 and 77 in 2024-2025, generating 440 workshop attendances. After incomplete questionnaires were excluded, 425 paired pre- and post-workshop evaluations were analyzed. Students reported improved learning, with self-reported knowledge scores increasing significantly from 6.1±2.6 to 8.7±1.6 (Δ=2.5; 95% confidence interval [CI], 2.29-2.71; Cohen's d=1.14; P<0.001). Male students showed greater knowledge gain than female students (Δ=2.8 vs. 2.3; 95% CI, 2.17-3.43 vs. 1.91-2.69; Cohen's d=1.17 vs. 1.15; P=0.034), and second-year students improved more than third-year students (Δ=2.9 vs. 2.1; 95% CI, 2.36-3.44 vs. 1.67-2.53; Cohen's d=1.21 vs. 1.11; P=0.001). Satisfaction was high, with mean scores above 4/5. CONCLUSION:Clinical simulation combined with peer tutoring was feasible and well accepted in an undergraduate medical curriculum in Spain, achieving sustained participation over 2 academic years and consistently high satisfaction ratings. The program was associated with significant immediate improvements in workshop-specific knowledge test scores.
Purpose: The clinical learning environment (CLE) is a crucial component of health professions education, providing the foundation for developing profession-specific clinical skills. This systematic review aimed to identify evaluated assessment tools for the CLE in health professions education and to report their measurement properties.Methods: This systematic review was preregistered (IDESR000098), and its protocol was published previously. Eligible studies were peer-reviewed articles in English that developed and validated tools for assessing the CLE among undergraduate health professions students and followed the COSMIN guidelines for systematic reviews of patient-reported outcome measures. Multiple electronic databases, including MEDLINE, the Cochrane Library, ERIC, Education Research Complete, and CINAHL, were searched; studies were independently screened, and data were extracted. Data were synthesized using best-evidence synthesis according to COSMIN guidelines.Results: Of the 6,236 articles included in title and abstract screening, 55 were eligible for full-text screening. A supplementary search identified 13 additional articles, resulting in 40 included articles. Overall, 28 tools were identified, with 4 tools (PET, PET-Midwifery, DECLEI, and MidSTEP) demonstrating sufficient content validity. Only MidSTEP demonstrated sufficient structural validity.Conclusion: Only a minority of the included tools provided sufficient evidence of content and structural validity according to the COSMIN criteria. This finding indicates a systemic need for higher standards in monitoring clinical placements and identifies tools that should be re-evaluated and supported by additional research. Limitations include the exclusion of EMBASE and gray literature and the reliance on studies that predominantly used psychometric-first rather than content-validity-first designs.
This scoping review examined research applying digital twins in nursing practice and education and summarized their application domains, methods, outcomes, and implications. A human digital twin is a virtual health replica modeled from real-world data. This study followed the 5-stage scoping review process proposed by Arksey and O’Malley. Two researchers independently conducted the literature search without restrictions on publication year. From April 1 to 15, 2026, the Cochrane Library, PubMed, Embase, CINAHL, ERIC, and RISS databases were searched, and 15 studies were ultimately included. Digital twin applications were identified in 3 major domains: clinical practice and patient-centered care, education and training, and decision-making and workflow management. Application methods and outcomes varied according to technological implementation and included (1) modeling and data-driven prediction, (2) development of immersive learning and practice-training environments, and (3) system integration and decision-support frameworks. In clinical settings, multimodal patient data can be analyzed using artificial intelligence and machine learning to generate a virtual persona resembling the patient, thereby facilitating real-time personalized nursing care and self-management. In educational settings, digital twins can provide realistic and safe learning environments that enhance training effectiveness. Digital twins show substantial potential to advance predictive and personalized nursing in both clinical practice and education. Their data-driven capabilities are expected to contribute to innovative applications in future nursing practice and educational environments.
PURPOSE:The objectives of this study were to develop a 2-item abbreviated version of the Program Sense of Belonging questionnaire (ProSBq) and to evaluate its ability to identify student physical therapists with relatively low valued competence and social acceptance. METHODS:A cross-sectional study was conducted using survey data from 634 students enrolled in physical therapist education programs across the United States. The 10-item ProSBq was used to assess 2 dimensions of belonging: valued competence and social acceptance. Principal component analysis was performed to identify representative items for each subscale, with 1 item selected per subscale. Pearson product-moment correlations were used to examine relationships between the single items and their corresponding subscale scores. Classification performance was evaluated by assessing how accurately the single-item responses classified students reporting a relatively low sense of valued competence and social acceptance, based on their full ProSBq subscale scores. Multiple single-item response thresholds were examined to assess classification accuracy. RESULTS:The single items demonstrated strong relationships with their corresponding subscale scores (r=0.63-0.80, with part-whole correction). For valued competence, sensitivity increased from 53.6% to 92.9%, whereas specificity decreased from 96.3% to 73.5% when a more inclusive threshold was used. A similar sensitivity-specificity tradeoff was observed for social acceptance. Receiver operating characteristic curve analyses demonstrated excellent discrimination (area under the curve ≥0.90). CONCLUSION:Single ProSBq items demonstrated strong relationships with full valued competence and social acceptance subscale scores and acceptable classification performance. The abbreviated 2-item ProSBq may provide a practical and efficient method for identifying students experiencing low valued competence or social acceptance.
PURPOSE:This systematic review aimed to identify and critically evaluate instruments assessing the ethics of teaching and related moral constructs among educators, with a focus on their psychometric properties and applicability to health professions education. METHODS:A systematic search was conducted in PubMed, ERIC, Scopus, and Emerald Insight databases through January 31, 2026. Only English-language studies were included. Measurement properties were evaluated using COSMIN (consensus-based standards for the selection of health measurement instruments) and COSMIN-modified GRADE (Grading of Recommendations Assessment, Development, and Evaluation) approaches. RESULTS:Of 246 records, 6 instruments met the inclusion criteria: ESSQ (Ethical Sensitivity Scale Questionnaire), ELS (Ethical Leadership Scale), EEQ (Ethical Evaluation Questionnaire), TEPI (Teaching-Profession Ethical Principles Inventory), TCPERSS (Teachers' Compliance with Professional Ethics in Relations with Students Scale), and MCI (Moral Competency Inventory). Psychometric properties were sufficiently reported for selected domains, primarily internal consistency and structural validity (Cronbach's α=0.74-0.97). However, construct validity (hypothesis testing), test-retest reliability, and cross-cultural validation were inconsistently reported. The quality of evidence was moderate because of limited cross-context validation. Notably, no tools were specifically developed for health professions education. Most identified instruments focused on classroom pedagogy, potentially overlooking clinical instruction, bedside teaching, and workplace-based learning, where power dynamics and clinical pressures coexist. Developing tools that capture the "ethics of the clinical encounter" would help more accurately reflect the realities faced by health professions educators. CONCLUSION:Existing instruments demonstrate sufficient psychometric properties in general education but reveal critical measurement gaps for health professions education. These findings provide an empirical basis for developing context-specific instruments to improve the evaluation of ethical teaching in clinical and healthcare settings.
Purpose: This study aimed to develop pilot clinical skills assessment (CSA) modules for Korean medicine-specific procedures and to examine their preliminary appropriateness, perceived necessity, and feasibility as a foundation for future licensing-related assessment development.Methods: A participatory action research framework, supplemented by qualitative interviews, was used to develop 4 CSA modules—acupuncture, Chuna manual therapy, pulse diagnosis, and constitutional diagnosis—in collaboration with expert evaluators, students, and standardized patients. The modules were implemented as formative examinations for third-year Korean medicine students, after which semi-structured interviews were conducted to obtain feedback on module content, implementation processes, and scoring procedures. Each module was also reviewed using the RUMBA checklist (Realistic, Understandable, Measurable, Behavioral, and Achievable), together with ratings of perceived necessity and feasibility for possible future use in licensing-related assessment. Interview data were analyzed inductively at the level of individual responses and then compared across modules and participant groups.Results: Qualitative analysis yielded 3 themes: content and scoring criteria, physical environment or simulators, and education or training. Participants emphasized the need to make key aspects of performance more observable, improve authenticity through simulators or task trainers, and strengthen the capacity of scoring systems to distinguish between levels of student performance. Across all modules, mean RUMBA scores were high in the understandable, behavioral, and achievable domains, whereas measurability was more problematic, especially for pulse diagnosis.Conclusion: These pilot findings clarify both the strengths and the limitations of Korean medicine-specific CSA modules. The modules received favorable ratings for understandability and achievability, whereas lower ratings for measurability and realism identified priorities for refinement before wider use. This study provides preliminary guidance for the continued development and broader evaluation of Korean medicine-specific performance assessments.
Developing objective structured clinical examination (OSCE) stations is time-consuming for medical teachers. We aimed to evaluate the ability of a large language model (LLM) to generate ready-to-use OSCE stations. Five OSCE stations generated by the LLM GPT-4o were evaluated by 7 expert assessors using a 5-point Likert scale and compared with 5 teacher-written stations targeting similar learning objectives. A station was considered to be of good quality if most assessors responded “agree” or “strongly agree” to the statement “The station is good enough to be used by students.” All teacher-written stations were rated as being of good quality, compared with only one GPT-4o-generated station. The LLM produced adequate clinical scenarios when reference knowledge was provided and tasks were clearly ordered, but it failed to generate reliable assessment grids. Careful review by teachers remained essential. GPT-4o failed to consistently produce fully ready-to-use OSCE stations.