Skill retention and decay are critical in robotic surgical simulation training. While performance decay has been studied in virtual reality platforms, its effects in high-fidelity biotissue drills remain underexplored. We evaluated both the impact of training breaks and the benefits of continued practice beyond proficiency among general surgery residents performing simulated robotic bowel anastomoses. This retrospective study analyzed 132 h of robotic simulation from 45 residents who reached proficiency (OSATS ≥ 28) on a high-fidelity bowel anastomosis drill. Two cohorts were included: a skill decay group (n = 30) who returned ≥ 1 month after achieving proficiency, and a continued training group (n = 15) who completed ≥ 3 additional sessions post-proficiency. OSATS scores and task times were compared between initial and follow-up sessions. Skill decay was defined as a ≥ 2-point OSATS drop or ≥ 6-min time increase. Statistical tests included paired comparisons, ROC analysis, and multivariable logistic regression. In the skill decay group, the median interval was 91 days (IQR 63–134). Task time increased from 30.5 to 34.5 min (p = 0.03), while OSATS declined from 30 to 28 (p = 0.001). By 6 months, OSATS scores dropped 17.6
Surgical Safety Checklist (SSC) consisting of Timeout and Debrief performed at appropriate time in the operating room (OR) has shown to be a critical and effective tool to reduce errors and improve patient safety. Simulation-based training of OR team members is critical for performing SSC but remains challenging using traditional methods. We have developed a VR simulator that utilizes the Operating Room Black Box (R) (ORBB) recordings, Natural Language Processing (NLP) techniques and Large Language Models (LLM) to simulate and evaluate realistic scenarios. Virtual Humanoid Avatars representing surgical team members, allowed trainees to interact with and learn from scenarios derived directly from ORBB Recordings. Trainees' speech transcribed using Whisper AI were analyzed in real-time using LLM fined tuned with SSC assessment rubrics. We compared zero-shot LLM-based evaluations against ORBB grades using accuracy and F1-score. For Timeout, GPT-OSS 20B achieved 0.88 accuracy / 0.91 F1, and Gemma 3-27B 0.87 / 0.90. For Debrief, both models achieved 0.90 accuracy, with GPT-OSS slightly higher (0.92 F1) than Gemma (0.91 F 1). The results show that the VR simulator for training in SSC with automated assessment model using LLM is highly realistic, reliable and can be used for improving operating room team performance.
Laparoscopic hiatal hernia repair is a common foregut procedure, yet high-fidelity simulation tools for training and assessment remain limited. Previously, task-specific metrics (TSMs) were developed and validated at a single institution to evaluate performance during simulated laparoscopic crural repair. This study assessed the external validity and reliability of these TSMs using a broader, international cohort. Participants at the 2024 SAGES Education and Innovation Center performed a simulated laparoscopic crural repair using an inanimate silicone model housed in a standard laparoscopic trainer box. Demographic data and surgical experience were collected via pre-survey. Performances were video-recorded and independently assessed by three blinded raters using both the Objective Structured Assessment of Technical Skills (OSATS) and TSMs. Participants were stratified into novice (PGY1–2) or experienced (PGY3–5, fellows, attendings). Inter-rater reliability was evaluated using intraclass correlation coefficient (ICC). Performance comparisons were analyzed using the Wilcoxon–Mann–Whitney test. Linear regression modeled training level as an ordered predictor, adjusting for simulation exposure. Correlation between OSATS and TSM scores was measured using Spearman’s rank correlation. Thirty-three participants were enrolled; 30 completed the task (attendings: 9; fellows: 3; PGY1–5: 18). 22 were U.S.-trained and 8 internationally trained. Inter-rater reliability was high for both scoring methods (ICC = 0.85; p < 0.001). Experienced participants scored significantly higher than novices on both OSATS (20 vs. 14; p = 0.039) and TSMs (39 vs. 18.5; p = 0.02). In adjusted models, robotic simulation training was independently associated with higher OSATS (β = 4.67, p = 0.021) and TSM scores (β = 15.01, p = 0.024). TSM and OSATS scores were strongly correlated (R = 0.88; p < 0.001). TSMs demonstrated strong external validity and reliability, effectively discriminating between experience levels and supporting their use as standardized assessment tool in surgical education and global training programs.
Simulation is an important aspect in surgical training as it pertains to gaining proficiency in technical maneuvers. Our previous work in laparoscopic hiatal hernia repair has established task-specific metrics by expert consensus, specifically related to the diaphragmatic crural closure. The aim of this study was to establish the validity and reliability of these task-specific metrics using an inanimate model for eventual incorporation in a virtual reality simulator. Study participants were video recorded performing a crural closure using an inanimate model compatible with the Advanced Training in Laparoscopic Suturing (ATLAS) training box. Two blinded raters evaluated performances using the Objective Structured Assessment of Technical Skills (global rating scale) and task-specific metrics. We assessed Interrater Reliability (IRR) using the Intraclass Correlation Coefficient (ICC). Between-group differences were evaluated using Wilcoxon Rank Sum tests. Additionally, Spearman’s rank correlation coefficient was computed to evaluate the correlation between global and task-specific scores. The study included 33 participants categorized by years of postgraduate training: (novices were PGY1–2, intermediate were PGY3–6 and experts were faculty). When compared to novices, intermediates demonstrated superior technical performance on both global (13.5 vs 23.0, p = 0.015) and task-specific metrics (30.5 vs 52.7, p = 0.006) and completed the task faster (17.39 vs 43.68, p = 0.004). IRR was high for global (ICC = 0.86, p < 0.001) and moderately high for task-specific metrics (ICC = 0.78, p < 0.001). There was a strong positive correlation between the global and task-specific scores (R = 0.95, p < 0.001). Significant discriminant validity was observed for both global and task-specific ratings in our novel crural closure model. These validated metrics of assessment can be incorporated into a future virtual reality simulator.
Robotic surgery demands specialized training to ensure proficiency and patient safety. Deliberate practice using Virtual Reality (VR) robotic simulation platforms provides a safe method for skill acquisition. The first part of our Initial curriculum (IC) consisted of 33 VR drills on the SimNow Simulator, which was refined in 2021 to 19 VR drills using a Simulation-Based Mastery Learning (SBML) approach. This study evaluates the feasibility and effectiveness of our refined curriculum (RC). A total of 87 general surgery residents were included. IC was completed by 41 residents, while 46 residents completed the RC. Metrics such as console time, pre- and post-test VR drill scores, and inanimate drill performance were assessed. Statistical analyses included independent or paired t-tests and Mann–Whitney U, or Wilcoxon Signed Rank tests for non-parametric data. In the IC, 83
In modern medical education, proficiency-based simulation has become a pivotal tool for safely developing critical clinical skills for healthcare professionals. In 2016, the surgery department at UT Southwestern introduced a suturing simulation curriculum as part of the Transition to Clerkship (TTC) course for medical students. This study examines the evolution of this simulation activity over an 8-year period, highlighting key adaptations made in response to the COVID-19 pandemic and ongoing improvements based on iterative annual curriculum evaluations. We hypothesize that iterative curriculum modifications informed by student feedback and designed to address challenges such as those posed by the COVID-19 pandemic would maintain curriculum effectiveness, defined as consistent student satisfaction and pass rates over the 8-year study period. In this Institutional Review Board approved mixed methods study, we conducted a retrospective examination of de-identified survey responses capturing quantitative and qualitative feedback from second-year medical students participating in the simulated suturing exercise (simple interrupted suture with instrument tying). Our analysis aimed to triangulate the association between student feedback and longitudinal variations in pass/fail rates of the simulated interrupted suture activity, specifically examining the periods before, during, and after the COVID-19 pandemic. Despite pandemic disruptions, feedback-driven adaptations maintained course effectiveness. Student perceptions remained stable, with positive responses noted for remote teaching and curriculum adjustments. Engagement increased over time, as evidenced by higher participation in optional activities. Over 8 years, this study showcases sustained curriculum refinement amidst COVID-19 challenges. Proactive collection, analysis, and implementation of feedback yielded consistent benefits in terms of student satisfaction and performance, with no significant shifts observed across different phases of the pandemic. Our findings underscore the importance of adapting educational approaches to meet the evolving needs and preferences of learners, highlighting a collaborative model that prioritizes feedback integration.
Many simulation centers are challenged in securing well-trained faculty who are able to teach learners using simulation because of competing professional priorities. In addition, surgeon educators interested in teaching simulation may need additional training to improve their teaching, assessment, and debriefing skills. To fill this gap, a course entitled Simulation in Surgical Education was developed by the American College of Surgeons (ACS) Division of Education. The course leveraged expert faculty, the co-branded ACS/Association for Surgical Education (ASE) Medical Student Simulation-Based Surgical Skills Curriculum and the resources of an ACS-Accredited Education Institute. This national 2-day experiential train-the-trainer course utilized seven core skills applicable to all medical students undergoing their surgical clerkship. Modules included suturing, knot-tying, nasogastric tube, Foley catheter, central line with ultrasound, intraosseous line, and basic airway. No previous simulation experience was necessary. Surgical faculty interested in simulation were invited and senior surgeons were encouraged to apply. Hands-on experiential skills training with feedback from experts in simulation was supplemented with didactics on the educational foundations of simulation-based training, deliberate practice, types of simulation modalities, prebriefing and debriefing, assessment tools, and various simulation-based curricula. Participants taught and assessed medical students using simulators as well as received student, peer and expert feedback regarding their newly acquired teaching and assessment skills. Course evaluations addressed the value and quality of didactics and simulation stations (scale of 5 = high to 1 = low). The course was fully subscribed with 28 participants, including 14 senior surgeons (> 25 years in practice) as well as clerkship and program directors, simulation champions, and surgical faculty. Fourteen medical students participated as subjects. Twenty-five participants completed all course evaluations. Overall score for meeting course objectives was 4.88. Didactics received value ratings of ≥ 4.57 with enhancing feedback (4.88) and deliberate practice (4.80) being highest rated. All simulation station activities received value ratings ≥ 4.46. The highly rated course was successful in training surgeons at various stages of their careers to use simulation in teaching and assessing skills, especially in the context of the ACS/ASE skills curriculum. The importance of experiential train-the-trainer courses to assure high-quality implementation of the curriculum was underscored. Future plans include evaluating participant success using simulation at their institutions and scaling the course to address the needs of additional surgeon educators.
Global assessment scales of surgical technical performance are well documented; however, evidence supporting task-specific metrics (TSM) is limited. This study evaluates the utility of TSM developed for simulated robotic pancreatojejunostomies on bio-tissue. Fifty videos of bio-tissue pancreatojejunostomies performed by residents as part of our robotic curriculum were graded using both the Objective Structured Assessment of Technical Skills (OSATS) and a TSM scale consisting of 6 domains by a trained blinded grader. Messick’s validity framework was used to assess the utility of the TSM. The internal structure validity was established by computing inter-rater reliability using the intraclass correlation coefficient (ICC) on a subset of 10 videos assessed by an additional grader. Using blinded graders established the response process. Weighted total TSM scores were calculated with partial least squares (PLS) regression, and the Spearman’s Rank correlation with OSATS established the relations to other variables. ICC was 0.74 for OSATS and 0.95 for TSM scores. Mean total scores for OSATS were 25.48 ± 3.27 and 51.32 ± 5.8 for TSM. Prior to PLS regression, a strong correlation was observed between OSATS and TSM scores (R = 0.86, R2 = 0.78; p < 0.001). After regression, the computed weights were 0.29 for posterior mattress suture, 0.49 for posterior row of pancreatic duct sutures, 1.07 for duct stent placement, 0.47 for anterior row of duct sutures, and 0.3 for anterior seromuscular stitches. After regression, the correlation between the TSM weighted sum and OSATS scores was significantly improved (R = 0.88, R2 = 0.82; p < 0.001). TSM scores can be an effective and reliable tool to assess resident performance. They can also be used to provide feedback to surgical trainees.
Automated assessment of surgical skills using artificial intelligence (AI) is valuable for trainees to obtain instantaneous feedback. After bimanual tool motions are captured, the derived kinematic metrics have shown to be reliable predictors of performance in laparoscopic tasks. Implementing automated tool tracking assessment requires time-intensive human annotation. We have developed AI-based tool tracking using the Segment Anything Model (SAM) to eliminate the need for human annotators. Here we describe a study to evaluate the usefulness of our tool tracking model in automated assessment of performance in a laparoscopic suturing task of the fundoplication procedure. An automated tool tracking model was applied to recorded videos of Nissen fundoplication on porcine bowel. The participating surgeons were grouped into novices (PGY1-2) and experts (PGY3-5, fellow and attendings). The beginning and the ending of each of the suturing steps were segmented and the motions of the left and right tools were extracted. A low-pass filter with a cut-off frequency of 24 Hz was then applied to remove any noise. Automated assessment of performance was implemented using both supervised and unsupervised models, and an ablation study was performed to compare the performance. For the supervised learning model, kinematic features included root mean square (RMS) velocity, RMS acceleration, RMS jerk, total path length, and Bimanual Dexterity in pixel coordinates (x, y) that were extracted and analyzed using Logistic Regression, Random Forest Classifier, Support Vector Classifier, and XGBoost. We further analyzed the performance with a reduced set of features selected using a Principal Component Analysis (PCA). For unsupervised learning, a Denoising Autoencoder (DAE) model with classifiers like a 1-D (Convolutional Neural Network) CNN and traditional Machine Learning models were trained for classification. Data was extracted for 28 participants out of an initial cohort of 38, categorized into 9 novices and 19 experts. The approach for supervised learning utilized kinematic features with Principal Component Analysis (PCA) using a Random Forest model (the best model among other Machine Learning models) and obtained an accuracy of 0.795 ± 0.065 and an F1 score of 0.778 ± 0.071. The approach for unsupervised learning employed a 1-D CNN achieved the best results with an accuracy of 0.817 ± 0.108 and an F1 score of 0.806 ± 0.110. This second approach is superior as it eliminates the need to compute kinematic features. We successfully demonstrated an AI model for automated classification of performance independent of human annotation of surgical videos that can be used to enhance surgical skill evaluation.
Crural repair involves reconstructing the esophageal hiatus. Automated AI-based assessment of intracorporeal suturing could improve trainee performance through objective feedback. We evaluated an AI system designed to assess performance in simulated laparoscopic crural repair without manual annotations. In this institutional review board-approved study, tool motion was tracked from 47 videos of 33 participants (15 novices, 18 experts) performing simulated crural repair. Each suture placement was segmented from needle grasp to final knot, and bilateral tool motion was analyzed. Kinematic features path length, root mean square (RMS) velocity, jerk, and bimanual dexterity were extracted. Noise was removed using a 24 Hz low-pass filter. Machine learning models (logistic regression, random forest, support vector classifier, XGBoost) were trained using tenfold cross validation. An ablation study identified the top-performing model, and group differences were evaluated with the Mann–Whitney U test. Data from all participants were successfully analyzed. Logistic regression with min–max scaling achieved the best performance (74
Accreditation bodies are driving competency-based education in healthcare, prompting curriculum reform. Simulation-based education (SBE) addresses challenges curriculum reform has uncovered, like lack of standardization in bedside teaching. This study explores the impact of an AI-powered Automated System Protocol (ASP) for grading students’ post-encounter notes in Clerkship OSCEs, comparing it to the legacy human grader system. The ASP, utilizing GPT-4, mapped rubric items to prompts. Analyzing post-encounter notes from 684 medical students across four academic years, we compared ASP with legacy Standardized Patient Evaluator (SPE) grades. Time efficiency, cost savings, and ROI analyses assessed educational and financial implications. Significant cost savings and efficiency gains were observed utilizing GPT-4 in comparison to SPEs. The Cost of Investment for ASP totaled 69,112 over 1,150 h. Comparing ASP to three SP graders yielded13,112 in increased costs and initial time investment was required. However, beyond development time ASP execution-only, compared to legacy, showed an ROI of 589.44
The Advanced Training in Laparoscopic Suturing (ATLAS) curriculum was created to address gaps in advanced laparoscopic suturing training and has been previously studied using in-person assessment (IPA). This study aimed to evaluate the reliability of a remote, asynchronous video-based assessment (VBA) protocol by comparing to IPA to build validity data for the ATLAS curriculum. Three surgeons (two trainees, one expert) at a single institution performed five repetitions of each ATLAS task. Each performance underwent IPA by a single rater and remote VBA by two independent raters. Videos were de-identified and randomized prior to scoring. Task scores were calculated using time and error measurements per the ATLAS scoring rubric. The VBA scores were averaged and compared to IPA scores by calculating intraclass correlation coefficients (ICC). A total of 90 task performances were reviewed. A higher number of errors were detected in VBA than in IPA (393.5 vs. 323) with higher detection frequency in nine of fifteen errors. Average VBA scores were lower than average IPA scores across all tasks, but with moderate-to-excellent reliability for each task; the ICC of all tasks combined was 0.93 (P < 0.001, 95
Laparoscopic hiatal hernia (HH) surgery is a complex procedure associated with high recurrence. The objective of this project was to develop a virtual reality (VR) simulator for HH surgery to ultimately enhance surgeon performance. A 48-item survey was created to assess demographics, prior experience, procedure specific training, preferences regarding procedural details, and desired simulator features. A REDCap survey was distributed to 250 expert U.S. foregut surgeons. Their importance ratings of procedural steps and simulation preferences were analyzed using a non-parametric Kruskal–Wallis test. There were 44 complete survey respondents (20
ObjectiveOur institution recently implemented a virtual reality (VR) skills curriculum for general surgery residents using the SimNow simulator. Based on a content alignment study, we revised the curriculum to include only 20 of 33 VR tasks and we added 3 previously validated inanimate tasks. The purpose of this study was to establish expert-derived proficiency levels for all tasks and to evaluate the validity of the scoring for the VR tasks.DesignTwo expert robotic surgeons performed 5 repetitions of each VR and inanimate task. The trimmed mean (lowest scoring attempt and outliers [>2 standard deviations] were eliminated) was defined as the expert level for each task. For the VR tasks, expert levels were compared to resident performance to evaluate validity.SettingThis study was conducted at the University of Texas Southwestern Medical Center (Dallas, TX), a tertiary care academic teaching hospital.ParticipantsTwo expert robotic surgeons participated in this study. The data from 42 residents (PGY2-4) who completed the original curriculum was used to represent novice performance.ResultsComparison of expert levels and resident performance was statistically significant for 15 VR tasks (supporting validity) and approached significance (p = 0.06, 0.09) for 2 VR tasks; expert levels were designated as proficiency levels for these 17 tasks. Group comparisons were clearly not significant (p = 0.2-0.8) for 3 VR tasks; 2 of these 3 tasks were retained as introductory exercises (with 3 repetitions required) and 1 was excluded. For the 3 inanimate tasks, expert levels minus 2 standard deviations were designated as proficiency levels.ConclusionsThis analysis generated validity evidence for 15 VR tasks and established expert-derived proficiency levels for 17 VR tasks and 3 inanimate tasks. Our proposed curriculum now consists of 19 VR and 3 inanimate tasks using the selected proficiency levels. We anticipate that this design will maximize curriculum efficiency and effectiveness.
The advanced training in laparoscopic suturing (ATLAS) program was developed by a working group within the Association for Surgical Education (ASE) simulation committee and is commercially available as of 2022. This study aimed to develop updated multicenter expert-derived proficiency scores for this proficiency-based training curriculum. Four laparoscopic suturing experts were identified through their affiliation with the ASE simulation committee, training in minimally invasive surgical techniques, and earlier work in developing the ATLAS project. After a brief optional warmup period on each task, each expert performed five consecutive repetitions which an independent observer scored. After each task, experts completed ratings using the NASA Task Load Index (TLX) scale across six domains (6–126; higher scores indicate increased workload). Outliers (± 2 standard deviations [SD]) were removed from the data sets; trimmed means and the new SD were used to create proficiency levels. Six of 120 data points were trimmed after identifying 0–2 outliers per task. Per group consensus, a tiered approach was taken to create multiple proficiency levels suitable for learners at various stages. Workload ratings ranged low to moderate for all tasks. Temporal (81 ± 23) and mental (76 ± 24) demand had the highest ratings and varied considerably among experts, whereas performance (44 ± 9) was consistently rated as favorable. We developed multicenter proficiency levels for the updated commercially available ATLAS curriculum using established methodology. Workload ratings reflected the rigorous nature of the ATLAS tasks. We anticipate that the Basic, Advanced, and Expert levels will facilitate longitudinal training for learners.
An Objective Structured Clinical Examination (OSCE) is a critical component of medical education whereby the data gathering, clinical reasoning, physical examination, diagnostic and planning capabilities of medical students are assessed in a simulated outpatient clinical setting with standardized patient actors (SPs) playing the role of patients with a predetermined diagnosis, or case. This study is the first to explore the zero-shot automation of physical exam grading in OSCEs by applying multimodal question answering techniques to the analysis of audiovisual recordings of simulated medical student encounters. Employing a combination of large multimodal models (llava-1.6 7B,13B,34B and GPT-4V), automatic speech recognition (Whisper v3), and large language models (LLMs), we assess the feasibility of applying these component systems to the domain of student evaluation without any retraining. Our approach converts video content into textual representations, encompassing the transcripts of the audio component and structured descriptions of selected video frames. These representations, referred to as "exam stories," are then used as context for an abstractive question-answering problem. A collection of 191 audiovisual recordings of medical student encounters with an SP for a single OSCE case was used as a test bed for exploring relevant features of successful exams. During this case, the students should have performed three physical exams: 1) mouth exam, 2) ear exam, and 3) nose exam. These examinations were each scored by two trained, non-faculty standardized patient evaluators (SPE) using the audiovisual recordings - an experienced, non-faculty SPE adjudicated disagreements. The percentage agreement between the described methods and the SPEs' determination of exam occurrence as measured by percentage agreement varied from 26% to 83%. The audio-only methods, which relied exclusively on the transcript for exam recognition, performed uniformly higher by this measure compared to both the image-only methods and the combined methods across differing model sizes. The outperformance of the transcript-only model was strongly linked to the presence of key phrases where the student-physician would "signpost" the progression of the physical exam for the standardized patient, either alerting when they were about to begin an examination or giving the patient instructions. Multimodal models offer tremendous opportunity for improving the workflow of the physical examinations' evaluation, for example by saving time and guiding focus for better assessment. While these models offer the promise of unlocking audiovisual data for downstream analysis with natural language processing methods, our findings reveal a gap between the off-the-shelf AI capabilities of many available models and the nuanced requirements of clinical practice, highlighting a need for further development and enhanced evaluation protocols in this area. We are actively pursuing a variety of approaches to realize this vision.### Competing Interest StatementThe authors have declared no competing interest.### Funding StatementThis study did not receive external funding.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:Ethics committee/IRB of UT Southwestern Medical Center gave ethical approval for this work.I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesRaw OSCE audio-video recording data of the students (FERPA protected) are not publicly available. Privacy sensitive analytical outputs may be made available upon request to authors.
Objective Accreditation bodies are driving competency-based education in healthcare, prompting curriculum reform. Simulation-based education (SBE) addresses challenges curriculum reform has uncovered, like lack of standardization in bedside teaching. This study explores the impact of an AI-powered Automated System Protocol (ASP) for grading students' post-encounter notes in Clerkship OSCEs, comparing it to the legacy human grader system. Methods The ASP, utilizing GPT-4, mapped rubric items to prompts. Analyzing post-encounter notes from 684 medical students across four academic years, we compared ASP with legacy Standardized Patient Evaluator (SPE) grades. Time efficiency, cost savings, and ROI analyses assessed educational and financial implications. Results Significant cost savings and efficiency gains were observed utilizing GPT-4 in comparison to SPEs. The Cost of Investment for ASP totaled $69,112 over 1,150 hours. Comparing ASP to three SP graders yielded $13,112 in increased costs and initial time investment was required. However, beyond development time ASP execution-only, compared to legacy, showed an ROI of 589.44%, saving $47,877 with 87.5% time efficiency. ASP-execution versus three MD graders demonstrated an even stronger ROI of 797.09%. Conclusion Implementing ASP in medical education provides substantial time and cost savings, enhancing ROI compared to legacy grading models. These findings highlight significant cost savings and efficiency improvements achievable through ASP implementation, positioning automated assessment as an innovative force shaping the future of medical education. By liberating human resources from manual grading and enhancing the immediacy of feedback, this approach contributes to a more efficient, effective, and engaging learning experience.
Simulation and video-based assessment (VBA) offer residents the opportunity to develop operative skills while ensuring patient safety. This study aims to determine whether simulation training can predict residents’ operative performance, focusing on the gastrojejunal (GJ) anastomosis during robotic pancreatoduodenectomy. Twenty-seven general surgery residents completed simulated robotic GJ drills and subsequently performed GJs in the operating room (OR). Both simulated and intraoperative performances were video recorded and retrospectively assessed by two blinded graders using the Objective Structural Assessment of Technical Skills (OSATS) scale, time to completion, and occurrence of errors. Intraoperative GJ OSATS scores were compared in cases with and without Clinically Relevant Delayed Gastric Emptying (CRDGE). Statistical analysis was performed using Spearman’s rho, Chi-square, and Kruskal–Wallis tests. For simulated GJs, the median OSATS score was 29 (IQR 27–33), time to completion was 30 min (IQR 27–35), and 11 cases had at least one error. Intraoperative GJs had a median OSATS of 30 (IQR 27–31), time to completion of 41 min (IQR 36–51), and errors occurred in nine cases. The OSATS score on the simulated GJs demonstrated a significant positive correlation to the OSATS score on the operative GJs (r = 0.74; p < 0.001) and less time to completion (r = − 0.68; p < 0.001). A shorter simulated GJ completion time significantly correlated with a higher intraoperative OSATS score (r = − 0.52; p < 0.01). Residents with at least one error in the simulated GJs had lower OSATS scores and higher times intraoperatively. Those cases with CRDGE had significantly lower intraoperative OSATS scores than those without CRDGE. Performance on a simulated robotic GJ environment is a robust predictor of OR GJ performance, demonstrating predictive validity. VBA of residents’ operative GJ performance is associated with the presentation of CRDGE. Simulation-based training may be crucial to optimizing surgical outcomes before operating on patients.
Background Our institution has established priorities for graduate medical education (GME) simulation which include increasing adoption of, garnering additional financial support for, and creating a core simulation curriculum. Better understanding of the Accreditation Council for Graduate Medical Education (ACGME) simulation requirements will inform our efforts and serve as a guide for other institutions. Objective The purpose of this study was to perform a structured review of ACGME simulation standards using a document analysis to guide GME simulation activities at an institutional level. Methods A document analysis was performed from May 2023 to June 2024 to select and search ACGME Institutional and Program Requirements corresponding to the primary specialties for 21 clinical departments that financially support our simulation center. Content relevant to simulation was identified, and iterative coding with investigator team consensus was performed to assign categories, characterize the requirements, and interpret the findings. Results Twenty-four documents included 120 simulation requirements that were assigned to 12 categories; 70 (58%) requirements were mandatory whereas 50 (42%) were not, and 48 (40%) were simulation-specific, whereas 72 (60%) were simulation-optional. All reviewed specialties had simulation requirements (average 5.4, range 2-12), but the ACGME Institutional Requirements did not. Moderate to strong evidence supported (1) simulation usage by all 21 departments; (2) the need for institutional resource support; and (3) institutional-level patient safety simulation curricula. Conclusions This study identified a large number of simulation requirements, including mandatory patient safety curricula requirements, for all specialties analyzed.
Evaluating medical students' physical skills is vital for assessing and validating the acquisition and utilization of their medical knowledge. However, the conventional method of human-based grading for standardized patient encounter videos is costly and error-prone, requires the participation of experts, and may suffer from inter-rater reliability issues and long processing times. Here we propose a deep learning pipeline to identify, extract and score ear exams in over a thousand Simulation Center COSCE medical encounter videos. Our three-stage approach consists of audio-based pre-segmentation with Whisper and Silero VAD, tool detection with CLIP and body node detection with Detectron2. The results of our pipeline are then compared to human graded output and used for automatic extraction of the most relevant video segments. This approach represents a first step toward our overarching goal of expediting and enhancing the quality of the debriefing process following standardized assessments.