PROBLEM:Medical educators pay increasing attention to the potential utility of narrative data for assessment, but lack of efficient and standardized ways of interpreting the data have limited its use. Natural language processing (NLP) algorithms could provide new insights for using narrative data within assessment especially identifying at-risk or low-performing students. APPROACH:Assessment data were reviewed from 16 cohorts of medical students from the University of Cincinnati College of Medicine (graduating classes of 2006-2022). A T-score Average (TSA) was calculated for each student based on clerkship assessment data. Narrative data from core clerkship evaluations in responses to the prompt "any opportunities for improvement" were utilized for analysis. The narrative data and calculated TSA were then used to train and test 4 NLP models with the goal of utilizing NLP to identify at-risk students as defined by the bottom 10% average TSA score. OUTCOMES:Based on typical NLP methods, the developed NLP models all performed adequately in identifying student performance with an overall accuracy of 0.8 across all 4 models. However, none of the NLP models was able to identify students within the bottom 10% of performance. During this process, we uncovered the presence of "copy/paste" behavior, a previously undocumented phenomenon within narrative data where preceptors duplicated comments from 1 student to the next. Training NLP models including "copy/paste" comments improved NLP ability to identify students within the bottom and top 10% of performance. NEXT STEPS:NLP models were unable to accurately identify at-risk students. Model accuracy increased with the inclusion of "copy/paste" comments, indicating there may be some discriminatory functionality within aspects of narrative data beyond keywords such as lexical diversity and word quantity. Future work will explore how these other aspects of narrative data associate with performance and utilize more advanced large language models of analysis.Based on typical natural language processing (NLP) methods, the tested NLP models performed adequately in identifying student performance with an overall accuracy of 0.8 across all models. However, none of the NLP models was able to identify students within the bottom 10% of performance, revealing widespread "copy/paste" behavior in evaluations that may limit narrative data's discriminatory utility.
Jailbreaking in Large Language Models (LLMs) threatens their safe use in sensitive domains like education by allowing users to bypass ethical safeguards. This study focuses on detecting jailbreaks in 2-Sigma, a clinical education platform that simulates patient interactions using LLMs. We annotated over 2,300 prompts across 158 conversations using four linguistic variables shown to correlate strongly with jailbreak behavior. The extracted features were used to train several predictive models, including Decision Trees, Fuzzy Logic-based classifiers, Boosting methods, and Logistic Regression. Results show that feature-based predictive models consistently outperformed Prompt Engineering, with the Fuzzy Decision Tree achieving the best overall performance. Our findings demonstrate that linguistic-feature-based models are effective and explainable alternatives for jailbreak detection. We suggest future work explore hybrid frameworks that integrate prompt-based flexibility with rule-based robustness for real-time, spectrum-based jailbreak monitoring in educational LLMs.
Clinical reasoning is an essential skill, yet physicians receive limited feedback. Artificial intelligence holds promise to fill this gap. We report the development of both named entity recognition (NER), logic-based and large language model (LLM)-based assessments of CR documentation in the electronic health record (EHR) across two institutions. Two note sets were retrieved from the EHR at each institution (NYU Grossman School of Medicine (NYU) and University of Cincinnati College of Medicine (UC)): 1) retrospective dataset comprised of internal medicine resident admission notes from July 2020-December 2021 (n=700 NYU notes, n=450 UC notes) and 2) prospective validation dataset from July 2023-December 2023 (n=155 NYU notes, n=92 UC notes). Using a validated human gold standard for assessment of CR documentation, the R-DEA tool, clinicians rated notes for D (differential diagnosis) and EA (explanation of reasoning) quality, each on 3-point scales (D0, D1, D2 and EA0, EA1, EA2). Model training occurred accordingly on the retrospective datasets: 1) NYU development of NER, logic-based model with validation at UC, 2) NYU fine tune training of LLM NYUTron (a BERT-like (Bidirectional Encoder Representation with Transformer) LLM with about 110 million parameters that has been pre-trained on 7.25 million clinical notes), 3) NYU fine tune training of LLM GatorTron (an open source LLM with 345 million parameters that was pre-trained on over 82 billion words of de-identified clinical text), 4) UC fine tune training of NYU fine-tuned GatorTron, and 5) UC fine tune training of GatorTron. The best performing models were validated with the prospective datasets and performance assessed with F1 scores for the NER, logic-based model and AUROC and AUPRC for the LLMs. At NYU, the NYUTron models were the best performing. The D0 and D2 models with an AUROC 0.87, AUPRC 0.79 and AUROC 0.89, AUPRC 0.86, respectively. The D1 model did not have sufficient performance for implementation. The EA0 and EA1 models also did not have adequate performance so the approach pivoted to create a binary EA2 model (i.e. EA2 vs not EA2) which had excellent performance with an AUROC 0.85 and AUPRC 0.80. At UC, the NER, D-logic-based model was the best performing D model. The F1-scores for the D model on the UC dataset were 0.80, 0.74, and 0.80 for D0, D1, D2, respectively. The UC fine tuning of NYU fine-tuned GatorTron EA2 model had an AUROC 0.75 and AUPRC 0.69. This is the first study to our knowledge to demonstrate the use of LLMs for assessment of CR documentation quality in the EHR across two institutions. Lessons learned can help promote implementation of these technologies across institutions with ranges of technical resources and enhance feedback on the essential skill of CR.
As hospitalists involved in internal medicine and pediatrics residency selection, each of us have read letters of recommendation (LORs) like the one in Box 1. As part of the process of selecting candidates for residency training, LORs hold significant implications for both applicants and programs and are tainted by deep-rooted flaws. These defects continue largely because of our collective failure to confront and address the glaring issues within the process. At best, LORs offer marginal benefit in selecting residents; at worst, they become overt channels for bias, inequity, inequality, and arbitrariness, often devolving into exercises of inanity, untruthfulness, obfuscation, and even propaganda. As such, they decrease the integrity and purpose of the residency selection process. Dear Sirs, As a distinguished professor with over two decades of experience in the medical field and numerous accolades to my name, I am well-versed in recognizing talent. My extensive work, including groundbreaking research and leadership of several high-profile projects, has given me a keen eye for potential. In this spirit, I wish to discuss a recent student, Ms. J. Smith, who was fortunate to rotate on my service for 1 week. While she was part of the team, her involvement was, for the most part, what one would expect from a student at her level. Ms. Smith showed a reasonable understanding of the basics, and she was generally punctual and present during her rotation. Several patients indicated that she was caring, compassionate, and nurturing, and the residents felt she was cooperative and supportive. As you can see from her CV, she was a Division 1 swimmer in college, and despite the rigors of medical school, she has kept her athletic figure. I believe this says a lot about her inner drive. In conclusion, Ms. Smith has completed her rotation under my supervision. I hope this letter assists you in making an assessment based on the comprehensive criteria you hold for potential candidates. Confidently, Dr. John Doe MD, PhD, MBA, FACP, SFHM, GPBS For example, gender bias in residency application LORs has been noted for Radiology, Orthopedic Surgery, Female Pelvic Medicine and Reconstructive Surgery, Cardiovascular Surgery, Emergency Medicine, Pediatrics, Anesthesiology, Radiation Oncology, Ophthalmology, Internal Medicine, and General Surgery, among others.1 Women often find themselves described in these letters with communal traits, such as helpful and caring.2 In contrast, men are more likely to be portrayed with agentic traits, including leader and taking initiative. LORs for women also tend to focus more on personal appearance (such as the misogynistic but real-life example about Ms. J. Smiths' figure in the above letter) and personal life. They also contain more doubt raisers (e.g., "it appears her health and personal life are stable"2), including hesitancy from the recommender, use of faint praise, potentially negative comments, unexplained comments, and irrelevancies.2 Ethnic and racial biases are also prominent in residency LORs, where differences in language can subtly influence readers' perceptions of candidates.3 As with gender, agentic and communal terms are used differently based on a candidate's ethnicity or race. Even apart from bias, LORs compound inequity. The process of obtaining LORs favors already advantaged groups who are more likely to have access to the most influential letter writers. Students often spend an inordinate amount of time searching for the "right" letter writer, often choosing those with titles or positions of power over those who know them best. LORs also tend to focus only on positive aspects of applicants, neglecting the comprehensive portrayal of a candidate's journey, struggles, and growth. This one-sided representation undermines the principle of holistic review (a balanced assessment of an applicant's experiences, attributes, and academic metrics4) by not fully acknowledging the resilience and perseverance shown in overcoming challenges, especially among disadvantaged applicants. Even worse, despite the purported value of holistic review, residency program directors (PDs) often view the demonstration of improvement or overcoming personal setbacks negatively and perceive narratives about growth as coded language for deficits.5 Finally, the interpretation of LORs varies significantly among readers. Studies suggest that readers are not able to discern from letters alone who the top performers are.6 In addition, the recent rise in plagiarism and potential use of artificial intelligence (AI) for generating LORs further undermines their credibility.7 An increasing number of LORs are being produced by generative AI, and readers are unable to reliably differentiate between human- and AI-authored versions.8 For all the reasons discussed above, it is not surprising that LORs have been shown to be poor predictors of residency performance.6, 9 Efforts to improve the process of writing LORs in various medical specialties have been undertaken, primarily through the introduction of standardized letters (SLORs). SLORs employ a uniform format designed to provide consistent and comparative information across all applicants. However, this approach predominantly relies on normative comparisons, where writers rank candidates based on flawed and incomplete data. This leads to grade inflation: in one study of otolaryngology residents, all 10 SLOR attributes for all candidates had a mean above the 80th percentile.10 Moreover, persistent issues such as gender bias, racial bias, and a general lack of validity evidence continue to mar SLOR effectiveness.11, 12 Some argue that occasionally accurate LORs are more compelling than the substantial evidence of their deep flaws. We believe this mindset is a manifestation of common cognitive biases in human reasoning. These include confirmation bias, framing effect, base-rate neglect, visceral bias, Semmelweis reflex, hindsight bias, and premature closure.13 As readers of LORs, we often believe in our own inherent ability to "read between the lines" and "determine the truth" but we attempt both at our own risk. Despite numerous workshops, papers, and initiatives aimed at improving letter-writing skills, it is unrealistic to expect significant behavioral changes from the vast number of letter writers and readers involved in residency selection. The reluctance to recognize the inherent flaws and unfairness in LORs is likely because doing so would lead to the inevitable conclusion: we should cease writing and relying on them. In reflecting on the evolution of the residency selection process, it is crucial to consider the historical context of assessment in medical education. There was a time when assessment amounted to little more than a cursory checkmark exercise, leaving graduate medical education (GME) with little faith in the integrity or quality assurance of graduates emerging from undergraduate medical education (UME). This lack of trust in assessment data led to a reliance on LORs from trusted and respected colleagues. The ethos of "big name" letter writers became a significant factor, compensating for untrustworthy assessment data. However, with the advent of Competency-Based Medical Education (CBME), the assessment landscape has dramatically transformed. Though UME assessment continues to have shortcomings, CBME has undeniably improved the process, offering a more reliable and equitable evaluation of applicants. CBME marks a significant advancement over LORs, now an obsolete tool with the advent of more sophisticated and trustworthy assessment methods. In clinical practice, we uphold the two principles of using evidence to guide decisions and monitoring biases to minimize harm. This ethos should extend to residency selection. Assessment, often used to safeguard societal interests, must be underpinned by credible evidence for the decisions made. LORs fall short in this regard, lacking the necessary evidence to substantiate their validity. If LORs were a medical procedure, they would not gain approval from regulatory bodies. This disparity highlights the urgent need to reevaluate and align residency application assessments, including LORs, with evidence-based standards. LORs serve as channels for bias and inequity, favor well-connected applicants, focus on selected positive attributes at the expense of true holistic assessment, and have little to no validity at predicting performance in residency. We should stop writing and reading LORs for residency selection, now. LORs are far from perfect but calling to remove them based on inequity, overemphasis on positive attributes, and poor validity evidence is like sweeping one leaf in a forest. In a systematic review, Lipman et al. summarize the literature on metrics used for resident recruitment. Study after study shows that grades, standardized test scores, additional degrees, interviews, Medical Student Performance Evaluations (MSPEs), personal statements, honors/awards, and more are all compromised by bias, as well as the potential for cheating and poor predictive validity.14 Calling for the removal of LORs oversimplifies a complex discussion and unnecessarily singles out one part of the residency selection process. LORs are not the problem, they are a symptom of a broken system. LORs are flawed, but one of the main suggestions to increase diversity, mitigate bias, and increase credible decision making in residency selection is holistic review.15, 16 The very idea of holistic review is that each piece of data is imperfect, and only through the systematic review of all data with a diversity of opinions (i.e., groups) can we begin to minimize bias in residency selection.4 So, will one less piece of data really lead us to less bias, or will it just shift our emphasis onto another piece of biased data? The argument to stop writing and reading LORs on the grounds of bias ignores the fact that all metrics and the entire process of residency selection are compromised by bias. Contrary to what the Point authors state, LORs might be a tool to increase equity. The Point authors have made the argument that LORs compound inequity since all applicants do not have the same access to letter writers. This is an example of equality, and we agree that equality is neither possible nor desirable. Equity seeks to get each person what they need, in hopes of reaching an equal outcome. With this framing in mind, LORs may provide a unique opportunity to promote equity in the residency selection process. LORs are a key tool that faculty can use to advocate for students and lift up those that are marginalized or underrepresented.17, 18 Individual faculty have limited control over grades, awards, or the specific opportunities a student may have. But as educators, we can promote applicants in unique ways, describing their passions, interests, and challenges they have overcome in a way they may not be able to highlight for themselves. Given the number of applicants, most experience the residency selection process as high stakes and impersonal, but a LOR is one of the few opportunities for students to connect with faculty. Removing LORs may improve equality, but it will eliminate one of the ways we can promote equity in a selection process where it is currently lacking.19 Empirical data shows residencies really value LORs and put them to good use. In the biennial survey conducted by the National Resident Matching Program, GME programs consistently cite LORs as a main factor in choosing applicants to interview (80%–90% of the time).20 In fact, LORs ranked higher in importance than standardized tests and personal statements. This is amplified in smaller specialties like Dermatology, Vascular Surgery, and Urology where LORs are almost unanimously perceived as important.20 Similar trends are seen in fellowship applications where LORs have magnified significance.21 The Point authors have stated that one shortcoming of LORs is their failure to comprehensively portray all the struggles, growth, and journey of an individual applicant. Setting aside whether this is even a reasonable expectation for one faculty member to comment upon, they go on to acknowledge that PDs may penalize an applicant when there is language in the LOR about growth or improvement. If the Point authors feel a transparent portrayal of each applicant's journey is lacking, we would point them to the MSPE rather than the LORs.22 Regardless, removing LORs will not make for a more transparent residency selection process. If LORs are useless, then why are so many continuing to use them? In a study of 150 Emergency Medicine PDs, only one advocated for removing LORs from residency selection. When asked for the most important characteristics in choosing who to interview, 139 ranked the LOR first.23 In another study of Anesthesia PDs, most agreed there is value in using LORs to choose who to interview and to look for important keywords and phrases.24 In a national survey of Pediatric PDs, commonly used phrases and keywords in LORs were found to be interpreted in a consistent manner. Importantly, this study found that almost 90% of PDs would consider a weaker candidate more favorably if they had a well-crafted LOR, once again underscoring the opportunity of using LORs to promote applicants.18 Clearly, LORs continue to be used because some key decision makers see potential value. Their depiction as rampant sources of bias should be interpreted with caution. In a 2023 systematic review on residency selection, the authors concluded that the case for bias in LORs is mixed and there is some data to support their predictive value. This led them to conclude that there is more evidence to continue using LORs, while USMLE scores, grades, national school ranking, additional degrees, and receipt of awards should have a limited role.14 Maybe we have chosen the wrong metric to debate. The Point authors would like you to believe that assessment in medical education has evolved with the arrival of CBME, but if it is all tainted by bias, has it really evolved?25 CBME was raised in the Point as a solution, rendering LORs as obsolete since we now have reliable, accurate, and trustworthy assessments. This could not feel further from reality in UME, where CBME is challenging to implement, normative assessments still dominate, and students are oriented primarily toward hiding their weaknesses to try and set themselves apart.26 CBME, as currently implemented, is not the solution. In fact, like holistic review, CBME is built on the idea that utilizing many flawed and imperfect assessments will allow for a more complete picture of a trainee's development.27 Therefore, should not their argument to remove LORs also extend to other forms of assessment that CBME holds with high esteem? If the Point authors want to remove any biased and flawed data from the residency selection process, this slippery slope leads to one solution: residents matched by a lottery. Does that feel extreme? The Netherlands have tried a lottery, stopped it, and are now bringing it back.28 The Point authors dismissed SLORs as a potential improvement, citing reasons such as bias and normative comparisons. However, there might be more to the story. In some contexts, the SLOR was more reliably interpreted and reduced the time that residency programs needed to review LORs.29 In a study of Pediatric applicants collecting validity evidence for SLORs, they found it to be moderately reliable, correlate to admission decisions, and differentiate among applicants even though faculty tended to inflate their ratings on the scale.30 The amount of variance attributed to the applicant in this study (i.e., ability for SLOR to differentiate between applicants) is much higher than most assessments found in medical education. Is not this the kind of validity evidence that the Point authors have called for? In fact, building upon the validity evidence, a decision study showed that by reading four SLORs, one could reliably differentiate between applicants.30 This is critical since residency programs need to complete a final rank of all applicants. Evidence that any piece of information will predict success in residency is lacking, but utilizing SLORs seems to be one way to improve residency selection, mitigate bias, and provide programs with the data they desire.31 Finally, the Point authors have beseeched readers to take an evidence-based approach to guide decisions in residency selection. Yet study after study has shown that no piece of data in the entire residency selection process seems to predict future performance.14, 31 LORs are consistently reported as valuable, with the potential to promote equity, while simultaneously being just as flawed as any other metric in the residency selection process. Removing LORs undermines the very idea of holistic review; we believe pulling one thread (i.e., LORs) from the tapestry (i.e., residency selection) has the potential to do more harm than good. The Counterpoint authors have asked: if the entire residency application process is flawed, why then focus solely on letters? Choosing LORs to remove first is not arbitrary. Rather, it is a strategic move to tackle a classic case of normalized deviance within the residency selection process. Normalized deviance, a phenomenon where deviant practices gradually become accepted as normal within an organization, often leads to a lowered standard of ethics and performance.32 To those entrenched in the system, these practices seem routine and acceptable, while they appear problematic to outsiders. In the case of residency selection, LORs are a prominent example of this deviance. They have become a routine part of the process, despite their inherent flaws and lack of fairness. The first step to ending normalized deviance is to acknowledge and make the problem visible. Removing LORs would do this in an instant. Once this step is taken, the focus can then shift to other aspects of the residency application and selection process. The goal is to create a system that is fair, equitable, accurate, valid, and valuable, rectifying not just a single flawed aspect but challenging a pattern of normalized deviance that has been accepted for too long. This approach is not just about removing a single problematic element; it is about taking a stand for greater integrity and effectiveness. Medical education assessment has evolved. CBME's narrative assessments are shared with a wide array of stakeholders including trainees, competency committees, PDs, and institutions, ensuring transparency and collective scrutiny. In contrast, LORs remain limited in visibility, accessible only to the authors and a select few reviewers. In their present form, they are anathema to CBME: isolated high-stakes assessments based on limited data with low-quality validity evidence. As such, all biases, inaccuracies, and inequities are heightened. Medical educators are still learning how to use CBME to clearly define and assist medical students in meeting criteria essential for graduation. These data, collected from many sources, should primarily facilitate formative assessments and feedback, rather than summative judgments.33 Once graduates meet these criteria, medical schools can confidently assert their readiness for residency, backed by concrete validity evidence. LORs would no longer be needed. We do not need to wait for this. LORs are causing harm now and we should stop writing and reading them for residency selection. The authors declare no conflict of interest.
Background: Physician communication during goals of care (GOC) discussions impact experiences for patients and families at end-of-life (EOL). Simulation allows training in a safe environment where feedback from simulated patients (SP), clinicians, and self-reflection can be incorporated. Objectives: To determine if multisource feedback from SP scenarios enriches feedback provided to trainees. Design: Fourth-medical students participated in two SP GOC discussions during an advanced care planning (ACP) curriculum. Students received feedback from SPs and faculty and completed a video review with self-reflection. Setting and Subjects: Forty-seven fourth-year medical students at the University of Cincinnati College of Medicine participated in the curriculum from 2019-2021. Measurements: An inductive thematic analysis of the narrative data was performed examining all sources of feedback from the SP sessions. Results: Six themes emerged from the feedback: the warning shot: words to say and why it helps; acknowledging emotion: verbal vs non-verbal responses; organization: necessity of a clear path; body language: adding to and distracting from the conversation; terminology to avoid: what jargon encompasses and how it impacts patients; and silence: perceived importance by everyone. SP feedback focused on the personal emotional impact of a student's word choice and body language. Faculty feedback focused on specific learning points through examples from the conversation and expanded to hypothetical scenarios. Student self-reflection after video review allowed students to see challenges that they did not notice while immersed in the encounter. Conclusion: Multisource feedback from simulated GOC discussions provides unique insights for students to guide their development in leading difficult conversations.
Purpose: As competency-based medical education (CBME) continues to advance in undergraduate medical education, students are expected to simultaneously pursue their competency development while also discriminating themselves for residency selection. During the foundational clerkship year, it is important to understand how these seemingly competing goals are navigated. Methods: In this phenomenological qualitative study, the authors describe the experience of 15 clerkship students taking part in a pilot pathway seeking to implement CBME principles. These students experienced the same clerkship curriculum and requirements with additional CBME components such as coaching, an entrustment committee to review their data, a dashboard to visualize their assessment data in real-time, and meeting as a community of practice. Results: Students shared their experiences with growth during the clerkship year. They conveyed the importance of learning from mistakes, but pushing past their discomfort with imperfect performance was a challenge when they feel pressure to perform well for grades. This tension led to significant effort spent on impression management while also trying to identify their role, clarify expectations, and learn to navigate feedback. Conclusions: Tension exists in the clinical environment for clerkship students between an orientation that focuses on maximizing grades versus maximizing growth. The former defined an era of medical education that is fading, while the latter offers a new vision for the future. The threats posed by continuing to grade and rank students seems incompatible with goals of implementing CBME.
The traditional undergraduate medical education curriculum focuses on bolstering knowledge for practice and building clinical skills. However, as future clinicians, medical students will be tasked with teaching throughout their careers, first as residents and then as attendings. Here, we describe teaching opportunities for students that foster their development as future teachers and potential clinician educators. These offerings are diverse in their focus and duration and are offered across various levels of the curriculum — including course-based learning, longitudinal electives, and extra-curricular opportunities for medical students who have a passion for teaching.
PROBLEM:Reflective practice is necessary for self-regulated learning. Helping medical students develop these skills can be challenging since they are difficult to observe. One common solution is to assign students' reflective self-assessments, which produce large quantities of narrative assessment data. Reflective self-assessments also provide feedback to faculty regarding students' understanding of content, reflective abilities, and areas for course improvement. To maximize student learning and feedback to faculty, reflective self-assessments must be reviewed and analyzed, activities that are often difficult for faculty due to the time-intensive and cumbersome nature of processing large quantities of narrative assessment data. APPROACH:The authors collected narrative assessment data (2,224 students' reflective self-assessments) from 344 medical students' reflective self-assessments. In academic years 2019-2020 and 2021-2022, students at the University of Cincinnati College of Medicine responded to 2 prompts (aspects that surprised students, areas for student improvement) after reviewing their standardized patient encounters. These free-text entries were analyzed using TopEx, an open-source natural language processing (NLP) tool, to identify common topics and themes, which faculty then reviewed. OUTCOMES:TopEx expedited theme identification in students' reflective self-assessments, unveiling 10 themes for prompt 1 such as question organization and history analysis, and 8 for prompt 2, including sensitive histories and exam efficiency. Using TopEx offered a user-friendly, time-saving analysis method without requiring complex NLP implementations. The authors discerned 4 education enhancement implications: aggregating themes for future student reflection, revising self-assessments for common improvement areas, adjusting curriculum to guide students better, and aiding faculty in providing targeted upcoming feedback. NEXT STEPS:The University of Cincinnati College of Medicine aims to refine and expand the utilization of TopEx for deeper narrative assessment analysis, while other institutions may model or extend this approach to uncover broader educational insights and drive curricular advancements.
Background: Physicians report inadequate training in advance care planning (ACP) discussions despite the importance of these skills for practicing physicians including new residents. Objectives: To evaluate the effectiveness of a novel curriculum to prepare graduating medical students to have ACP discussions. Design: An ACP curriculum was implemented within a new fourth-year medical student elective with a focus on interactive educational methods and simulated experiences. Setting/Subjects: Forty-seven students received the curriculum over 3 years at a medium-sized, urban medical school. Measurements: Students were surveyed regarding attitudes and comfort related to ACP discussions and end-of-life (EOL) topics before and after the course. Additionally, students were asked about baseline experiences in the pre-course survey and perceived effectiveness of the educational methods in the post-course survey. Results: Comfort discussing EOL care decisions without supervision rose from 4% to 36% after the course with none of the students feeling they needed maximal help from a supervisor after the course compared to 51% before the course. All students agree or strongly agreed (Likert 4 or 5) that they felt prepared to discuss patient’s wishes and values in EOL care with a real patient or family after the course. Conclusions: An ACP curriculum can increase student comfort and preparedness to have these conversations as residents. Students found small group discussions and the chance for direct practice with simulated patients to be most helpful. These findings can help guide implementation of ACP curricula in medical education.
BackgroundThe rapid trajectory of artificial intelligence (AI) development and advancement is quickly outpacing society's ability to determine its future role. As AI continues to transform various aspects of our lives, one critical question arises for medical education: what will be the nature of education, teaching, and learning in a future world where the acquisition, retention, and application of knowledge in the traditional sense are fundamentally altered by AI? ObjectiveThe purpose of this perspective is to plan for the intersection of health care and medical education in the future. MethodsWe used GPT-4 and scenario-based strategic planning techniques to craft 4 hypothetical future worlds influenced by AI's integration into health care and medical education. This method, used by organizations such as Shell and the Accreditation Council for Graduate Medical Education, assesses readiness for alternative futures and effectively manages uncertainty, risk, and opportunity. The detailed scenarios provide insights into potential environments the medical profession may face and lay the foundation for hypothesis generation and idea-building regarding responsible AI implementation. ResultsThe following 4 worlds were created using OpenAI’s GPT model: AI Harmony, AI conflict, The world of Ecological Balance, and Existential Risk. Risks include disinformation and misinformation, loss of privacy, widening inequity, erosion of human autonomy, and ethical dilemmas. Benefits involve improved efficiency, personalized interventions, enhanced collaboration, early detection, and accelerated research. ConclusionsTo ensure responsible AI use, the authors suggest focusing on 3 key areas: developing a robust ethical framework, fostering interdisciplinary collaboration, and investing in education and training. A strong ethical framework emphasizes patient safety, privacy, and autonomy while promoting equity and inclusivity. Interdisciplinary collaboration encourages cooperation among various experts in developing and implementing AI technologies, ensuring that they address the complex needs and challenges in health care and medical education. Investing in education and training prepares professionals and trainees with necessary skills and knowledge to effectively use and critically evaluate AI technologies. The integration of AI in health care and medical education presents a critical juncture between transformative advancements and significant risks. By working together to address both immediate and long-term risks and consequences, we can ensure that AI integration leads to a more equitable, sustainable, and prosperous future for both health care and medical education. As we engage with AI technologies, our collective actions will ultimately determine the state of the future of health care and medical education to harness AI's power while ensuring the safety and well-being of humanity.
Inequity in assessment has been described as a "wicked problem"-an issue with complex roots, inherent tensions, and unclear solutions. To address inequity, health professions educators must critically examine their implicit understandings of truth and knowledge (i.e., their epistemologies) with regard to educational assessment before jumping to solutions. The authors use the analogy of a ship (program of assessment) sailing on different seas (epistemologies) to describe their journey in seeking to improve equity in assessment. Should the education community repair the ship of assessment while sailing or should the ship be scrapped and built anew? The authors share a case study of a well-developed internal medicine residency program of assessment and describe efforts to evaluate and enable equity using various epistemological lenses. They first used a postpositivist lens to evaluate if the systems and strategies aligned with best practices, but found they did not capture important nuances of what equitable assessment entails. Next, they used a constructivist approach to improve stakeholder engagement, but found they still failed to question the inequitable assumptions inherent to their systems and strategies. Finally, they describe a shift to critical epistemologies, seeking to understand who experiences inequity and harm to dismantle inequitable systems and create better ones. The authors describe how each unique sea promoted different adaptations to their ship, and challenge programs to sail through new epistemological waters as a starting point for making their own ships more equitable.
BACKGROUND AND OBJECTIVESWritten discharge instructions help to bridge hospital-to-home transitions for patients and families, though substantial variation in discharge instruction quality exists. We aimed to assess the association between participation in an Institute for Healthcare Improvement Virtual Breakthrough Series collaborative and the quality of pediatric written discharge instructions across 8 US hospitals. METHODSWe conducted a multicenter, interrupted time-series analysis of a medical records-based quality measure focused on written discharge instruction content (0-100 scale, higher scores reflect better quality). Data were from random samples of pediatric patients (N = 5739) discharged from participating hospitals between September 2015 and August 2016, and between December 2017 and January 2020. These periods consisted of 3 phases: 1. a 14-month precollaborative phase; 2. a 12-month quality improvement collaborative phase when hospitals implemented multiple rapid cycle tests of change and shared improvement strategies; and 3. a 12-month postcollaborative phase. Interrupted time-series models assessed the association between study phase and measure performance over time, stratified by baseline hospital performance, adjusting for seasonality and hospital fixed effects. RESULTSAmong hospitals with high baseline performance, measure scores increased during the quality improvement collaborative phase beyond the expected precollaborative trend (+0.7 points/month; 95% confidence interval, 0.4-1.0; P < .001). Among hospitals with low baseline performance, measure scores increased but at a lower rate than the expected precollaborative trend (-0.5 points/month; 95% confidence interval, -0.8 to -0.2; P < .01). CONCLUSIONSParticipation in this 8-hospital Institute for Healthcare Improvement Virtual Breakthrough Series collaborative was associated with improvement in the quality of written discharge instructions beyond precollaborative trends only for hospitals with high baseline performance.
An interprofessional pulmonary standardized patient simulation case was used for medical and social work students to work together using motivational interviewing. Following the case, 68 medical students and 20 social work students responded to open-ended questions eliciting reflection on the encounter. Using thematic text analysis, results indicated students learned how different professions approached healthcare issues and communicated to patients, the importance of collaboration, and the need to be patient-centered with an integrated care approach. This case met all four Interprofessional Education Collaborative Competencies. Our results suggest interprofessional standardized patient simulations can assist students in addressing substance use using a team-based approach. With barriers such as time, instructor availability, and limited funding for interprofessional activities, meaningful high-yield activities are essential.
September 27, 2028.The Match died today. The cause of death was indifference.The Match was born in 1951 to Francis Joseph Mullin and John Marshall Stalnaker.1 Early on, it had several congenital disorders corrected by future pediatric surgeon W. Hardy Hendren III,2 leading to the National Resident Matching Program (NRMP) that we know today.3 In 1998, The Match received care from Drs. Alvin E. Roth and Elliott Peranson, allowing it to accommodate couples, myriad specialties, and DO applicants.4 For his work, Dr. Roth won the 2012 Nobel Prize in Economics.5During its heyday, the Match served nearly 50 000 registrants, boasting decades-long streaks of year-over-year increases in applications and positions. Success came with controversy. In 2003, residents sued the NRMP in a monopoly lawsuit. However, Congress gave the Match antitrust protection akin to that of Major League Baseball.6 Afterwards, the Match consolidated its power, showing swagger with its 2012 all-in policy that required programs to fill every first-year position either exclusively through or outside of it. The Match justified its actions by asserting it was protecting residents and that participation was voluntary.7Then, the Match developed health problems, starting with a severe case of application fever.3,8 As the number of applicants increased, students, institutions, and programs found themselves trapped in a series of prisoner's dilemmas where strategies such as overapplication, obfuscation of performance data, and use of flawed and discriminatory screening heuristics seemed logical.9 Sickness spread as participants suffered from perverse incentives, excessive and progressive costs, inequities, time pressures, waste, and mistrust.10Forces outside the Match began taking shape, including competency-based time-variable education,11 situational judgement testing,12 pass/fail grading,13 resident unions,14 affinity matching,15 virtual interviews,16 and implicit bias training,17 all driven by the desire for high-quality skill development, veracity, fairness, and autonomy. Unfortunately, the Match came to be seen in direct opposition to these important principles and turned ill.Physicians suggested many remedies. Signal preferencing initially showed promise, but students realized that it helped programs more than them,18 and they lobbied for increasing choices until signaling became meaningless. Supplemental applications initially provided more signal than noise,19 but students quickly used AI chatbots to create the perfect essay for each program.20,21 Application caps were thought to be a cure,22 but students rebelled when their choices became limited. Prominent specialists recommended application phases,23 but this increased everyone's workload past capacity, and the idea was abandoned. Ultimately, participants in the Match rejected all remedies despite endless calls for more.24 Unfortunately, treatment for the fever did not cure the infection.9The beginning of the end began outside of organized medicine when Ivy League law schools pulled out of the US News & World Report rankings.25 Once the first school acted, the rest followed. On the surface, law school Deans stated they quit the rankings because of flawed methodology. However, they also calculated that forces beyond their control could lead to unfavorable ranking changes, and they asserted their power to declare themselves elite, etching this judgement in history, impossible to broach.25Important academic physicians also called for the end of formal rankings,26 and just like law schools, medical schools followed.27 These events led residencies to realize they could take back their own power by removing themselves from the Match. Several community programs had been “all out” of the Match for years, shortening interview seasons and offering applicants positions on the spot, thereby gaining advantage over their competitors. Once the first Ivy League residency opted out, others had to follow. Students, frustrated by the inflexibility, inequity, restrictions, and workload of the Match, began opting out to join these programs. Remaining programs saw fewer and fewer candidates, and one by one opted out to compete. Yesterday there were no applications in the Electronic Residency Application Service system, and the Match officially died.The Match, like paper charts, celluloid x-ray film, and analog stethoscopes, will be fondly remembered for being important and effective in its time.It is survived by chaos and remorse.
BACKGROUND:Clinical documentation is a key component of practice. Trainees rarely receive formal training in documentation or assessment of their documentation. Effective methods of improving documentation remain unknown.OBJECTIVE:The objective of this study was to determine if the implementation of a documentation curriculum led to improvement in admission note quality.DESIGNS:Admission notes written prior to implementation of the curriculum and after the curriculum intervention were assessed. Notes were assessed from two-time frames for both years to account for improvement with time not associated with the intervention.SETTINGS AND PARTICIPANTS:Admission notes written by University of Cincinnati interns were assessed.INTERVENTIONS:The documentation curriculum consisted of educational sessions and routine admission note assessments with feedback.MAIN OUTCOMES AND MEASURES:Admission notes were assessed via the 16 checklist items and two global assessment items of the Admission Note Assessment Tool (ANAT).RESULTS:Six ANAT items showed statistically significant differences. The review of systems item improved with the intervention only (odds ratio: 3.61, p < .001) while the assessment and plan item 1 and global assessment item 2 improved with time only (β = .08, p = .03 and β = .25, p = .02, respectively) in univariate models. In univariate models the physical exam item, diagnostic data item 2, and global assessment item 1 showed improvement with both intervention and time, respectively, with additive effects seen in models with both intervention and time.CONCLUSION:Several aspects of documentation can improve with a formal documentation curriculum which includes a routine assessment with feedback, and some aspects of documentation improve with time.
PURPOSE Few medical schools offer electives with the goal of teaching medical students to be effective teachers prior to residency. We developed a novel year-long, longitudinal course, the Clinical Teaching Elective (CTE), that develops fourth-year medical students as student teachers within Clinical Skills (CS). APPROACH/METHODS The elective was designed by Clinical Skills (CS) Course Directors and two fourth-year medical students (M4) as a longitudinal elective. The elective involves teaching in the Simulation Center where M4 student instructors teach first and second-year medical students. Each session, in addition to simulated patient case topics, emphasizes application of a key topic within medical education (ie clinical reasoning, reflective practice, dual process reasoning). DISCUSSION Six “teaching takeaways” were crafted to summarize common themes experienced by near-peer medical student educators. Teaching is not about the destination, but rather the diagnostic journey. Students thrive when learning is co-produced. A little bit of praise goes a long way. You can't please every learner. When students struggle, there is more to teach than just the answer. Facilitating learner independent thinking promotes future autonomy. SIGNIFICANCE A novel CTE for fourth-year medical students that emphasizes medical education pedagogy prepares students to serve as educators in residency. The CTE provides an opportunity for medical students to develop into effective clinical educators prior to residency. The focus of our elective on medical education pedagogy furthers medical student understanding of adult learning theory and fosters professional development in teaching clinical reasoning.
BackgroundCommunication failures occur often in the inpatient setting. Efforts to understand and improve communication often exclude patients or are siloed by discipline. ObjectiveWe aimed to identify barriers and facilitators to effective communication within interdisciplinary inpatient internal medicine (IM) teams using a participatory research approach. DesignWe conducted a single-center participatory mixed methods study using group-level assessment (GLA) and concept mapping to iteratively engage stakeholders. Stakeholder groups included patients/families, IM faculty, IM residents, nurses and ancillary staff, and care managers. Stakeholder-specific GLA sessions were conducted. Participants responded to prompts addressing interdisciplinary communication then worked in small groups to synthesize the qualitative data into unique ideas. A subset of each stakeholder group then sorted ideas through a concept mapping exercise. Multidimensional scaling and hierarchical cluster analysis were used to generate a concept map of the data. ResultsParticipants generated 97 unique ideas that were then sorted. The research team chose an eight-cluster concept map representing patient inclusion and engagement, processes and resources, team morale and inclusive dynamics, attitudes and behaviors, effective communication, barriers to communication, the culture of healthcare, and clear expectations. Three larger domains of patient inclusion and engagement, organizational conditions and role clarity, and team dynamics and behaviors were noted. ConclusionUse of a participatory research approach made it feasible to engage diverse stakeholders including patients. Our results highlight the need to identify context-specific facilitators and barriers of interdisciplinary communication. The importance of clear expectations was identified as a prioritized area to target communication improvement efforts.