ISSUE:Despite rapid innovation, health care systems face a persistent 17-year gap between evidence discovery and implementation, undermining efforts to deliver value-based care. Bridging this "know-do gap" is essential to improving outcomes and reducing waste. Existing Learning Health System (LHS) frameworks often lack mechanisms to institutionalize learning at speed and scale. CRITICAL THEORETICAL ANALYSIS:We propose an AI-enabled LHS framework that leverages artificial intelligence (AI) to connect micro-level clinical learning with macro-level organizational decision-making. Grounded in organizational learning theory, our model illustrates how AI accelerates knowledge capture, conversion, and institutionalization via continuous, bidirectional feedback loops. AI enables real-time learning cycles, linking patient-provider data ("micro") to system-wide insights and policy adjustments ("macro"), and back to point-of-care decision support. INSIGHT/ADVANCE:Our framework advances the LHS paradigm by adding speed, scale, and micro↔macro integration. Unlike earlier models, it centers AI not as an adjunct but as a foundational learning engine. Case examples from UCHealth and Mass General Brigham show how AI can drive real-time operational learning and institutional memory through structured governance and data infrastructure. PRACTICE IMPLICATIONS:To implement an AI-LHS, organizations should (1) assess readiness and align on value-based goals; (2) invest in data infrastructure and interoperability; (3) cultivate a learning culture by engaging clinicians and staff; (4) embed AI into continuous improvement cycles with interdisciplinary governance; (5) adopt a sociotechnical approach integrating people, processes, and technology; and (6) ensure safeguards for equity, privacy, and security. These steps allow systems to reduce lag between insight and impact, accelerating value-based care transformation.
We're sold on SoLD (the science of learning and development) for medical educators to use in times of crisis … or anytime. #MedEd
Background: Current telehealth reimbursement guidelines, based primarily on evaluation and management duration, inadequately capture the true clinical complexity and cognitive effort involved in provider responses to secure patient messages. There is a critical need for reimbursement frameworks that reflect the nuanced reality of clinical care delivered through asynchronous telehealth messaging. Materials and Methods: We analyzed 149,499 secure messages (2023-2024) from a large health system. Given the impracticality of manually annotating all threads, we used supervised pseudo-labeling informed by rigorous annotation guidelines (interrater reliability: Cohen's κ = 0.816) applied to a representative subset. Clinical complexity was quantified through features of providers' electronic health records engagement (Domain Engagement [DE]) and message linguistic complexity (Cognitive Judgment [CJ]). A gradient boosting classifier leveraging these complexity-based features was developed and validated via stratified cross-validation, temporal validation, and leave-one-provider-out testing. Interpretability was assessed using permutation importance and Shapley Additive Explanations values. Results: The complexity-based model achieved strong predictive performance (area under the curve [AUC] ∼0.82), substantially outperforming a time-based baseline (AUC ∼0.57). Key predictors included unique medical concepts, breadth of clinical content, and emotional urgency. Traditional metrics like message length and duration were poor predictors. Billable threads showed significantly higher medical density (∼55 vs. ∼33 concepts) and emotional intensity (∼0.6 vs. ∼0.3), supporting DE and CJ validity. Discussion: Complexity-based features more accurately captured providers' cognitive effort than elapsed time alone, consistent with Cognitive Load, Complexity, and Sensemaking theories. Conclusion: This AI-enabled complexity-based model provides a practical, accurate, and explainable foundation for telehealth billing reform, facilitating fairer provider reimbursement and improved telehealth sustainability.
PURPOSE:Electronic health records (EHRs) can provide valuable insights into workflow, clinical reasoning, and personal attributes; however, the indicators for how an individual acts within the EHR (EHR use metrics) are not frequently analyzed. This study examines whether EHR use metrics are associated with internal medicine resident clinical performance. METHOD:In this retrospective cohort study, data on EHR use metrics and achievement of 22 clinical performance measures (CPMs) were collected between November 2021-October 2022 from University of Cincinnati internal medicine residents during a year dedicated to ambulatory care. The CPMs were sorted on an attribution-contribution continuum for subgroup analysis. The EHR use metrics were used for agglomerative hierarchical clustering to group residents with similar EHR behaviors. RESULTS:Thirty residents (11 [37%] male and 19 [63%] female) were included. Clustering with a subset of 10 EHR use metrics resulted in 3 clusters with different clinical performance as indicated by achievement of CPMs. The clusters were characterized as lower-performing (n = 5; mean [SD] CPMs achieved, 11.4 [2.3]; 95% CI, 9.4-13.4), middle-performing (n = 23; mean [SD] CPMs achieved, 15.8 [2.1]; 95% CI, 14.9-16.6), and higher-performing (n = 2; mean [SD] CPMs achieved, 22 [0]; 95% CI, 22-22). After sorting the CPMs on an attribution-contribution continuum, the clusters performed differently in actions (F 2 ,27 = 7.73, P = .002) and screenings (F 2,27 = 9.60, P < .001) but not lab testing (F 2,27 = 2.88, P = .07) or disease control (F 2,27 = 1.01, P = .38). The lower-performing cluster had longer response times and incomplete work, whereas the higher-performing cluster was most responsive and communicative. CONCLUSIONS:Hierarchical cluster analysis of EHR use metrics can identify EHR use patterns associated with resident clinical performance. Clustering provides a framework that will enable programs to apply EHR use metrics to augment resident assessment and feedback.
Background: Although depression is one of the most common mental health disorders outpacing other diseases and conditions, poor access to care and limited resources leave many untreated. Secure messaging (SM) offers patients an online means to bridge this gap by communicating nonurgent medical questions. We focused on self-care health management behaviors and delved into SM initiation as the initial act of engagement and SM exchanges as continuous engagement patterns. This study examined whether those with depression might be using SM more than those without depression.Methods: Patient portal data were obtained from a large academic medical center's electronic health records spanning 5 years, from January 2018 to December 2022. We organized and analyzed SM initiations and exchanges using the linear mixed-effects modeling technique.Results: Our predictors correlated with SM initiations, accounting for 25.1% of variance explained. In parallel, 24.9% of SM exchanges were attributable to these predictors. Overall, our predictors demonstrate stronger associations with SM exchanges.Discussion: We examined patients with and without depression across 2,629 zip codes over five years. Our findings reveal that the predictors affecting SM initiations and exchanges are multifaceted, with certain predictors enhancing its utilization and others impeding it.Conclusions: SM telehealth service provided support to patients with mental health needs to a greater extent than those without. By increasing access, fostering better communication, and efficiently allocating resources, telehealth services not only encourage patients to begin using SM but also promote sustained interaction through ongoing SM exchanges.
ABSTRACT:The shift to pass/fail grading in undergraduate medical education was designed to reduce medical students' stress. However, this change has given rise to a "shadow economy of effort," as students move away from traditional didactic and clinical learning to engage in increasing numbers of research, volunteer, and work experiences to enhance their residency applications. These extracurricular efforts to secure a residency position are subphenomena of the hidden curriculum. Medical schools do not officially require all the activities students need to be most competitive for residency selection; therefore, students, as rational actors, participate in the activities they think will most help them succeed.Here, the authors frame residency application and selection as a complex adaptive system (CAS), which self-organizes without centralized control or hierarchical intent. Individuals in a CAS operate in environments marked by volatility, randomness, and uncertainty-all of which are abundant in the residency selection process. Outcomes in such systems, like the development of a shadow economy, are novel, emergent, and cannot always be anticipated. To address these challenges, the authors suggest the need for deep understanding of the system's elements, interrelationships, and dynamics, including feedback loops and emergent properties. Optimizing the results of a CAS requires incentivizing outcomes over activities, ensuring open information flow, and engaging in continuous monitoring and evaluation.The current pass/fail era and resultant shadow economy of effort risk creating a triple harm by devaluing clinical excellence, burning out medical students, and potentially producing superficial or, worse, inauthentic academic and community work. Medical educators must optimize residency application and selection for cooperative outcomes and design incentives to ensure the outputs of medical education align student, institutional, patient, and societal goals. Without a set of predictive "answers," the authors suggest a process of determining actions to advance this ultimate aim and reduce harm.
High-quality precision education (PE) aims to enhance outcomes for learners and society by incorporating longitudinal data and analytics to shape personalized learning strategies. However, existing educational data collection methods often suffer from fragmentation, leading to gaps in understanding learner and program performance. In this article, the authors present a novel approach to PE at the University of Cincinnati, focusing on the Ambulatory Long Block, a year-long continuous ambulatory group-practice experience. Over the last 17 years, the Ambulatory Long Block has evolved into a sophisticated data collection and analysis system that integrates feedback from various stakeholders, as well as learner self-assessment, electronic health record utilization information, and clinical throughput metrics. The authors detail their approach to data prioritization, collection, analysis, visualization, and feedback, providing a practical example of PE in action. This model has been associated with improvements in both learner performance and patient care outcomes. The authors also highlight the potential for real-time data review through automation and emphasize the importance of collaboration in advancing PE. Generalizable principles include designing learning environments with continuity as a central feature, gathering both quantitative and qualitative performance data from interprofessional assessors, using this information to supplement traditional workplace-based assessments, and pairing it with self-assessments. The authors advocate for criterion referencing over normative comparisons, using user-friendly data visualizations, and employing tailored coaching strategies for individual learners. The Ambulatory Long Block model underscores the potential of PE to drive improvements in medical education and health care outcomes.
Entrustable professional activities (EPAs) and entrustment decision-making have become common language in competency-based education in the health professions. Since its introduction, several other related concepts have been introduced, which has made it more complex to get an overview of the domain. This chapter sets out to discuss OPAs (observable practice activities), EPA specifications, nested EPAs, core EPAs versus elective EPAs, transdisciplinary EPAs, Practice Activities as used by the WHO, retrospective versus prospective entrustment–supervision scales, STARs (statements of awarded responsibility), microcredentials, and hospital privileging. The concepts are defined and elaborated with examples.
As hospitalists involved in internal medicine and pediatrics residency selection, each of us have read letters of recommendation (LORs) like the one in Box 1. As part of the process of selecting candidates for residency training, LORs hold significant implications for both applicants and programs and are tainted by deep-rooted flaws. These defects continue largely because of our collective failure to confront and address the glaring issues within the process. At best, LORs offer marginal benefit in selecting residents; at worst, they become overt channels for bias, inequity, inequality, and arbitrariness, often devolving into exercises of inanity, untruthfulness, obfuscation, and even propaganda. As such, they decrease the integrity and purpose of the residency selection process. Dear Sirs, As a distinguished professor with over two decades of experience in the medical field and numerous accolades to my name, I am well-versed in recognizing talent. My extensive work, including groundbreaking research and leadership of several high-profile projects, has given me a keen eye for potential. In this spirit, I wish to discuss a recent student, Ms. J. Smith, who was fortunate to rotate on my service for 1 week. While she was part of the team, her involvement was, for the most part, what one would expect from a student at her level. Ms. Smith showed a reasonable understanding of the basics, and she was generally punctual and present during her rotation. Several patients indicated that she was caring, compassionate, and nurturing, and the residents felt she was cooperative and supportive. As you can see from her CV, she was a Division 1 swimmer in college, and despite the rigors of medical school, she has kept her athletic figure. I believe this says a lot about her inner drive. In conclusion, Ms. Smith has completed her rotation under my supervision. I hope this letter assists you in making an assessment based on the comprehensive criteria you hold for potential candidates. Confidently, Dr. John Doe MD, PhD, MBA, FACP, SFHM, GPBS For example, gender bias in residency application LORs has been noted for Radiology, Orthopedic Surgery, Female Pelvic Medicine and Reconstructive Surgery, Cardiovascular Surgery, Emergency Medicine, Pediatrics, Anesthesiology, Radiation Oncology, Ophthalmology, Internal Medicine, and General Surgery, among others.1 Women often find themselves described in these letters with communal traits, such as helpful and caring.2 In contrast, men are more likely to be portrayed with agentic traits, including leader and taking initiative. LORs for women also tend to focus more on personal appearance (such as the misogynistic but real-life example about Ms. J. Smiths' figure in the above letter) and personal life. They also contain more doubt raisers (e.g., "it appears her health and personal life are stable"2), including hesitancy from the recommender, use of faint praise, potentially negative comments, unexplained comments, and irrelevancies.2 Ethnic and racial biases are also prominent in residency LORs, where differences in language can subtly influence readers' perceptions of candidates.3 As with gender, agentic and communal terms are used differently based on a candidate's ethnicity or race. Even apart from bias, LORs compound inequity. The process of obtaining LORs favors already advantaged groups who are more likely to have access to the most influential letter writers. Students often spend an inordinate amount of time searching for the "right" letter writer, often choosing those with titles or positions of power over those who know them best. LORs also tend to focus only on positive aspects of applicants, neglecting the comprehensive portrayal of a candidate's journey, struggles, and growth. This one-sided representation undermines the principle of holistic review (a balanced assessment of an applicant's experiences, attributes, and academic metrics4) by not fully acknowledging the resilience and perseverance shown in overcoming challenges, especially among disadvantaged applicants. Even worse, despite the purported value of holistic review, residency program directors (PDs) often view the demonstration of improvement or overcoming personal setbacks negatively and perceive narratives about growth as coded language for deficits.5 Finally, the interpretation of LORs varies significantly among readers. Studies suggest that readers are not able to discern from letters alone who the top performers are.6 In addition, the recent rise in plagiarism and potential use of artificial intelligence (AI) for generating LORs further undermines their credibility.7 An increasing number of LORs are being produced by generative AI, and readers are unable to reliably differentiate between human- and AI-authored versions.8 For all the reasons discussed above, it is not surprising that LORs have been shown to be poor predictors of residency performance.6, 9 Efforts to improve the process of writing LORs in various medical specialties have been undertaken, primarily through the introduction of standardized letters (SLORs). SLORs employ a uniform format designed to provide consistent and comparative information across all applicants. However, this approach predominantly relies on normative comparisons, where writers rank candidates based on flawed and incomplete data. This leads to grade inflation: in one study of otolaryngology residents, all 10 SLOR attributes for all candidates had a mean above the 80th percentile.10 Moreover, persistent issues such as gender bias, racial bias, and a general lack of validity evidence continue to mar SLOR effectiveness.11, 12 Some argue that occasionally accurate LORs are more compelling than the substantial evidence of their deep flaws. We believe this mindset is a manifestation of common cognitive biases in human reasoning. These include confirmation bias, framing effect, base-rate neglect, visceral bias, Semmelweis reflex, hindsight bias, and premature closure.13 As readers of LORs, we often believe in our own inherent ability to "read between the lines" and "determine the truth" but we attempt both at our own risk. Despite numerous workshops, papers, and initiatives aimed at improving letter-writing skills, it is unrealistic to expect significant behavioral changes from the vast number of letter writers and readers involved in residency selection. The reluctance to recognize the inherent flaws and unfairness in LORs is likely because doing so would lead to the inevitable conclusion: we should cease writing and relying on them. In reflecting on the evolution of the residency selection process, it is crucial to consider the historical context of assessment in medical education. There was a time when assessment amounted to little more than a cursory checkmark exercise, leaving graduate medical education (GME) with little faith in the integrity or quality assurance of graduates emerging from undergraduate medical education (UME). This lack of trust in assessment data led to a reliance on LORs from trusted and respected colleagues. The ethos of "big name" letter writers became a significant factor, compensating for untrustworthy assessment data. However, with the advent of Competency-Based Medical Education (CBME), the assessment landscape has dramatically transformed. Though UME assessment continues to have shortcomings, CBME has undeniably improved the process, offering a more reliable and equitable evaluation of applicants. CBME marks a significant advancement over LORs, now an obsolete tool with the advent of more sophisticated and trustworthy assessment methods. In clinical practice, we uphold the two principles of using evidence to guide decisions and monitoring biases to minimize harm. This ethos should extend to residency selection. Assessment, often used to safeguard societal interests, must be underpinned by credible evidence for the decisions made. LORs fall short in this regard, lacking the necessary evidence to substantiate their validity. If LORs were a medical procedure, they would not gain approval from regulatory bodies. This disparity highlights the urgent need to reevaluate and align residency application assessments, including LORs, with evidence-based standards. LORs serve as channels for bias and inequity, favor well-connected applicants, focus on selected positive attributes at the expense of true holistic assessment, and have little to no validity at predicting performance in residency. We should stop writing and reading LORs for residency selection, now. LORs are far from perfect but calling to remove them based on inequity, overemphasis on positive attributes, and poor validity evidence is like sweeping one leaf in a forest. In a systematic review, Lipman et al. summarize the literature on metrics used for resident recruitment. Study after study shows that grades, standardized test scores, additional degrees, interviews, Medical Student Performance Evaluations (MSPEs), personal statements, honors/awards, and more are all compromised by bias, as well as the potential for cheating and poor predictive validity.14 Calling for the removal of LORs oversimplifies a complex discussion and unnecessarily singles out one part of the residency selection process. LORs are not the problem, they are a symptom of a broken system. LORs are flawed, but one of the main suggestions to increase diversity, mitigate bias, and increase credible decision making in residency selection is holistic review.15, 16 The very idea of holistic review is that each piece of data is imperfect, and only through the systematic review of all data with a diversity of opinions (i.e., groups) can we begin to minimize bias in residency selection.4 So, will one less piece of data really lead us to less bias, or will it just shift our emphasis onto another piece of biased data? The argument to stop writing and reading LORs on the grounds of bias ignores the fact that all metrics and the entire process of residency selection are compromised by bias. Contrary to what the Point authors state, LORs might be a tool to increase equity. The Point authors have made the argument that LORs compound inequity since all applicants do not have the same access to letter writers. This is an example of equality, and we agree that equality is neither possible nor desirable. Equity seeks to get each person what they need, in hopes of reaching an equal outcome. With this framing in mind, LORs may provide a unique opportunity to promote equity in the residency selection process. LORs are a key tool that faculty can use to advocate for students and lift up those that are marginalized or underrepresented.17, 18 Individual faculty have limited control over grades, awards, or the specific opportunities a student may have. But as educators, we can promote applicants in unique ways, describing their passions, interests, and challenges they have overcome in a way they may not be able to highlight for themselves. Given the number of applicants, most experience the residency selection process as high stakes and impersonal, but a LOR is one of the few opportunities for students to connect with faculty. Removing LORs may improve equality, but it will eliminate one of the ways we can promote equity in a selection process where it is currently lacking.19 Empirical data shows residencies really value LORs and put them to good use. In the biennial survey conducted by the National Resident Matching Program, GME programs consistently cite LORs as a main factor in choosing applicants to interview (80%–90% of the time).20 In fact, LORs ranked higher in importance than standardized tests and personal statements. This is amplified in smaller specialties like Dermatology, Vascular Surgery, and Urology where LORs are almost unanimously perceived as important.20 Similar trends are seen in fellowship applications where LORs have magnified significance.21 The Point authors have stated that one shortcoming of LORs is their failure to comprehensively portray all the struggles, growth, and journey of an individual applicant. Setting aside whether this is even a reasonable expectation for one faculty member to comment upon, they go on to acknowledge that PDs may penalize an applicant when there is language in the LOR about growth or improvement. If the Point authors feel a transparent portrayal of each applicant's journey is lacking, we would point them to the MSPE rather than the LORs.22 Regardless, removing LORs will not make for a more transparent residency selection process. If LORs are useless, then why are so many continuing to use them? In a study of 150 Emergency Medicine PDs, only one advocated for removing LORs from residency selection. When asked for the most important characteristics in choosing who to interview, 139 ranked the LOR first.23 In another study of Anesthesia PDs, most agreed there is value in using LORs to choose who to interview and to look for important keywords and phrases.24 In a national survey of Pediatric PDs, commonly used phrases and keywords in LORs were found to be interpreted in a consistent manner. Importantly, this study found that almost 90% of PDs would consider a weaker candidate more favorably if they had a well-crafted LOR, once again underscoring the opportunity of using LORs to promote applicants.18 Clearly, LORs continue to be used because some key decision makers see potential value. Their depiction as rampant sources of bias should be interpreted with caution. In a 2023 systematic review on residency selection, the authors concluded that the case for bias in LORs is mixed and there is some data to support their predictive value. This led them to conclude that there is more evidence to continue using LORs, while USMLE scores, grades, national school ranking, additional degrees, and receipt of awards should have a limited role.14 Maybe we have chosen the wrong metric to debate. The Point authors would like you to believe that assessment in medical education has evolved with the arrival of CBME, but if it is all tainted by bias, has it really evolved?25 CBME was raised in the Point as a solution, rendering LORs as obsolete since we now have reliable, accurate, and trustworthy assessments. This could not feel further from reality in UME, where CBME is challenging to implement, normative assessments still dominate, and students are oriented primarily toward hiding their weaknesses to try and set themselves apart.26 CBME, as currently implemented, is not the solution. In fact, like holistic review, CBME is built on the idea that utilizing many flawed and imperfect assessments will allow for a more complete picture of a trainee's development.27 Therefore, should not their argument to remove LORs also extend to other forms of assessment that CBME holds with high esteem? If the Point authors want to remove any biased and flawed data from the residency selection process, this slippery slope leads to one solution: residents matched by a lottery. Does that feel extreme? The Netherlands have tried a lottery, stopped it, and are now bringing it back.28 The Point authors dismissed SLORs as a potential improvement, citing reasons such as bias and normative comparisons. However, there might be more to the story. In some contexts, the SLOR was more reliably interpreted and reduced the time that residency programs needed to review LORs.29 In a study of Pediatric applicants collecting validity evidence for SLORs, they found it to be moderately reliable, correlate to admission decisions, and differentiate among applicants even though faculty tended to inflate their ratings on the scale.30 The amount of variance attributed to the applicant in this study (i.e., ability for SLOR to differentiate between applicants) is much higher than most assessments found in medical education. Is not this the kind of validity evidence that the Point authors have called for? In fact, building upon the validity evidence, a decision study showed that by reading four SLORs, one could reliably differentiate between applicants.30 This is critical since residency programs need to complete a final rank of all applicants. Evidence that any piece of information will predict success in residency is lacking, but utilizing SLORs seems to be one way to improve residency selection, mitigate bias, and provide programs with the data they desire.31 Finally, the Point authors have beseeched readers to take an evidence-based approach to guide decisions in residency selection. Yet study after study has shown that no piece of data in the entire residency selection process seems to predict future performance.14, 31 LORs are consistently reported as valuable, with the potential to promote equity, while simultaneously being just as flawed as any other metric in the residency selection process. Removing LORs undermines the very idea of holistic review; we believe pulling one thread (i.e., LORs) from the tapestry (i.e., residency selection) has the potential to do more harm than good. The Counterpoint authors have asked: if the entire residency application process is flawed, why then focus solely on letters? Choosing LORs to remove first is not arbitrary. Rather, it is a strategic move to tackle a classic case of normalized deviance within the residency selection process. Normalized deviance, a phenomenon where deviant practices gradually become accepted as normal within an organization, often leads to a lowered standard of ethics and performance.32 To those entrenched in the system, these practices seem routine and acceptable, while they appear problematic to outsiders. In the case of residency selection, LORs are a prominent example of this deviance. They have become a routine part of the process, despite their inherent flaws and lack of fairness. The first step to ending normalized deviance is to acknowledge and make the problem visible. Removing LORs would do this in an instant. Once this step is taken, the focus can then shift to other aspects of the residency application and selection process. The goal is to create a system that is fair, equitable, accurate, valid, and valuable, rectifying not just a single flawed aspect but challenging a pattern of normalized deviance that has been accepted for too long. This approach is not just about removing a single problematic element; it is about taking a stand for greater integrity and effectiveness. Medical education assessment has evolved. CBME's narrative assessments are shared with a wide array of stakeholders including trainees, competency committees, PDs, and institutions, ensuring transparency and collective scrutiny. In contrast, LORs remain limited in visibility, accessible only to the authors and a select few reviewers. In their present form, they are anathema to CBME: isolated high-stakes assessments based on limited data with low-quality validity evidence. As such, all biases, inaccuracies, and inequities are heightened. Medical educators are still learning how to use CBME to clearly define and assist medical students in meeting criteria essential for graduation. These data, collected from many sources, should primarily facilitate formative assessments and feedback, rather than summative judgments.33 Once graduates meet these criteria, medical schools can confidently assert their readiness for residency, backed by concrete validity evidence. LORs would no longer be needed. We do not need to wait for this. LORs are causing harm now and we should stop writing and reading them for residency selection. The authors declare no conflict of interest.
OBJECTIVE:We proposed adopting billing models for secure messaging (SM) telehealth services that move beyond time-based metrics, focusing on the complexity and clinical expertise involved in patient care. MATERIALS AND METHODS:We trained 8 classification machine learning (ML) models using providers' electronic health record (EHR) audit log data for patient-initiated non-urgent messages. Mixed effect modeling (MEM) analyzed significance. RESULTS:Accuracy and area under the receiver operating characteristics curve scores generally exceeded 0.85, demonstrating robust performance. MEM showed that knowledge domains significantly influenced SM billing, explaining nearly 40% of the variance. DISCUSSION:This study demonstrates that ML models using EHR audit log data can improve and predict billing in SM telehealth services, supporting billing models that reflect clinical complexity and expertise rather than time-based metrics. CONCLUSION:Our research highlights the need for SM billing models beyond time-based metrics, using EHR audit log data to capture the true value of clinical work.
Precision education (PE) leverages longitudinal data and analytics to tailor educational interventions to improve patient, learner, and system-level outcomes. At present, few programs in medical education can accomplish this goal as they must develop new data streams transformed by analytics to drive trainee learning and program improvement. Other professions, such as Major League Baseball (MLB), have already developed extremely sophisticated approaches to gathering large volumes of precise data points to inform assessment of individual performance. In this perspective, the authors argue that medical education-whose entry into precision assessment is fairly nascent-can look to MLB to learn the possibilities and pitfalls of precision assessment strategies. They describe 3 epochs of player assessment in MLB: observation, analytics (sabermetrics), and technology (Statcast). The longest tenured approach, observation, relies on scouting and expert opinion. Sabermetrics brought new approaches to analyzing existing data in a way that better predicted which players would help the team win. Statcast created precise, granular data about highly attributable elements of player performance while helping to account for nonplayer factors that confound assessment such as weather, ballpark dimensions, and the performance of other players. Medical education is progressing through similar epochs marked by workplace-based assessment, learning analytics, and novel measurement technologies. The authors explore how medical education can leverage intersectional concepts of MLB player and medical trainee assessment to inform present and future directions of PE.
Purpose: As competency-based medical education (CBME) continues to advance in undergraduate medical education, students are expected to simultaneously pursue their competency development while also discriminating themselves for residency selection. During the foundational clerkship year, it is important to understand how these seemingly competing goals are navigated. Methods: In this phenomenological qualitative study, the authors describe the experience of 15 clerkship students taking part in a pilot pathway seeking to implement CBME principles. These students experienced the same clerkship curriculum and requirements with additional CBME components such as coaching, an entrustment committee to review their data, a dashboard to visualize their assessment data in real-time, and meeting as a community of practice. Results: Students shared their experiences with growth during the clerkship year. They conveyed the importance of learning from mistakes, but pushing past their discomfort with imperfect performance was a challenge when they feel pressure to perform well for grades. This tension led to significant effort spent on impression management while also trying to identify their role, clarify expectations, and learn to navigate feedback. Conclusions: Tension exists in the clinical environment for clerkship students between an orientation that focuses on maximizing grades versus maximizing growth. The former defined an era of medical education that is fading, while the latter offers a new vision for the future. The threats posed by continuing to grade and rank students seems incompatible with goals of implementing CBME.
Holistic review has become the gold standard for residency selection. As a result, many programs are de-emphasizing standardized exam scores and other normative metrics. However, if standardized exam scores predict passing of an initial certifying exam, this may lead to an increase in board failure rates within specific residency training programs who do not emphasize test scores on entry. Currently, the board pass rates of residency programs from many of the American Board of Medical Subspecialities (ABMS) are publicly reported as a rolling average. In theory, this should create accountability but may also create pressure and distort the way residency program selects applicants. The risk to programs of having a lower board pass rate publicly reported incentivizes programs to focus increasingly on standardized test scores, threatening holistic review. All programs do not recruit students entering residency with an identical chance of passing boards. Therefore, we believe the ABMS member boards should stop publicly reporting raw certifying exam rates above a certain threshold for normative comparison. We strongly encourage the use of learning analytics to create a residency “expected board pass rate” that would be a better metric for program evaluation and accreditation.
We conducted a usability assessment of a resident clinical competency dashboard among eleven Clinical Competency Committee members using the System Usability Scale (SUS) and open-ended questions. Although the average SUS score was 53.47, indicating below-average usability, qualitative feedback highlighted strengths in data integration and interface layout but identified technical issues and requested features for improvement. Findings will inform future development efforts for the dashboard.
Despite the numerous calls for integrating quality improvement and patient safety (QIPS) curricula into health professions education, there are limited examples of effective implementation for early learners. Typically, pre-clinical QIPS experiences involve lectures or lessons that are disconnected from the practice of medicine. Consequently, students often prioritize other content they consider more important. As a result, they may enter clinical settings without essential QIPS skills and struggle to incorporate these concepts into their early professional identity formation. In this paper, we present twelve tips aimed at assisting educators in developing QIPS education early in the curricula of health professions students. These tips address various key issues, including aligning incentives, providing longitudinal experiences, incorporating real-world care outcomes, optimizing learning environments, communicating successes, and continually enhancing education and care delivery processes.
Competency-based medical education (CBME) is an outcomes-based approach to education and assessment that focuses on what competencies trainees need to learn in order to provide effective patient care. Despite this goal of providing quality patient care, trainees rarely receive measures of their clinical performance. This is problematic because defining a trainee's learning progression requires measuring their clinical performance. Traditional clinical performance measures (CPMs) are often met with skepticism from trainees given their poor individual-level attribution. Resident-sensitive quality measures (RSQMs) are attributable to individuals, but lack the expeditiousness needed to deliver timely feedback and can be difficult to automate at scale across programs. In this eye opener, the authors present a conceptual framework for a new type of measure - TRainee Attributable & Automatable Care Evaluations in Real-time (TRACERs) - attuned to both automation and trainee attribution as the next evolutionary step in linking education to patient care. TRACERs have five defining characteristics: meaningful (for patient care and trainees), attributable (sufficiently to the trainee of interest), automatable (minimal human input once fully implemented), scalable (across electronic health records [EHRs] and training environments), and real-time (amenable to formative educational feedback loops). Ideally, TRACERs optimize all five characteristics to the greatest degree possible. TRACERs are uniquely focused on measures of clinical performance that are captured in the EHR, whether routinely collected or generated using sophisticated analytics, and are intended to complement (not replace) other sources of assessment data. TRACERs have the potential to contribute to a national system of high-density, trainee-attributable, patient-centered outcome measures.
BackgroundThe rapid trajectory of artificial intelligence (AI) development and advancement is quickly outpacing society's ability to determine its future role. As AI continues to transform various aspects of our lives, one critical question arises for medical education: what will be the nature of education, teaching, and learning in a future world where the acquisition, retention, and application of knowledge in the traditional sense are fundamentally altered by AI? ObjectiveThe purpose of this perspective is to plan for the intersection of health care and medical education in the future. MethodsWe used GPT-4 and scenario-based strategic planning techniques to craft 4 hypothetical future worlds influenced by AI's integration into health care and medical education. This method, used by organizations such as Shell and the Accreditation Council for Graduate Medical Education, assesses readiness for alternative futures and effectively manages uncertainty, risk, and opportunity. The detailed scenarios provide insights into potential environments the medical profession may face and lay the foundation for hypothesis generation and idea-building regarding responsible AI implementation. ResultsThe following 4 worlds were created using OpenAI’s GPT model: AI Harmony, AI conflict, The world of Ecological Balance, and Existential Risk. Risks include disinformation and misinformation, loss of privacy, widening inequity, erosion of human autonomy, and ethical dilemmas. Benefits involve improved efficiency, personalized interventions, enhanced collaboration, early detection, and accelerated research. ConclusionsTo ensure responsible AI use, the authors suggest focusing on 3 key areas: developing a robust ethical framework, fostering interdisciplinary collaboration, and investing in education and training. A strong ethical framework emphasizes patient safety, privacy, and autonomy while promoting equity and inclusivity. Interdisciplinary collaboration encourages cooperation among various experts in developing and implementing AI technologies, ensuring that they address the complex needs and challenges in health care and medical education. Investing in education and training prepares professionals and trainees with necessary skills and knowledge to effectively use and critically evaluate AI technologies. The integration of AI in health care and medical education presents a critical juncture between transformative advancements and significant risks. By working together to address both immediate and long-term risks and consequences, we can ensure that AI integration leads to a more equitable, sustainable, and prosperous future for both health care and medical education. As we engage with AI technologies, our collective actions will ultimately determine the state of the future of health care and medical education to harness AI's power while ensuring the safety and well-being of humanity.