BACKGROUND AND OBJECTIVE:Low volume has been recognized as a problem when benchmarking hospitals due to outcome rate instability. We asked if low-volume hospital outcomes, using matching to control for many clinical and sociodemographic characteristics, would expose quality problems not observed with CMS methods. RESEARCH DESIGN:Matched cohort study. Grades derive from mortality differences between all patients at the low-volume hospital and their matched controls. SUBJECTS:Medicare patients admitted with Acute Myocardial Infarction, Heart Failure and Pneumonia in 78 low-volume Pennsylvania acute care hospitals (combined condition volume=75≤N≤750 for the 3 y, 2017-2019), using Medicare's Virtual Research Data Center. MEASURES:Thirty-day mortality. RESULTS:Using matching, 10 of 78 reportable low-volume hospitals had significantly higher mortality versus matched typical controls and 16 low-volume hospitals displayed significantly higher mortality versus well-resourced controls. In contrast, Medicare reported that only 3 of these same 78 hospitals had significantly higher mortality than "the national rate" on AMI, HF, or pneumonia. CONCLUSIONS:We find that some low-volume hospitals performed well. Other low-volume hospitals had significantly worse outcomes than both well-resourced and typical hospitals; and some displayed significantly worse mortality compared with well-resourced controls but did not reach significant differences from typical controls. In short, performing "no different from the national rate," as is almost always reported for low-volume hospitals when using CMS methods, does not imply a low-volume hospital has acceptable outcomes. Reports based on matching can expose low-volume hospital quality problems not apparent using standard methods. Low-volume hospitals have more quality problems than generally reported.
Importance Understanding the association of postoperative delirium with adverse outcomes and the hospital-level variation of postoperative delirium is important for efforts to improve perioperative brain health. Objective To examine (1) the association of postoperative delirium with 30-day mortality and complications and (2) hospital-level variation in postoperative delirium. Design, Setting, and Participants This retrospective cohort study examined hospitalizations among patients aged 65 years and older who underwent noncardiac surgery in US hospitals between January 1, 2017, and December 31, 2020. Data were analyzed between August 28, 2024, and April 10, 2025. Exposure Postoperative delirium. Main Outcomes and Measures The association of the composite of death and major complications with postoperative delirium was examined using multivariable logistic regression. Variability in the hospital incidence of postoperative delirium was evaluated using multilevel logistic regression analysis. Results Among 5 530 054 inpatient admissions for major noncardiac surgery in 3169 hospitals, the mean (SD) patient age was 74.7 (7.0) years, and 3 161 054 admissions (57.2%) were of female patients. The incidence of postoperative delirium was 3.6% (197 921 admissions). Compared with patients without postoperative delirium, patients with postoperative delirium were more likely to experience death or major complications (adjusted OR [aOR], 3.47; 95% CI, 3.41-3.53; P < .001), 30-day mortality (aOR, 2.77; 95% CI, 2.71-2.83; P < .001), and nonhome discharges (aOR, 3.96; 95% CI, 3.88-4.04; P < .001). Controlling for patient characteristics, the odds of postoperative delirium were higher for patients undergoing surgery in hospitals with a higher rate of postoperative delirium compared with hospitals with lower rates of postoperative delirium (median OR, 1.53; 95% CI, 1.50-1.56). Conclusions and Relevance In this national retrospective cohort study of more than 5.5 million hospitalizations, older individuals undergoing major noncardiac surgery who experienced postoperative delirium had 3.5-fold higher odds of death or major complications, 2.8-fold higher odds of death, and 4.0-fold higher odds of nonhome discharge. There was substantial variation in the hospital rate of postoperative delirium after accounting for patient risk, which suggests that this complication may be an appropriate target for hospital efforts to improve perioperative brain health, provided that delirium screening and coding accuracy are improved.
Objective:. Develop a new hospital surgery report card for use in performance improvement. Background:. When evaluating quality, a surgical program is aided by benchmark comparisons with outcomes achieved at other hospitals. To be credible, benchmarking should be based on the same surgical procedures and patient risk, despite there being many types of patients and procedures. Methods:. Using Medicare patients undergoing general, orthopedic, or vascular surgery, each patient in a hospital is closely matched to 10 control patients from typical hospitals and to 10 control patients from well-resourced hospitals throughout the United States. Patients were matched on 200 characteristics, including procedure, comorbidities, socio-demographics, and the presence of multimorbidity. Hospitals were graded based on the differences in outcomes between matched sets of patients. As an illustration, we examine the 20 highest volume hospitals in Pennsylvania and provide detailed report cards on 2 example hospitals. Results:. The hospitals studied differed in quality and grades, with better outcomes than matched controls for Hospital A and significantly worse outcomes than controls for Hospital B, depending on the type of surgery and patient. For the 20 largest hospitals in Pennsylvania, 5 had significantly elevated mortality, and 2 had significantly lower mortality than matched controls. Conclusions:. Surgical programs benefit from knowing how their outcomes compare with those of other hospitals, both their overall outcomes and their outcomes for subsets of patients, such as patients with or without multimorbidity. Detailed reports based on matching can help identify meaningful deficiencies and strengths in programs concerning specific surgeries and patient types.
BACKGROUND AND OBJECTIVES:To improve upon existing hospital grading systems, we developed a new report card based on multivariate matching. RESEARCH DESIGN:Matched cohorts. For each focal hospital patient, we match 10 control patients treated at "well-resourced" hospitals with excellent hospital characteristics from across the nation, and 10 control patients treated at "typical" hospitals, on over 300 patient characteristics from Medicare Claims. Grades were based on outcome differences between patients at the focal hospital and their matched controls. We also create an "Analogous" match that is comprised of multiple control patients matched to each focal hospital patient with similar patient characteristics who were treated at hospitals with similar characteristics to the focal hospital, answering the question, "How would patients who looked like my patients and who were treated at hospitals like my hospital fare, compared to how my patients fared." We also report outcomes by multimorbidity status. SUBJECTS:Medicare admissions from 2017 to 2019 for heart attack, heart failure and pneumonia. To illustrate our methods, we report on 4 hospitals in the same region: a well-known "Flagship" teaching Hospital, an Affiliated Hospital within the same flagship system, a Poor-Performing Hospital that is not part of the flagship system, and a Small Hospital with unstable estimates. MEASURES:Thirty-day mortality and revisit rates. RESULTS:Report cards for each example hospital. CONCLUSIONS:Matched report cards allow users to better benchmark hospitals and see those types of patients where a specific hospital is performing poorly compared to other hospitals treating very similar patients.
Importance:Delaying elective noncardiac surgery after a recent acute myocardial infarction is associated with better outcomes, but current American Heart Association recommendations are based on data that are more than 20 years old. Objective:To examine the association between the time since a non-ST-segment elevation myocardial infarction (NSTEMI) and the risk of postoperative major adverse cardiovascular and cerebrovascular events (MACCE). Design, Setting, and Participants:This cross-sectional study examined Medicare claims data between 2015 and 2020 for patients 67 years or older who had major noncardiac surgery. Data were analyzed from September 21, 2023, to February 1, 2024. Exposure:Time elapsed between a prior NSTEMI and surgery. Main Outcomes and Measures:MACCE (30-day mortality, in-hospital myocardial infarction, heart failure, or stroke) and all-cause 30-day mortality. Multivariable logistic regression was used to estimate the association between outcomes and time since a prior NSTEMI. Results:The sample included 5 227 473 surgeries. The mean (SD) age was 75.7 (6.6) years; 2 981 239 (57.0%) were female, and 2 246 234 (43%) were male. There were 42 278 patients (0.81%) with a previous NSTEMI. Compared with patients without a prior NSTEMI, patients with an NSTEMI within 30 days of elective surgery had higher odds of MACCE, regardless of whether they had undergone coronary revascularization (adjusted odds ratio [aOR], 2.15; 95% CI, 1.09-4.23; P = .03) or not (aOR, 2.04; 95% CI, 1.31-3.16; P = .001). The odds of postoperative MACCE leveled off after 30 days in patients who had undergone any coronary revascularization procedure (and after 90 days in patients with drug-eluting stents) and then increased after 180 days (any revascularization at 181-365 days: aOR, 1.46; 95% CI, 1.25-1.71; P < .001; patients with drug-eluting stents at 181-365 days: aOR, 1.73; 95% CI, 1.42-2.12; P < .001). The odds of MACCE did not level off for patients who did not have revascularization. Findings for all-cause 30-day mortality were similar to those for MACCE, except that the odds of mortality in patients with previous NSTEMI who had revascularization leveled off after 60 days in elective surgeries and 90 days for nonelective surgeries (elective 30-day: aOR, 2.88; 95% CI, 1.30-6.36; P = .009; elective 61- to 90-day: aOR, 1.03; 95% CI, 0.57-1.86; P = .92; nonelective 30-day: aOR, 1.91; 95% CI, 1.52-2.40; P < .001; nonelective 91- to 120-day: aOR, 1.00; 95% CI, 0.73-1.37; P = .99). Conclusions and Relevance:This study found that among older patients undergoing noncardiac surgery who had revascularization, the odds of postoperative MACCE and mortality leveled off between 30 and 90 days and then increased after 180 days. The odds did not level off for patients who did not have revascularization. Delaying elective noncardiac surgery to occur between 90 and 180 days after an NSTEMI may be reasonable for patients who have had revascularization.
BACKGROUND:The use of real-world data (RWD) in artificial intelligence (AI) applications for healthcare offers unique opportunities but also poses complex challenges related to interpretability, transparency, safety, efficacy, bias, equity, privacy, ethics, accountability, and stakeholder engagement. METHODS:A multi-stakeholder expert panel comprising healthcare professionals, AI developers, policymakers, and other stakeholders was assembled. Their task was to identify critical issues and formulate consensus recommendations, focusing on the responsible use of RWD in healthcare AI. The panel's work involved an in-person conference and workshop and extensive deliberations over several months. RESULTS:The panel's findings revealed several critical challenges, including the necessity for data literacy and documentation, the identification and mitigation of bias, privacy and ethics considerations, and the absence of an accountability structure for stakeholder management. To address these, the panel proposed a series of recommendations, such as the adoption of metadata standards for RWD sources, the development of transparency frameworks and instructional labels likened to "nutrition labels" for AI applications, the provision of cross-disciplinary training materials, the implementation of bias detection and mitigation strategies, and the establishment of ongoing monitoring and update processes. CONCLUSION:Guidelines and resources focused on the responsible use of RWD in healthcare AI are essential for developing safe, effective, equitable, and trustworthy applications. The proposed recommendations provide a foundation for a comprehensive framework addressing the entire lifecycle of healthcare AI, emphasizing the importance of documentation, training, transparency, accountability, and multi-stakeholder engagement.
ObjectivesThe extent to which care quality influenced outcomes for patients hospitalised with COVID-19 is unknown. Our objective was to determine if prepandemic hospital quality is associated with mortality among Medicare patients hospitalised with COVID-19.DesignThis is a retrospective observational study. We calculated hospital-level risk-standardised in-hospital and 30-day mortality rates (risk-standardised mortality rates, RSMRs) for patients hospitalised with COVID-19, and correlation coefficients between RSMRs and pre-COVID-19 hospital quality, overall and stratified by hospital characteristics.SettingShort-term acute care hospitals and critical access hospitals in the USA.ParticipantsHospitalised Medicare beneficiaries (Fee-For-Service and Medicare Advantage) age 65 and older hospitalised with COVID-19, discharged between 1 April 2020 and 30 September 2021.Intervention/exposurePre-COVID-19 hospital quality.OutcomesRisk-standardised COVID-19 in-hospital and 30-day mortality rates (RSMRs).ResultsIn-hospital (n=4256) RSMRs for Medicare patients hospitalised with COVID-19 (April 2020–September 2021) ranged from 4.5% to 59.9% (median 18.2%; IQR 14.7%–23.7%); 30-day RSMRs ranged from 12.9% to 56.2% (IQR 24.6%–30.6%). COVID-19 RSMRs were negatively correlated with star rating summary scores (in-hospital correlation coefficient −0.41, p<0.0001; 30 days −0.38, p<0.0001). Correlations with in-hospital RSMRs were strongest for patient experience (−0.39, p<0.0001) and timely and effective care (−0.30, p<0.0001) group scores; 30-day RSMRs were strongest for patient experience (−0.34, p<0.0001) and mortality (−0.33, p<0.0001) groups. Patients admitted to 1-star hospitals had higher odds of mortality (in-hospital OR 1.87, 95% CI 1.83 to 1.91; 30-day OR 1.46, 95% CI 1.43 to 1.48) compared with 5-star hospitals. If all hospitals performed like an average 5-star hospital, we estimate 38 000 fewer COVID-19-related deaths would have occurred between April 2020 and September 2021.ConclusionsHospitals with better prepandemic quality may have care structures and processes that allowed for better care delivery and outcomes during the COVID-19 pandemic. Understanding the relationship between pre-COVID-19 hospital quality and COVID-19 outcomes will allow policy-makers and hospitals better prepare for future public health emergencies.
BACKGROUND:Integrating artificial intelligence (AI) in healthcare settings has the potential to benefit clinical decision-making. Addressing challenges such as ensuring trustworthiness, mitigating bias, and maintaining safety is paramount. The lack of established methodologies for pre- and post-deployment evaluation of AI tools regarding crucial attributes such as transparency, performance monitoring, and adverse event reporting makes this situation challenging. OBJECTIVES:This paper aims to make practical suggestions for creating methods, rules, and guidelines to ensure that the development, testing, supervision, and use of AI in clinical decision support (CDS) systems are done well and safely for patients. MATERIALS AND METHODS:In May 2023, the Division of Clinical Informatics at Beth Israel Deaconess Medical Center and the American Medical Informatics Association co-sponsored a working group on AI in healthcare. In August 2023, there were 4 webinars on AI topics and a 2-day workshop in September 2023 for consensus-building. The event included over 200 industry stakeholders, including clinicians, software developers, academics, ethicists, attorneys, government policy experts, scientists, and patients. The goal was to identify challenges associated with the trusted use of AI-enabled CDS in medical practice. Key issues were identified, and solutions were proposed through qualitative analysis and a 4-month iterative consensus process. RESULTS:Our work culminated in several key recommendations: (1) building safe and trustworthy systems; (2) developing validation, verification, and certification processes for AI-CDS systems; (3) providing a means of safety monitoring and reporting at the national level; and (4) ensuring that appropriate documentation and end-user training are provided. DISCUSSION:AI-enabled Clinical Decision Support (AI-CDS) systems promise to revolutionize healthcare decision-making, necessitating a comprehensive framework for their development, implementation, and regulation that emphasizes trustworthiness, transparency, and safety. This framework encompasses various aspects including model training, explainability, validation, certification, monitoring, and continuous evaluation, while also addressing challenges such as data privacy, fairness, and the need for regulatory oversight to ensure responsible integration of AI into clinical workflow. CONCLUSIONS:Achieving responsible AI-CDS systems requires a collective effort from many healthcare stakeholders. This involves implementing robust safety, monitoring, and transparency measures while fostering innovation. Future steps include testing and piloting proposed trust mechanisms, such as safety reporting protocols, and establishing best practice guidelines.