Objective:. The goal of this study was to assess 2 analytic strategies for comparing hospital outcomes among those with emergency general surgery (EGS) conditions, comparing a conventional risk stratification method with a less utilized, but equally informative strategy. Background:. EGS is a complex set of heterogeneous, time-sensitive conditions that require expeditious treatment. Patients need a mechanism to evaluate how hospitals perform for similar populations treated within the hospital and a reliable metric that benchmarks outcomes across institutions. Methods:. We performed a retrospective cohort study assessing hospital outcomes for EGS Medicare beneficiaries from July 1, 2015, to June 30, 2018. Using direct standardization with balancing weights and indirect standardization with logistic regression, we compare hospital performance on a risk-adjusted composite adverse event rate. Performance based on each standardization modality was correlated using the Spearman rank coefficient. Results:. There were 536,284 patients with a median (interquartile interval) age of 74.2 (72.9, 75.6) years treated at 1866 study hospitals. Direct and indirect standardization showed agreement on 92 low- and 76 high-performing hospitals. Adverse event rates for hospital rankings were strongly correlated between the 2 methods of standardization (0.83, P < 0.001). Rankings based on operative (0.75) and nonoperative (0.77) groups were also highly correlated (all P < 0.001). Conclusions:. Significant variation exists in EGS outcomes. Hospital performance is inconsistent between operative and nonoperative treatment. A small number of hospitals can be distinguished based on risk-adjusted outcomes regardless of analytic technique, suggesting opportunities for optimized care standardization and quality improvement.
OBJECTIVE:Using patient outcomes to provide feedback with benchmarking has been shown to lead to improved performance among attending surgeons. Residents seldom receive information about their patient outcomes during training, which may limit their ability to develop a reflective practice. This study aimed to compare patient outcomes across individual senior residents in complex surgical oncology. We hypothesized that patient outcomes would vary by resident assignment. DESIGN:This was a retrospective cohort study of senior residents who contributed surgical oncology operations (colectomy, hepatectomy, pancreatectomy, or thyroidectomy) to the National Surgical Quality Improvement Program database. Resident pseudo-identifier was the exposure and outcomes included the presence of any adverse event including post-operative complications, length of stay (LOS), operative duration, and 30-day readmissions. Mixed effects regression was used to estimate expected outcome probabilities for each patient. Observed-minus-expected (O-E) rates were then calculated at the resident level. SETTING:Single university-based hospital with a general surgery residency program (2018-2025). PARTICIPANTS:Clinical year 4 and 5 general surgery residents contributing more than 10 surgical oncology operations to the NSQIP registry. RESULTS:Fifty-one senior residents contributed a median of 41 operations per resident (IQR: 25, 65) (n = 2,294 patients). After adjustment for potential confounders, two residents had higher (O-E rates: 18.3%-20.0%) and one resident had a lower adverse event rate than expected (O-E rate: -21.1%). Fifteen residents were outliers in LOS with five residents demonstrating poor performance and 10 residents demonstrating exemplary performance. A total of 15 residents were outliers in operative duration. Two residents were poor performers, and one resident was an exemplary performer across these three metrics. No performance outliers in 30-day readmission were identified. CONCLUSIONS:Risk-adjusted patient outcomes can be used to identify performance outliers at the individual resident level. These data provide residents with valuable feedback necessary to develop practice-based learning and improvement skills.
Background Medullary thyroid cancer (MTC) is rare and requires specialty knowledge. Medicare Advantage (MA) is associated with reduced access to top-rated cancer hospitals and worse outcomes for complex cancer surgery than Traditional Medicare (TM). Whether this applies to MTC is unknown. We described access to accredited cancer centers and compared outcomes between TM and MA beneficiaries with MTC. Methods This retrospective cohort study of Medicare beneficiaries who underwent surgery for a new diagnosis of MTC in the SEER-Medicare database (2018-2021) examined the association between Medicare type (TM or MA) and utilization of a Commission on Cancer (CoC)-accredited center with stratification by state. Logistic regression models evaluated risk-adjusted CoC hospital utilization. Results We identified 381patients of which 231 (60.6%) were enrolled in TM and 150 (39.4%) in MA. In the overall analysis, there were no significant differences in accredited cancer center utilization by Medicare type (TM: 83.6% vs MA: 78.0%; p=0.22). When stratified by state, MA beneficiaries were less often treated at accredited cancer centers in California and New York. Postoperative complications, delayed reoperation and disease-specific survival did not differ significantly by Medicare type. Conclusion Medicare type is associated with state-level variation in the utilization of accredited cancer centers for MTC surgical care, Patients with MTC should consider access issues in their local context when selecting Medicare plan types. Outcomes studies of health care utilization by Medicare plan type must consider the geographic setting as a component of their analysis to responsibly report results.
PURPOSE:Extramural funding is critical to career success in academic surgery and most federal and societal funding opportunities for trainees require application submission early in residency. Single institution studies have demonstrated that surgical resident self-efficacy in grantsmanship is limited, but there are no national data on trainee experience and comfort with grant writing. The objective of this study was to perform a multi-institutional needs assessment to explore surgical resident grantsmanship self-efficacy. METHODS:This was a multi-institutional survey of all general surgery residents enrolled in the Association for Surgical Education list-serv. A previously validated, 19-item grantsmanship self-efficacy inventory was distributed electronically. The survey included 3 domains of grantsmanship: project conceptualization, study design, and knowledge of the funding process. Each item was scored on a scale from 0 (no confidence) to 10 (complete confidence). Junior and senior residents as well as residents with and without prior grant writing experience were compared using Student t-tests, 1 way Analysis of Variance (ANOVA), and Pearson's Chi-square tests. RESULTS:The 77 resident respondents represented all training levels and multiple program types. The mean resident grantsmanship self-efficacy score was 6.0 (standard deviation-SD: 2.3, range: 0 to 8.9) out of 10. Senior residents (PGY3-5) had significantly higher mean scores than junior residents (PGY 1-3) (5.5 vs. 3.5, p = 0.002). Grantsmanship self-efficacy was also higher among residents who completed research time, submitted or funded a grant, and/or attended formal grantsmanship training compared to those that did not (all p < 0.001). A significantly higher proportion of residents who received formal grant training had a grant funded compared to those who had no formal training (43.2% vs. 7.5%, p < 0.001). CONCLUSIONS:General surgery resident grantsmanship self-efficacy is low with some evidence of improvement later in training, after key funding opportunities have passed. Residents with exposure to both formal and informal grant writing training have higher reported comfort with the process. In an increasingly challenging funding environment, formal curricula that support mastery of the grant writing and funding processes are needed early in residency.
This cross-sectional study examines the association between surgery residents’ sense of belonging and their performance on the American Board of Surgery In-Service Training Examination.
INTRODUCTION:Malignant obstructions are associated with morbidity, mortality, and impaired quality of life with limited population-based data on practice patterns. We characterized the association between prior cancer-related hospitalizations and receipt of operative treatment for obstruction. METHODS:We performed a retrospective cohort study of patients hospitalized with malignant gastrointestinal obstructions from colorectal, gynecologic (GYN), and hepato-pancreatico-biliary (HPB) cancers in 12 states using the Healthcare Cost and Utilization Project State Inpatient Database, 2016-2020. The primary exposure was the number of cancer-related hospitalizations in the year prior to the first hospitalization for an obstruction. Risk-adjusted use of operative treatment during the index hospitalization, enrollment in hospice, length of stay, and inpatient mortality were examined using regression. RESULTS:Of 8665 patients, the median age was 67 years [Interquartile Range: 58,76]. There were 3789 patients with colorectal (43.7%), 2593 patients with GYN (29.9%), and 2337 patients (26.9%) with HPB cancers. The operative rate was 24.1%. Patients with prior hospitalizations were less likely to undergo operative treatment than those who were not previously hospitalized (operative rate difference: -7.8%, p < 0.001). Operative treatment was associated with longer median length of stay ( + 8.6 days, p < 0.001), lower rates of hospice enrollment (-3.0%, p = 0.002) and lower rates of inpatient mortality (operative rate difference -2.3%, p < 0.001). The time between readmissions also decreased with each recurrent obstruction in colorectal and GYN cancers. CONCLUSION:A greater number of previous cancer-related hospitalizations was associated with lower rates of operative treatment during the first hospitalization for a malignant obstruction. These data provide insight into the treatment trends for patients with malignant obstructions and may be helpful when counseling patients about the natural course of their disease.
ABSTRACT Background and Objectives Serious mental illness and dementia‐related disorders are associated with worse postoperative outcomes. We evaluated the impact of pre‐existing neuropsychiatric disorders on post‐parathyroidectomy healthcare utilization. Methods Adult patients who underwent parathyroidectomy for primary hyperparathyroidism were identified in the Healthcare Cost and Utilization Project State Inpatient and Ambulatory Surgery Databases (2016–2021). Neuropsychiatric disorders were classified. The primary outcome was 30‐day readmission. Secondary outcomes were length of stay and costs (2021USD). Balancing weights were used for covariate adjustment. Subgroup analyses were performed in serious mental illness and dementia‐related disorder groups. Results Of 10,733 parathyroidectomy patients, 1,249 (11.6%) had neuropsychiatric disorders: 60.5% mood, 58.8% anxiety, 5.1% dementia‐related, and 1.8% schizophrenia‐type disorders. Patients with neuropsychiatric disorders were younger and more likely to undergo emergent/urgent parathyroidectomy ( p < 0.001). There was no difference in adjusted odds of readmission between patients with and without neuropsychiatric disorders (0.91[95%CI 0.55, 1.49], p = 0.710). Adjusted length of stay was 0.20 days longer among patients with neuropsychiatric disorders ([95%CI: + 0.11, + 0.30]; p < 0.001). Adjusted encounter costs were similar between groups (median[IQI]: neuropsychiatric disorder $6767[5290,8589] vs. no neuropsychiatric disorder $6511[5259, 8550]; p = 0.080). Subgroup analyses showed similar results. Conclusions Patients with neuropsychiatric disorders have similar parathyroidectomy readmission and cost. They have slightly longer length of stay that may ameliorate the risk of readmission.
Importance:Surgeon compensation models influence physician productivity, care quality, and engagement in nonclinical activities. Information on compensation plans across surgical specialties and settings is often difficult to obtain. Objective:To describe surgeon compensation models in the US, differentiate models by practice setting, and evaluate their association with clinical productivity and nonclinical contributions. Evidence Review:A systematic review was conducted following PRISMA guidelines. PubMed and Embase were queried for articles published between January 1, 2014, and August 20, 2024, reporting on surgeon compensation models. Two reviewers independently screened and abstracted data on study characteristics, practice settings, compensation structures, and outcomes. Risk of bias was assessed, and studies were synthesized by compensation model. Qualitative analysis was performed to determine themes across the literature. Findings:Of 3268 screened records, 39 studies met inclusion criteria, encompassing 13 surgical specialties. Articles reported on compensation models, including salary (n = 8), work relative value unit (wRVU)-based (n = 8), hybrid (n = 7), fee-for-service (n = 5), and value-based models (n = 3). Hybrid models include a blend of financial incentives and base salary. Salary-based models provided financial stability, promoted team-based care and were associated with lower clinical volume. wRVU and fee-for-service models strongly incentivized productivity and often failed to account for case complexity, patient outcomes, or nonclinical work. Hybrid models offered flexibility by combining base salaries with incentives for volume, quality, and academic contributions despite greater administrative complexity. Value-based models, rarely used, may have unintentional consequences. Across models, there was wide variability in financial compensation for teaching, research, and administrative duties. Conclusions and Relevance:Surgeon compensation models in the US remain heterogeneous. While productivity-based systems dominate, emerging hybrid and value-based approaches aim to support broader professional obligations. Transparent, adaptable frameworks that balance clinical output with quality and nonclinical contributions are needed to sustain surgeon engagement and align compensation across the totality of surgical practice.
The American College of Surgeons (ACS) developed the Quality In-Training Initiative (QITI) to link surgeon trainee participation with patient outcomes within ACS National Surgical Quality Improvement Program (NSQIP). As surgical education shifts toward competency-based and outcomes-informed training, QITI may offer a platform not only for data collection, but also for resident learning and preparation for practice. We evaluated QITI as a system for collecting resident-linked clinical outcomes, supporting meaningful clinical education, and preparing trainees for data-driven surgical practice. This is a retrospective study of NSQIP cases with completed QITI fields from July 2013 through July 2025. Descriptive analyses evaluated case volume, participation across institutions, operative case mix, and postoperative outcomes by PGY level. Multivariable regression models were used to assess the association between PGY level and postoperative outcomes after adjustment for standard NSQIP preoperative risk factors. From 190 participating institutions, 1,412,068 cases met inclusion criteria, including 1,098,264 PGY1–5 cases. While operative case mix increased in complexity across PGY levels, some procedures such as laparoscopic appendectomy and cholecystectomy remained common at all levels. Unadjusted complication rates increased stepwise with PGY level (all p < 0.001). After risk adjustment, most differences attenuated, although cases involving PGY5 residents remained significantly associated with higher odds of select outcomes, including intubation, prolonged ventilation, renal complications, cardiac complications, readmission, morbidity, and mortality. Sensitivity analysis performed for years 2022–2024 showed fewer significant associations. These findings suggest that QITI is a feasible and sustainable system for collecting resident-linked clinical outcomes at national scale. The data provide meaningful educational value by characterizing progression in operative complexity and linking cases with outcomes. In addition, risk-adjusted comparisons and benchmarking reflect the types of outcome reports surgeons encounter in practice, creating an opportunity to expose residents to the types of data-driven quality reports they may encounter in independent practice. QITI may therefore serve as a valuable platform for competency-based surgical education and future trainee feedback systems.
OBJECTIVES:Older adults with abdominal pain present diagnostic uncertainty due to less informative histories/exams, broader etiologies, and higher morbidity. Whether ED imaging decisions are calibrated to this risk is unclear. The objective of this study was to compare age-stratified clinical features, CT utilization, and CT diagnostic yield, and to assess how history/physical and clinician pretest suspicion relate to adverse outcomes. METHODS:This was a retrospective cohort analysis of data from a prospective cohort collected from March 2016-January 2017 at a single community teaching hospital emergency department in southwest Baltimore. We analyzed 1169 visits of adults presenting with nontraumatic abdominal pain including 229 (19.6%) aged ≥ 60 years. Patients < 18 years were excluded. Age groups were 18-39, 40-59, ≥ 60 years. Outcomes were CT ordering, acute actionable CT findings, admission, surgery, and a composite of adverse outcomes (any actionable CT finding, admission, surgery, or Emergency General Surgical diagnosis). History and physical examination operating characteristics (e.g., sensitivity/specificity of tenderness, rebound) were also calculated. RESULTS:Of 1169 visits, 19.6% were aged ≥ 60 years. CT ordering increased with age (41.7%, 66.2%, 70.7% for 18-39, 40-59, ≥ 60; p < 0.001), as did CT yield (18.4%, 31.2%, 37.7%; p < 0.001). Admissions (12.1%, 28.0%, 37.6%) and surgeries (4.6%, 9.0%, 10.6%) also rose with age. Clinician pretest suspicion was similar across age groups. Abdominal tenderness was less sensitive for adverse outcomes in older adults (sensitivity 0.58 in ≥ 60 vs. 0.73 in 18-39 and 0.73 in 40-59), while rebound tenderness was highly specific across ages (specificity 0.98, 0.96, 0.98). The number of potential diagnoses to consider rose with age. CONCLUSION:In this cohort, CT use and positivity increased with age and key exam findings (e.g., tenderness) being less informative in older adults, despite similar reported clinician pretest suspicion. These results support age-aware imaging decisions and motivate reframing ED abdominal pain as a geriatric-specific chief complaint.
Objective:. To develop a machine learning model that predicts surgical case length and benchmark its performance against an embedded electronic health record (EHR) model. Background:. Surgical care accounts for one-third of U.S. healthcare expenditure. Current case length prediction models are generally overly simplistic and inaccurate or too specialized to have a broad impact, contributing to operating room (OR) inefficiency and dissatisfaction for patients and providers. Methods:. Retrospective analysis of 55,495 surgical cases performed by 299 surgeons between January 2022 and April 2024 at a metropolitan, quaternary care hospital. The dataset was split temporally for training (46,767 cases) and holdout validation (8728 cases). Three separate machine learning models predicted preprocedure, operative, and postprocedure times using patient and provider characteristics, operation details, and hospital features available at least 1 day before surgery. Approximately 22% of cases lacked historical time averages and relied on procedural time heuristics. Results:. The machine learning model significantly outperformed the embedded EHR model, achieving lower root mean squared error (61.0 vs 91.0 minutes; P < 0.01), lower mean average error (39.6 vs 51.8 minutes; P < 0.01), and higher R2 (0.78 vs 0.50; P < 0.01). The model predicted 213 more cases within ±30 minutes of actual duration. In cases without historical time averages, the model increased cases within ±30 minutes of actual duration (35% vs 29%; P < 0.01). Conclusions:. A machine learning model leveraging comprehensive preoperative data significantly improved surgical case length prediction compared to an embedded EHR model. Future implementation has the potential to improve OR efficiency and patient and provider satisfaction.
BACKGROUND:Post-colectomy adverse events occur in up to one third of patients. We used machine learning (ML) models to predict complications and inform optimal surgeon-hospital assignment. STUDY DESIGN:Adults ≥18 undergoing colon or rectal resection in an academic system from 2018 to 2024 were included. Multiple ML algorithms trained on patient and provider features predicted postoperative complications. Model performance was compared by the scaled Brier score. Patients were simulated to their optimal surgeon-hospital dyad to estimate risk reduction. RESULTS:Among 4689 cases, 1562 (33.3%) experienced ≥1 adverse event. LightGBM outperformed alternate ML models (p < 0.05). LightGBM achieved a scaled Brier score of 0.11 (95% CI 0.08-0.16). Top predictors included pre-operative diagnosis, surgeon identifier, and ostomy. Simulations suggested a 4.6% (CI 4.2-4.9%) net absolute complication risk reduction and 6.9% (CI 6.5-7.3%) reduction with reassignment to an optimal surgeon-hospital dyad. CONCLUSIONS:ML models could enable a proactive referral strategy to promote optimal surgical outcomes.
Hernia repairs have differences in outcomes based on hernia type. Information regarding hernia burden in the emergency setting is lacking. Among older adults, who have the greatest prevalence of hernia and the need for emergent repair, little data on the impact of multimorbidity on outcomes exist. We aim to define the burden of emergency hernia on hospitals and to compare outcomes of older adults with and without multimorbidity. This was a nationwide retrospective cohort study of Medicare beneficiaries admitted emergently from 2015–2018 with a principal diagnosis of an umbilical, ventral, parastomal, femoral, or inguinal hernia. The primary outcome was all-cause inpatient mortality. Multivariable logistic regression was performed. Among 47,687 hospitalized patients, there were 4,612 (9.7
Team dynamics influence team performance and patient outcomes in surgery, yet data on resident-led teams are scarce. This study aimed to compare patient outcomes across resident-led teams in complex surgical oncology. We hypothesized that patient outcomes would vary by team assignment. This was a retrospective cohort study of resident-led teams who contributed more than 10 surgical oncology operations (colectomy, hepatectomy, pancreatectomy, or thyroidectomy) to the National Surgical Quality Improvement Project registry at a single university-based hospital (2018–2025). The primary outcome was presence of any adverse event, including mortality and postoperative complications. Length of stay and 30-day readmissions were also examined. Mixed-effects regression estimated expected outcome probabilities for each patient. For each team, observed-minus-expected (O-E) outcome rates were calculated to assess performance. In total, 145 teams cared for a median of 22 patients (interquartile interval 16– 27; n = 2919). Five teams demonstrated poor performance based on risk-adjusted adverse event rates (O-E rates: 6.9
Introduction: Analyzing general surgeons' operative case mix can provide an update on contemporary practice patterns and inform pragmatic residency training. Methods: We performed a retrospective cohort study of general surgeons in Florida, Iowa, and Maryland, 2016-2020. Cases were identified using billing codes. The Cochran-Armitage test of trends was used to evaluate the proportion of practice devoted to specific case types and operative setting over time. Results: General surgeons (n = 1300) performed 1,287,745 cases. The mean (+/- SD) annual volume per surgeon for all procedures was 356 (+/- 250), with 198 (+/- 152) general surgery operations, 57 (+/- 142) endoscopic procedures, and 101 (+/- 109) other cases. On average, surgeons operated on 7.1 (+/- 2.6) different organ systems. Trends toward a lower proportion of general surgery operations, and a greater proportion of subspecialty procedures and surgery in the outpatient setting over time were demonstrated (p < 0.001). Conclusion: The practice pattern of the general surgeon continues to be heterogeneous, reflecting the persistent need for a broad training paradigm that permits specialization.
OBJECTIVE:To review the current state of research training during surgical residency and make recommendations commensurate with current surgical training and academic environment. BACKGROUND:Research training has been a mainstay of academic surgical programs, yet the scientific disciplines have evolved significantly from the traditional years of bench research. It is time to reconsider how research training should prepare surgeons for future academic practice and ensure the foundational knowledge of research evidence. METHODS:As part of the Blue Ribbon Committee II, a research subcommittee was tasked to make recommendations on research training during surgical residency. Our 8-member panel brought diverse perspectives on the roles and goals of research training. We also sought input from a convenience sample of current and recent surgical residents on the impact of research training during their residency. RESULTS:We identified a lack of a common framework and foundational research training for all surgical residents. Participation in dedicated years of scholarly activity helped trainees meet several professional and personal goals. The lack of an integrated, dedicated research track may dissuade some medical school graduates from pursuing surgery. CONCLUSIONS:We recommend incorporating a minimum standard for all trainees and flexibility in dedicated scholarly training to meet the needs of future academic surgeons.
ABSTRACTBackground and MethodsColorectal cancer (CRC) treatment can influence health‐related quality of life (HRQOL). This study examined HRQOL among older adults undergoing CRC treatment, and the conditional effects of race, ethnicity, and primary language. We conducted a retrospective cohort study of Medicare Advantage enrollees ≥ 65 years old who completed the Medicare Health Outcomes Survey (MHOS) (2016−2020). The exposure group answered “Yes” to the current CRC treatment and the control group answered “No.” The primary outcomes were physical component summary (PCS) and mental component summary (MCS) scores. Conditional effects by race and ethnicity were analyzed using interaction terms.ResultsAmong 184 486 adults, 676 (0.4%) reported current CRC treatment. Those receiving treatment had significantly lower PCS scores (β coefficient −1.98, p < 0.001) and lower MCS scores (β coefficient −0.81, p = 0.018), compared to nontreatment. In the treatment group, Hispanic respondents and Spanish speakers had higher PCS scores (β coefficient 1.96, p = 0.019 and 3.19, p = 0.023, respectively), and respondents identifying as American Indian or Alaska Native had higher MCS scores (β coefficient 8.72, p = 0.016).ConclusionIndividuals receiving CRC treatment exhibit worse HRQOL. Outcomes differed by race and ethnicity. This study suggests the need to invest in targeted interventions to improve overall HRQOL during treatment for CRC.