Background Artificial Intelligence (AI) in the selection of residency program applicants is a new tool that is gaining traction, with the aim of screening high numbers of applicants while introducing objectivity and mitigating bias in a traditionally subjective process. This study aims to compare applicants screened by an AI software to a single Program Director (PD) for interview selection. Methods A single PD at an ACGME-accredited, academic general surgery program screened applicants. A parallel screen by AI software, programmed by the same PD, was conducted on the same pool of applicants. Weighted preferences were assigned in the following order: personal statement, research, medical school rankings, letters of recommendation, personal qualities, board scores, graduate degree, geographic preference, past experiences, program signal, honor society membership, and multilingualism. Statistical analyses were conducted by chi-square, ANOVA, and independent two-sided t-tests. Results Out of 1235 applications, 144 applications were PD-selected and 150 AI-selected (294 top applications). Twenty applications (7.3%) were both PD and AI selected for a total analysis cohort of 274 prospective residents. We performed two analyses: 1) PD-selected vs. AI-selected vs. Both and 2) PD-selected vs. AI-selected with the overlapping applicants censored. For the first analysis, AI selected significantly: more White/Hispanic applicants (p < 0.001), less signals (p < 0.001), more AOA honors society (p = 0.016), and more publications (p < 0.001). When censoring overlapping PD and AI selection, AI selected significantly: more White/Hispanic applicants (p < 0.001), less signals (p < 0.001), more US medical graduates (p = 0.027), less applicants needing visa sponsorship (p = 0.01), younger applicants (p = 0.024), higher USMLE Step 2 CK scores (p < 0.001), and more publications (p < 0.001). Conclusions There was only a 7% overlap between PD-selected and AI-selected applicants for interview screening in the same applicant pool. Despite the same PD educating the AI software, the 2 application pools differed significantly. In its present state, AI may be utilized as a tool in resident application selection but should not completely replace human review. We recommend careful analysis of the performance of each AI model in the respective environment of each institution applying it, as it may alter the group of interviewees.
OBJECTIVE: Intentionally self-driven professional development of surgical resident physicians is a hallmark of surgical training and is expected to gain further traction as Entrustable Professional Activities (EPAs) become the new paradigm for surgical education. We aimed to ana-lyze how surgical residents rate themselves as compared to the evaluation of the Clinical Competency Committee using ACGME Milestones Version 1 (M1.0) and Version 2 (M2.0).DESIGN: We asked 22 general surgical trainees for self-evaluation of Milestones (both M1.0 and M2.0) from 2017 semiannually to 2022. ACGME-required Milestone evaluations by the Clinical Competency Committee (CCC) were independently performed after the time window for resident self-evaluation. Neither trainees nor CCC were aware of the other party's evaluations. There were 1552 paired data available for evaluating individual competencies by both trainees and CCC. Paired Wil-coxon signed-rank tests were then performed among the corresponding pairs.SETTING: MercyOne Des Moines Medical Center, Des Moines, IA; Teaching tertiary referral center.PARTICIPANTS: Twenty-two general surgical trainees at this hospital and 28 faculty surgeons participated in this study.RESULTS: The average self-evaluation of surgical resi-dents was lower in the M1.0 cohort compared to the cor-responding CCC evaluation (1.96 +/- 0.72 vs. 2.11 +/- 0.67; p < 0.001). M1.0 self-assessments and CCC-assessments were statistically similar for ICS (p = 0.548) and PROF (p = 0.554) competencies and differed for MK (p < 0.001), PBLI (p < 0.001), PC (p < 0.001), SBP (p = 0.008). On the contrary, the M2.0 cohort demon-strated higher average self-evaluation of surgical resi-dents compared to the corresponding CCC evaluation (2.75 +/- 0.87 vs. 2.12 +/- 0.97; p < 0.001). Significant differences were observed for all 6 ACGME competencies using M2.0 self-assessments and CCC-assessments (all < 0.001). Multivariate regression modeling (p < 0.001, R2 = 0.255) predicted the degree of discordance between self-assessment and CCC-assessed achievement of compe-tencies with a significant effect of gender (baseline male: coef =-0.232, p < 0.001), PGY level (-0.083 per year, < 0.001) and Milestone version (0.831, p < 0.001). A sig-nificant interaction exists for all gender/Milestone combi-nations except for the female trainees with M1.0.CONCLUSIONS: The difference between self-evaluated Milestone achievement and faculty-driven CCC evalua-tion of surgical resident physician performance is more evident in Milestones 2.0 than in Milestones 1.0. Resi-dents self-evaluate higher compared to faculty using Milestones 2.0. This discrepancy is seen among both genders and is more pronounced among male residents overestimating core competencies with M2.0 self-evalua-tion than formal CCC assessment. ( J Surg Ed 80:1378-1384. (c) 2023 Published by Elsevier Inc. on behalf of Association of Program Directors in Surgery.)
We characterized the peritoneal immune cellular profile during cytoreductive surgery and hyperthermic intraperitoneal chemotherapy (HIPEC) in this pilot study. We prospectively performed flow cytometric analysis of peritoneal fluid collected at laparotomy and during HIPEC at 0, 30, 60, and 90 min. Analysis consisted of standard flow cytometric leukocyte gating and the use of antibodies for stem cells, B lymphocytes, T-helper, T-suppressor, and natural killer (NK) cells. The mean peritoneal carcinomatosis index (PCI) score was 19.8 ± 11.5 (median 19). Twelve patients had a completeness of cytoreduction (CCR) score of 0–1, and three patients had a CCR score of ≥ 2 (20%). The proportion of peritoneal NK cells remained stable (p = 0.655) throughout perfusion. The CD4/CD8 ratio (p = 0.019) and granulocyte/lymphocyte ratio (p = 0.018) evolved during cytoreduction, with no further change during HIPEC. Two distinct temporal patterns of peritoneal T lymphocytes became evident (the ‘high’ and ‘low’ CD4/CD8 ratio groups) and patients maintained their high versus low peritoneal CD4/CD8 ratio status throughout the duration of HIPEC. High CD4/CD8 was associated with longer cytoreduction (p = 0.019) and borderline higher PCI score (p = 0.058). No association was identified with age (p = 0.131), sex (p = 1.000), CCR status (p = 0.580), occurrence of complication (p = 0.282), or ascites volume (p = 0.713). The cellular immunoprofile of peritoneal fluid during HIPEC is stable but changes during cytoreduction. Two distinct immune groups emerged, based on CD4/CD8 ratios in the peritoneal perfusate. Further studies are warranted to evaluate peritoneal immunity and the clinical significance of novel peritoneal immune phenotype.