Neurosurgery trainees from high-income countries (HICs) can learn valuable clinical, technical, and public health skills via international observerships at partner institutions in low- and middle-income countries (LMICs). Implementing a program that is mutually beneficial to both HIC and LMIC collaborators requires thoughtful planning and consideration of best practices in global partnerships. This report shares one such experience, including perspectives of both the HIC rotator and host LMIC institution, providing guidance for similar initiatives elsewhere. A senior neurosurgery resident from the United States (US) completed a 6-week pediatric neurosurgery observership at the Philippine General Hospital (PGH) in Manila, Philippines from September to October 2023. To evaluate the experience, the resident gave a formal presentation and submitted an exit essay, while PGH trainees completed a survey, and the consultant faculty participated in a focus group discussion. Pre-observership preparation involved forming partnerships, securing funding, setting mutual objectives, creating schedules, designing research projects, arranging travel and lodging, and pursuing cultural humility training. During the observership, the rotator was fully immersed in the PGH resident schedule, including team rounds, surgeries, clinics, and on-call shifts. Academic activities included a quality improvement project on external ventricular drain management and a pilot study on telementorship in the cadaver lab. PGH neurosurgery faculty supervised the rotator, supported by weekly virtual conferences and visits from U.S.- and Philippines-based pediatric neurosurgery faculty. Post-observership benefits included clinical, surgical, and public health insights, with positive feedback from PGH trainees and faculty. This international observership was perceived as highly educational and mutually beneficial, due to shared goals, structured mentorship, and emphasis on bidirectional learning. Clear communication, well-defined clinical and research duties, cultural humility, and equitable engagement are key to overcoming challenges previously described in HIC-LMIC collaborations.
Glioblastoma is the most common form of primary brain tumor in adults, characterized by rapid progression and poor prognosis—despite the standard of care treatment including maximal safe resection, radiotherapy, and chemotherapy. Cancer vaccination has emerged as a promising strategy to harness the patient's immune system against glioblastoma. Cancer vaccination strategies can broadly be divided into cell-based or tumor antigen only (TAO), depending on whether they incorporate the use of viable immune cells. Here, we reviewed data from clinical trials that tested TAO cancer vaccination strategies for glioblastoma treatment, including personalized vaccines. Clinical safety and efficacy profiles for each vaccination strategy are summarized. Insights gained from these clinical trials are reviewed to identify opportunities for future therapeutic advancement.
Measuring physician productivity is a central function of health care management because clinical work generates the majority of revenue, including that of academic medical centers. Although numerous methods exist to quantify clinical productivity, most focus on the assessment of individual providers. The authors argue that individual-focused metrics - whether these are direct financial measures, such as billings or collections, or indirect estimates of work, such as relative value units - create misaligned incentives that compromise the best interests of patients, physicians, and the institution. Drawing from their experience in the Department of Neurosurgery at Brown University Health, the authors describe the rationale for, and benefits of, a group-based productivity model implemented in cooperation with institutional leadership. Specifically, shifting from individual to group metrics has created a collaborative rather than competitive environment that enables more expert, more effcient, and, ultimately, more ethical patient care. The authors contend this approach aligns interests and incentives to better promote optimal care for individual patients while also serving the myriad financial and nonfinancial needs of an academic program. The authors explain how, in principle and in practice, this system better addresses the core task of delivering high-quality patient care while providing significant benefits for surgeons, the department, education and research, and the institution. The potential limitations of such an approach are discussed, and the generalizability of this model beyond neurosurgery is considered.
OBJECTIVE Innovations in robotics continue to reshape the landscape of neurosurgery. Here, the authors evaluated the safety and efficacy of the ExcelsiusGPS robot in the treatment of neuro-oncological, intracranial lesions. METHODS The authors conducted a retrospective analysis of 19 consecutive adult patients with a neuro-oncological diagnosis who underwent intracranial biopsy and/or laser interstitial thermal therapy (LITT) with the assistance of the ExcelsiusGPS robot and intraoperative CT. Demographic and clinical data were collected from the electronic medical record and the robot software. RESULTS All 19 patients harbored lesions that were deep seated, involving the eloquent cortex, or subcentimeter. Definitive tissue diagnosis was achieved in all cases involving stereotactic biopsy (n = 16), with glioblastoma as the most common diagnosis. The mean +/- SD time for setting up the robotic stereotaxis system was 57.4 +/- 10.7 minutes. The mean procedural time after that was 71.6 +/- 41.0 minutes for stereotactic needle biopsy and 188.4 +/- 61.2 minutes for procedures involving LITT. The mean radial errors of the actual trajectory relative to the planned trajectory at the entry and target points were 0.625 +/- 0.443 mm and 0.745 +/- 0.472 mm, respectively. There were no procedural complications or new postoperative deficits, although routine postoperative CT showed new hyperdensity at the target site in 3/19 patients (15.7%). All patients who underwent elective procedures were discharged by postoperative day 3 (mean 1.38 +/- 0.619 days). There were two 30-day readmissions (pulmonary embolus and general weakness), and neither was attributable to the surgical procedure. CONCLUSIONS The authors' pilot experience with the ExcelsiusGPS robot in neuro-oncology procedures indicates a favorable efficacy and safety profile.
BACKGROUND Suprasellar masses commonly include craniopharyngiomas and pituitary adenomas. Suprasellar glioblastoma is exceedingly rare with only a few prior case reports in the literature. Suprasellar glioblastoma can mimic craniopharyngioma or other more common suprasellar etiologies preoperatively. OBSERVATIONS A 65-year-old male with no significant history presented to the emergency department with a subacute decline in mental status. Work-up revealed a large suprasellar mass with extension to the right inferior medial frontal lobe and right lateral ventricle, associated with significant vasogenic edema. The patient underwent an interhemispheric transcallosal approach subtotal resection of the interventricular portion of the mass. Pathological analysis revealed glioblastoma, MGMT partially methylated, with a BRAF V600E mutation. LESSONS Malignant glioblastomas can mimic benign suprasellar masses and should remain on the differential for a diverse set of brain masses with a broad range of radiological and clinical features. For complex cases accessible from the ventricle where the pituitary complex cannot be confidently preserved via a transsphenoidal approach, an interhemispheric approach is also a practical initial surgical option. In addition to providing diagnostic value, molecular profiling may also reveal therapeutically significant gene alterations such as BRAF mutations.
OBJECTIVE:To establish whether or not a natural language processing technique could identify two common inpatient neurosurgical comorbidities using only text reports of inpatient head imaging.MATERIALS AND METHODS:A training and testing dataset of reports of 979 CT or MRI scans of the brain for patients admitted to the neurosurgery service of a single hospital in June 2021 or to the Emergency Department between July 1-8, 2021, was identified. A variety of machine learning and deep learning algorithms utilizing natural language processing were trained on the training set (84% of the total cohort) and tested on the remaining images. A subset comparison cohort (n = 76) was then assessed to compare output of the best algorithm against real-life inpatient documentation.RESULTS:For "brain compression", a random forest classifier outperformed other candidate algorithms with an accuracy of 0.81 and area under the curve of 0.90 in the testing dataset. For "brain edema", a random forest classifier again outperformed other candidate algorithms with an accuracy of 0.92 and AUC of 0.94 in the testing dataset. In the provider comparison dataset, for "brain compression," the random forest algorithm demonstrated better accuracy (0.76 vs 0.70) and sensitivity (0.73 vs 0.43) than provider documentation. For "brain edema," the algorithm again demonstrated better accuracy (0.92 vs 0.84) and AUC (0.45 vs 0.09) than provider documentation.DISCUSSION:A natural language processing-based machine learning algorithm can reliably and reproducibly identify selected common neurosurgical comorbidities from radiology reports.CONCLUSION:This result may justify the use of machine learning-based decision support to augment provider documentation.
OBJECTIVES:Symptomatic carotid web is an increasingly recognized cause of acute ischemic stroke with a high risk of recurrent ischemic events despite aggressive medical interventions. Surgical interventions including transfemoral carotid artery stenting (TFCAS) and carotid endarterectomy have been described to reduce this risk, but transcarotid arterial revascularization (TCAR) has not been evaluated for this purpose. MATERIALS AND METHODS:Patients with cerebral ischemia from carotid web underwent TCAR with flow reversal. Patients were monitored for periprocedural complications and assessed at follow-up for clinical evidence of recurrent ischemia. RESULTS:Six cases over the course of 21 months were identified, 2 males and 4 females with a median age of 59.5 (interquartile range of 39). All underwent technically successful TCAR without periprocedural complications no post-procedural cerebral ischemia over a median follow-up time of 21 months. CONCLUSIONS:In this small series of patients, TCAR provided a safe and effective treatment of carotid webs that had previously caused cerebral ischemia.
Importance The progression of artificial intelligence (AI) text-to-image generators raises concerns of perpetuating societal biases, including profession-based stereotypes. Objective To gauge the demographic accuracy of surgeon representation by 3 prominent AI text-to-image models compared to real-world attending surgeons and trainees. Design, Setting, and Participants The study used a cross-sectional design, assessing the latest release of 3 leading publicly available AI text-to-image generators. Seven independent reviewers categorized AI-produced images. A total of 2400 images were analyzed, generated across 8 surgical specialties within each model. An additional 1200 images were evaluated based on geographic prompts for 3 countries. The study was conducted in May 2023. The 3 AI text-to-image generators were chosen due to their popularity at the time of this study. The measure of demographic characteristics was provided by the Association of American Medical Colleges subspecialty report, which references the American Medical Association master file for physician demographic characteristics across 50 states. Given changing demographic characteristics in trainees compared to attending surgeons, the decision was made to look into both groups separately. Race (non-White, defined as any race other than non-Hispanic White, and White) and gender (female and male) were assessed to evaluate known societal biases. Exposures Images were generated using a prompt template, "a photo of the face of a [blank]", with the blank replaced by a surgical specialty. Geographic-based prompting was evaluated by specifying the most populous countries on 3 continents (the US, Nigeria, and China). Main Outcomes and Measures The study compared representation of female and non-White surgeons in each model with real demographic data using χ2, Fisher exact, and proportion tests. Results There was a significantly higher mean representation of female (35.8% vs 14.7%; P < .001) and non-White (37.4% vs 22.8%; P < .001) surgeons among trainees than attending surgeons. DALL-E 2 reflected attending surgeons' true demographic data for female surgeons (15.9% vs 14.7%; P = .39) and non-White surgeons (22.6% vs 22.8%; P = .92) but underestimated trainees' representation for both female (15.9% vs 35.8%; P < .001) and non-White (22.6% vs 37.4%; P < .001) surgeons. In contrast, Midjourney and Stable Diffusion had significantly lower representation of images of female (0% and 1.8%, respectively; P < .001) and non-White (0.5% and 0.6%, respectively; P < .001) surgeons than DALL-E 2 or true demographic data. Geographic-based prompting increased non-White surgeon representation but did not alter female representation for all models in prompts specifying Nigeria and China. Conclusion and Relevance In this study, 2 leading publicly available text-to-image generators amplified societal biases, depicting over 98% surgeons as White and male. While 1 of the models depicted comparable demographic characteristics to real attending surgeons, all 3 models underestimated trainee representation. The study suggests the need for guardrails and robust feedback systems to minimize AI text-to-image generators magnifying stereotypes in professions such as surgery.
Background Aneurysmal subarachnoid hemorrhage (aSAH) is a major source of morbidity and mortality, and its management has undergone foundational changes over thepast 2 decades. We reviewed the National Inpatient Sample to outline the changes in severity of illness, surgical management, and patient outcomes over time. Methods A retrospective cohort of admissions for spontaneous SAH in the National Inpatient Sample from 2001 to 2020 was reviewed, including those who underwent microsurgical or endovascular surgery to secure a ruptured aneurysm. National incidence was calculated, and multivariable regression was used to identify changes in incidence and outcome through time, segmented by epoch. Results Review of the National Inpatient Sample identified 448 655 patients with SAH, of whom 181 590 underwent surgical aneurysm treatment. The incidence of spontaneous SAH fell −0.097 per 100 000 person‐years each year (95%CI, −0.144 to −0.049). Among patients surgically treated for aneurysmal SAH, the proportion of patients <50 years old fell from 40% to 30% between first and final epochs, and the proportion of those in the lowest stroke scale category, roughly equivalent to Hunt and Hess grade 1 or 2, fell from 66% to 45%. The proportion treated by microsurgery fell from 70% to 23% in favor of endovascular surgery. Hospital mortality among these treated cases was stable at 13% throughout the study period despite increasing illness severity indices. After adjustment, there was a 42% reduction of odds of hospital mortality in the final epoch compared with the first. Conclusion The incidence of hospitalization for spontaneous SAH fell between 2001 and 2020. Patients undergoing surgery to secure an aneurysm were more severely ill through time yet experienced a stable hospital mortality rate.
Despite the importance of informed consent in healthcare, the readability and specificity of consent forms often impede patients’ comprehension. This study investigates the use of GPT-4 to simplify surgical consent forms and introduces an AI-human expert collaborative approach to validate content appropriateness. Consent forms from multiple institutions were assessed for readability and simplified using GPT-4, with pre- and post-simplification readability metrics compared using nonparametric tests. Independent reviews by medical authors and a malpractice defense attorney were conducted. Finally, GPT-4’s potential for generating de novo procedure-specific consent forms was assessed, with forms evaluated using a validated 8-item rubric and expert subspecialty surgeon review. Analysis of 15 academic medical centers’ consent forms revealed significant reductions in average reading time, word rarity, and passive sentence frequency (all P < 0.05) following GPT-4-faciliated simplification. Readability improved from an average college freshman to an 8th-grade level (P = 0.004), matching the average American’s reading level. Medical and legal sufficiency consistency was confirmed. GPT-4 generated procedure-specific consent forms for five varied surgical procedures at an average 6th-grade reading level. These forms received perfect scores on a standardized consent form rubric and withstood scrutiny upon expert subspeciality surgeon review. This study demonstrates the first AI-human expert collaboration to enhance surgical consent forms, significantly improving readability without sacrificing clinical detail. Our framework could be extended to other patient communication materials, emphasizing clear communication and mitigating disparities related to health literacy barriers.
Informed consent is integral to the practice of medicine. Most informed consent documents are written at a reading level that surpasses the reading comprehension level of the average American. Large language models, a type of artificial intelligence (AI) with the ability to summarize and revise content, present a novel opportunity to make the language used in consent forms more accessible to the average American and thus, improve the quality of informed consent. In this study, we present the experience of the largest health care system in the state of Rhode Island in implementing AI to improve the readability of informed consent documents, highlighting one tangible application for emerging AI in the clinical setting.
To the Editor: Among the quality measures used for Comprehensive Stroke Center Certification by the Joint Commission is the proportion of patients presenting with an intracerebral hemorrhage (ICH) or subarachnoid hemorrhage (SAH) for whom an initial ICH score or Hunt and Hess (HH) grade, respectively, is documented.1,2 While these measures may seem straightforward, there are challenges in consistently adhering to them. The responsibility to document often falls on house staff, who can come from various specialties and who frequently turnover. In addition, in the setting of the patient with acute ICH or SAH, the documentation of an initial ICH score or HH grade may be overlooked. Because these 2 clinical entities are among the most staple consults overseen by neurosurgery residents and pathologies managed by the neurosurgery service line, there exists a need for a streamlined solution that can improve adherence to the Joint Commission's metrics for documentation of initial ICH scores and HH grades. To improve adherence to these Joint Commission specifications, our institution, a Comprehensive Stroke Center in New England, created an in-house, customized stroke admissions navigator on our Epic Systems electronic health record. Admissions navigators are tools that can facilitate more streamlined intake of admitted patients. Among the features built into our in-house, customized stroke admissions navigator was a panel to document ICH score for patients with ICHs and a panel to document HH grade for patients with SAHs. Once inputted, these values autopopulate into an H&P note template, ensuring their accurate documentation. All neurosurgery and neurology residents at our institution were instructed on how to use the stroke admissions navigator in March of 2019, and the use of the navigator went live in April of 2019. For statistical analysis, percentage rates in documentation before and after the intervention were compared with χ2 tests. In addition, an interrupted time series analysis using an autoregressive integrated moving average linear spline model was performed to assess for changes in documentation percentage after the intervention while adjusting for temporal variations, including across the study period and within individual years.3 The nonparametric Conover two-sample ranked square tests were used to query changes in documentation variance before and after the intervention.4 The study population was composed of 621 total patients presenting for ICH or SAH from January 1, 2018, to December 31, 2021, including 155 patients with SAH (25.0%) and 466 patients with ICH (75.0%). After the implementation of the intervention, there was a significant increase in Comprehensive Stroke Center Certification–associated documentation for the overall study population (82.1% vs 63.9%, P < .001; Figure 1A). This increase was also independently observed for patients with SAH (89.2% vs 75.0%, P = .027; Figure 1B) and ICH patients with (79.1% vs 61.1%, P < .001; Figure 1C). In an autoregressive integrated moving average time series model adjusting for longitudinal and intrayear changes in documentation during the study timeline, the intervention was associated with a significant increase in documentation for all (+11.7%, P = .004) and patients with ICH (+12.4%, P = .001; Figure 1D-1F). While documentation rates for patients with SAH underwent an increase approaching significance (+20.2%, P = .090), the intervention was nevertheless associated with a significant decrease in month-by-month variance in documentation (SD = 15.9% vs 36.7%, P = .005; Figure 1E).FIGURE 1.: Improvement in adherence to Joint Commission specifications after implementation of stroke admissions navigator. Changes to adherence to Joint Commission specifications for documenting Hunt and Hess scale for patients with SAH and ICH scores for patients with ICH after implementation of a stroke admissions navigator. A-C, Monthly percentage rates of documentation of scores for overall A, patients with stroke, B, Hunt and Hess scale for SAH and C, ICH scores for ICH from January 2018 to December 2021. D-F, Box and whisker plots for documentation of scores for D, overall patients with stroke, E, Hunt and Hess scale for SAH, and F, ICH scores for ICH before (red) and after (blue) implementation of the navigator on April of 2019. ICH, intracerebral hemorrhage; SAH, subarachnoid hemorrhage.In summation, using an in-house, customized stroke navigator, our institution improved adherence to Joint Commission benchmarks for reporting of initial ICH score and HH grade. To the best of our knowledge, this is the first publication demonstrating improvement in key documentation in stroke care through the utilization of an admission navigator. Based on this work, institutions with Comprehensive Stroke Center status should consider the utilization of a similar stroke navigator to improve documentation of this and other key metrics. Moreover, the implications of this work are extensible beyond cerebrovascular and endovascular neurosurgery, as similar admissions navigators may improve documentation of metrics and scoring systems of value to other neurosurgical subspecialties and clinical entities.
BACKGROUND:The relationship between marital status and overall survival (OS) in adult patients with craniopharyngioma has not been explored in depth. We aimed to elucidate the impact of marital status on the prognosis of craniopharyngioma patients excluding bias from baseline demographics and treatment. METHODS:We extracted 1539 patients diagnosed with craniopharyngioma between 2000 and 2019 from the Surveillance, Epidemiology, and End Results database and divided patients into 4 marital subgroups: married, single, divorced/separated, and widowed. Kaplan-Meier curves with a log-rank test were used to discern differences in OS between marital subgroups. Univariate and multivariate Cox regression were used to identify independent prognostic factors of mortality. RESULTS:There were 1539 eligible patients: 863 (56.1%) were married, 466 (30.3%) were single, 135 (8.8%) were divorced/separated, and 75 (4.9%) were widowed. Widowed patients had the worst mean OS, 5-year OS and 10-year OS at 84.2 months, 58.0% and 26.9%, respectively. After stratifying patients by age, our multivariate analysis showed that marital status was an independent predictor of mortality in middle-aged craniopharyngioma patients (40-60 years, P < 0.001), but not in young adults (18-39 years, P = 0.646) or elderly patients (>60 years, P = 0.076). Among middle-aged patients, single (hazard ratio 1.72, confidence interval 1.19-2.47, P = 0.004) and divorced/separated patients (hazard ratio = 2.29, confidence interval = 1.49-3.54, P < 0.001) showed a higher risk of mortality compared to married patients (reference). CONCLUSIONS:Marital status is an independent prognostic factor predicting OS for middle-aged patients with craniopharyngioma. Providing additional social and psychological support for single and divorced/separated patients may improve outcomes for this vulnerable population.
BACKGROUND AND OBJECTIVES: Interest surrounding generative large language models (LLMs) has rapidly grown. Although ChatGPT (GPT-3.5), a general LLM, has shown near-passing performance on medical student board examinations, the performance of ChatGPT or its successor GPT-4 on specialized examinations and the factors affecting accuracy remain unclear. This study aims to assess the performance of ChatGPT and GPT-4 on a 500-question mock neurosurgical written board examination. METHODS: The Self-Assessment Neurosurgery Examinations (SANS) American Board of Neurological Surgery Self-Assessment Examination 1 was used to evaluate ChatGPT and GPT-4. Questions were in single best answer, multiple-choice format. χ 2 , Fisher exact, and univariable logistic regression tests were used to assess performance differences in relation to question characteristics. RESULTS: ChatGPT (GPT-3.5) and GPT-4 achieved scores of 73.4% (95% CI: 69.3%-77.2%) and 83.4% (95% CI: 79.8%-86.5%), respectively, relative to the user average of 72.8% (95% CI: 68.6%-76.6%). Both LLMs exceeded last year's passing threshold of 69%. Although scores between ChatGPT and question bank users were equivalent ( P = .963), GPT-4 outperformed both (both P < .001). GPT-4 answered every question answered correctly by ChatGPT and 37.6% (50/133) of remaining incorrect questions correctly. Among 12 question categories, GPT-4 significantly outperformed users in each but performed comparably with ChatGPT in 3 (functional, other general, and spine) and outperformed both users and ChatGPT for tumor questions. Increased word count (odds ratio = 0.89 of answering a question correctly per +10 words) and higher-order problem-solving (odds ratio = 0.40, P = .009) were associated with lower accuracy for ChatGPT, but not for GPT-4 (both P > .005). Multimodal input was not available at the time of this study; hence, on questions with image content, ChatGPT and GPT-4 answered 49.5% and 56.8% of questions correctly based on contextual context clues alone. CONCLUSION: LLMs achieved passing scores on a mock 500-question neurosurgical written board examination, with GPT-4 significantly outperforming ChatGPT.
BACKGROUND:Tenosynovial giant cell tumor (TGCT) occurs most commonly in the appendicular skeleton and is only rarely found in the vertebral column. Lesions of the craniocervical junction are particularly rare, with only 4 cases reported in the literature. The authors describe the case of a diffuse-type TGCT at the craniocervical junction.OBSERVATIONS:A patient presented with a 1-year history of right-sided neck pain and bilateral neurological symptoms in the distribution of the right occipital nerve. A 20-mm homogeneously contrast-enhancing mass in the suboccipital and posterior C1 region was discovered on magnetic resonance imaging of the cervical spine. The tumor was operated on via a posterior approach, and gross-total resection (GTR) was achieved. Immunohistochemical (IHC) examination revealed a diffuse-type TGCT. The patient had an uneventful recovery.LESSONS:TGCT can arise at the craniocervical junction and is easily misdiagnosed because of its rare occurrence. IHC examination of a tumor specimen should be done to confirm the diagnosis. GTR is the objective when treating these tumors, especially when they are the diffuse type, as they have a high recurrence rate. Radiation and small-molecule therapies are viable postoperative therapies if GTR cannot be achieved or in cases of recurrence.
Background This study investigates the accuracy of three prominent artificial intelligence (AI) text-to-image generators—DALL-E 2, Midjourney, and Stable Diffusion—in representing the demographic realities in the surgical profession, addressing raised concerns about the perpetuation of societal biases, especially profession-based stereotypes. Methods A cross-sectional analysis was conducted on 2,400 images generated across eight surgical specialties by each model. An additional 1,200 images were evaluated based on geographic prompts for three countries. Images were generated using a prompt template, “A photo of the face of a [blank]”, with blank replaced by a surgical specialty. Geographic-based prompting was evaluated by specifying the most populous countries for three continents (United States, Nigeria, and China). Results There was a significantly higher representation of female (average=35.8% vs. 14.7%, P<0.001) and non-white (average=37.4% vs. 22.8%, P<0.001) surgeons among trainees than attendings. DALL-E 2 reflected attendings’ true demographics for female surgeons (15.9% vs. 14.7%, P=0.386) and non-white surgeons (22.6% vs. 22.8%, P=0.919) but underestimated trainees’ representation for both female (15.9% vs. 35.8%, P<0.001) and non-white (22.6% vs. 37.4%, P<0.001) surgeons. In contrast, Midjourney and Stable Diffusion had significantly lower representation of images of female (0% and 1.8%, respectively) and non-white (0.5% and 0.6%, respectively) surgeons than DALL-E 2 or true demographics (all P<0.001). Geographic-based prompting increased non-white surgeon representation (all P<0.001), but did not alter female representation (P=0.779). Conclusions While Midjourney and Stable Diffusion amplified societal biases by depicting over 98% of surgeons as white males, DALL-E 2 depicted more accurate demographics, although all three models underestimated trainee representation. These findings underscore the necessity for guardrails and robust feedback systems to prevent AI text-to-image generators from exacerbating profession-based stereotypes, and the importance of bolstering the representation of the evolving surgical field in these models’ future training sets.
To the Editor: Artificial intelligence's (AI's) increasing viability as an aide to clinical practice marks an inflection point in health care. In the setting of neurosurgery, our group has shown that the general Large Language Model (LLM) GPT-4 (OpenAI), the successor to ChatGPT, outperformed average human users in a mock neurosurgical written boards examination.1 Neurosurgery has a history of embracing new technologies, so the dawning era of AI in medicine will likely play to our strengths. However, neurosurgeons should not just function as passive observers and adopters but rather strive to become leaders in health care systems, industry, and policy to ensure that the benefits of AI for patients and neurosurgical practice are fully realized. As AI's role in health care grows, it is crucial for neurosurgeons to be facile and proficient with these technologies and help shape their impact on patient care. To this end, we outline 6 goals to guide neurosurgeons' response to AI advancements. Goal 1: Neurosurgeons bear the ultimate responsibility for patient care and should continually evaluate AI systems' readiness for patient-facing use. As medical professionals, we have an ethical duty to ensure that AI technologies adhere to the highest safety and efficacy standards, regardless of the technical capabilities of such systems. Neurosurgical cases often have substantive clinical equipoise, and existing AI models can struggle with higher-order management questions or even produce "hallucinations."2,3 Accordingly, our recent work has elucidated that certain AI systems not infrequently confabulate fabricated or incorrect answer rationales, especially when lacking contextual data.3 However, even if AI models did not hallucinate or struggle with higher-order reasoning, the onus would still rest on neurosurgeons to assess suitability for clinical application. Superimposing human judgement remains critical in achieving optimal patient outcomes. For our patients' well-being, it is vital that neurosurgeons continue to rigorously assess AI performance, take ownership of the integration process, and establish guardrails to minimize risk of adverse events. Goal #2: Neurosurgeons should use AI for nonclinical tasks to focus more on patient care. The advent of AI raises the possibility of increased self-sufficiency and efficiency in administrative tasks like the growing burden on physicians of clinical documentation and navigation of prior authorization.4 Recognizing this potential, Epic—a software electronic medical records company holding nearly 80% of Americans' health care data—has already announced a partnership with OpenAI.5 Incorporation of AI into health care may assist with reining in administrative bloat and planning quality improvement initiatives. LLMs may also function as educational resources for learning topics such as health care finance and performance assessment. Finally, AI may improve multidisciplinary collaboration in health care, such as incorporating novel clinical technologies. By integrating AI tools into daily workflows for nonclinical tasks, neurosurgeons can save time and energy for providing the best possible care.6 Goal #3: Neurosurgeons should act as leaders bridging industry and patients. As domain-focused LLMs emerge,7 neurosurgeons should guide the development, testing, and validation of specialized applications of AI for their field. This may involve backend development, frontend optimization, and ethical advising. In compliance with patient privacy laws, neurosurgeons may also proactively create and share open-source registries of clinical data to accelerate the development of tools addressing critical management concerns in our specialty. Goal #4: Neurosurgeons should advocate for appropriate policies, investment, and reimbursement structures to ensure safe and efficient AI adoption in health care. Policy and reimbursement structures can significantly influence the success of emerging neurosurgical technologies. For example, the US Food and Drug Administration's authorization of Gamma Knife and its incorporation into Medicare reimbursement policies accelerated this technology's spread in the 1980–1990s. As AI's potential benefits grow, neurosurgeons must engage in the political arena to facilitate its utilization. This may involve adding an AI-focused arm to existing political apparatuses, such as the American Association of Neurological Surgeons/Congress of Neurological Surgeons Washington Committee, to ensure that the right structures are in place for effective AI integration.8 Nevertheless, while doing so, neurosurgeons must heed prior cautionary tales, such as computer-aided detection in mammograms, which was widely used after gaining the US Food and Drug Administration approval in 1998 and adopted by hospitals to increase profits through extra charges but ultimately failed and even potentially worsened diagnostic accuracy.9 Goal #5: Organized neurosurgical societies should lead in addressing AI advancements. Organized societies may promote standards, develop best practice guidelines, create educational materials for their members, and fund AI research endeavors. Actionable next steps for organized neurosurgery may include forming task forces to study AI implementation, fostering collaboration with AI developers, and advocating for policies that ensure equitable AI integration into health care systems. Goal #6: Neurosurgeons should ensure AI in health care does not worsen existing inefficiencies or inequities in clinical decision-making. Imperfect training processes and data sets can result in biased recommendations, limiting generalizability to diverse patient populations.10,11 As one striking example of how algorithms may inadvertently learn sociocultural biases, Amazon scrapped an experimental AI recruiting tool after it consistently gave female job candidates lower ratings, reflective of historical gender-based hiring disparities captured by the training data set.12 This example underscores the importance of having standards for openness and transparency with respect to the data sets used. For instance, GPT-4 is trained on an unknown data set, which raises concerns about the potential biases it may have inherited. It is not difficult to envision a scenario where technologies trained on biased data may recommend against surgery for patients possessing specific characteristics that have been historically associated with poorer perioperative outcomes.13 Neurosurgeons must identify and address these biases to ensure effective and equitable deployment of AI into clinical care. In this piece, we have outlined just 6 areas for potential technical leadership in the domain of AI. Nevertheless, it remains to be seen how AI can further enable neurosurgeons in nontechnical, adaptive forms of leadership where social and emotional intelligence are increasingly recognized and valued as key drivers for influencing change, innovation, and the social proof needed to bring technological advances to the mainstream. The rise of AI has immeasurable potential impacts on health care and society; neurosurgeons must do our part to align these new tools with our collective needs and interests.
INTRODUCTION: To date, limited research has focused on how charting styles can affect quality metrics in neurosurgery. METHODS: We performed a comparative interrupted time-series analysis at an urban, academic, level I trauma center. The intervention consisted of adjusting the charting style from system-based to problem-based charting on an inpatient neurosurgery service line and providing a one-hour educational session targeting improved documentation of three problems with high clinical relevance to neurosurgery. RESULTS: Among patients whose providers were exposed to the charting improvement intervention, there was a significant increase in the E-MR (p = 0.029). Average E-LOS also increased to a degree that approached significance (p = 0.085). In a multivariate model adjusting for elective status, length of stay, cost of admission, and hospital associated complications, the increases in both E-MR (p = 0.029) and E-LOS (p = 0.091) persisted. The switch from SBC to PBC resulted in noninferior rates of documentation (p>0.05) for the six most common complications not targeted as documentation opportunities during the educational intervention. Among patients in the reference group, whose providers did not receive the charting intervention, there was no significant change in either the E-MR (p = 0.198) or the E-LOS (p = 0.341), both of which slightly decreased over the course of the study period. CONCLUSIONS: Switching from SBC to specialty-focused PBC on a neurosurgery service line at an urban, academic, level I trauma center led to a statistically significant increase in the average patient-specific E-MR and a trend toward significant increase in the average patient-specific E-LOS. The study illustrates the degree to which a low-cost intervention on documentation patterns can impact neurosurgical risk-adjusted quality metrics.
Delayed cerebral ischemia (DCI) is a major etiology of poor neurologic outcomes after aneurysmal subarachnoid hemorrhage (aSAH). Although the development of DCI is certainly multifactorial, the presence of vasospasm is strongly correlated with it. Cerebral angiography remains the gold standard for evaluation of vasospasm, though it is not always practical or cost-effective. In this study, the authors assess the utility of automated MRI Perfusion imaging, with or without MR Angiography (MRA), as a confirmatory tool for suspected angiographic vasospasm. All patients admitted to a single institution with aneurysmal subarachnoid hemorrhage between January 2014 and February 2020 and who underwent MR Perfusion imaging with or without MRA for suspected vasospasm no >24 h prior to an angiogram were identified. 43 subjects were identified. 29 of these patients (67%) underwent simultaneous MRA. 25 patients (53%) received intra-arterial treatment for symptomatic vasospasm. The sensitivity, specificity, PPV, and NPV of MR Perfusion were 43%, 82%, 53%, and 75% for any angiographic vasospasm and 57%, 81%, 42%, and 89% for treated vasospasm. The sensitivity, specificity, PPV, and NPV of MR Perfusion in conjunction with MRA were 61%, 81%, 59%, and 82% for any angiographic vasospasm and 62%, 74%, 35%, and 89% for treated vasospasm. The sensitivity, specificity, PPV, and NPV of transcranial Dopplers (TCDs) in these patients were 35%, 93%, 71%, and 75% for angiographic vasospasm and 42%, 90%, 47%, and 88% for treated vasospasm. Automated MR Perfusion imaging demonstrated relatively low sensitivity and PPV for detection of angiographic and treated vasospasm in this subset of patients after aSAH.
OBJECTIVES:Opioids are frequently used for analgesia in patients with acute subarachnoid hemorrhage (SAH) due to a high prevalence of headache and neck pain. However, it is unclear if this practice may pose a risk for opioid dependence, as long-term opioid use in this population remains unknown. We sought to determine the prevalence of opioid use in SAH survivors, and to identify potential risk factors for opioid utilization. METHODS:We analyzed a cohort of consecutive patients admitted with non-traumatic and suspected aneurysmal SAH to an academic referral center. We included patients who survived hospitalization and excluded those who were not opioid-naïve. Potential risk factors for opioid prescription at discharge, 3 and 12 months post-discharge were assessed. RESULTS:Of 240 SAH patients who met our inclusion criteria (mean age 58.4 years [SD 14.8], 58% women), 233 (97%) received opioids during hospitalization and 152 (63%) received opioid prescription at discharge. Twenty-eight patients (12%) still continued to use opioids at 3 months post-discharge, and 13 patients (6%) at 12-month follow up. Although patients with poor Hunt and Hess grades (odds ratio 0.19, 95% CI 0.06-0.57) and those with intraventricular hemorrhage (odds ratio 0.38, 95% CI 0.18-0.87) were less likely to receive opioid prescriptions at discharge, we did not find significant differences between patients who had long-term opioid use and those who did not. CONCLUSION:Opioids are regularly used in both the acute SAH setting and immediately after discharge. A considerable number of patients also continue to use opioids in the long-term. Opioid-sparing pain control strategies should be explored in the future.