The Mount Sinai Health System is a hospital network in New York City. It was formed in September 2013 by merging the operations of Continuum Health Partners and the Mount Sinai Medical Center.The Health System is structured around eight hospital campuses, the Icahn School of Medicine at Mount Sinai and Phillips School of Nursing at Mount Sinai Beth Israel. The eight hospitals are: Mount Sinai Beth Israel, Mount Sinai Brooklyn, Mount Sinai Hospital (including Kravis Children's Hospital), Mount Sinai Queens, Mount Sinai Morningside (formerly Mount Sinai St. Luke's), Mount Sinai West (formerly Mount Sinai Roosevelt), New York Eye and Ear Infirmary of Mount Sinai, and Mount Sinai South Nassau.The Health System includes more than 6,600 primary and specialty care physicians and 13 ambulatory surgical centers. It has ambulatory practices throughout the five boroughs of New York City, Westchester County, and Long Island, along with more than 30 affiliated community health centers.In the 2017-2018 fiscal year, the Health System employed more than 39,000 people and the Icahn School of Medicine at Mount Sinai had 33 multidisciplinary research, educational, and clinical institutes. In addition, the Health System reported 3,360 beds among its seven hospitals as well as 136,528 inpatient admissions, 500,901 Emergency Department visits, and more than 14,700 babies delivered.
BACKGROUND:Large language models (LLMs) are increasingly used in health care but remain vulnerable to medical misinformation. We aimed to evaluate how often these models accept or reject fabricated medical content, and how framing that content as a logical fallacy changes results. METHODS:In this cross-sectional benchmarking analysis, we probed 20 LLMs with more than 3·4 million prompts that all contained health misinformation drawn from three sources: public-forum and social-media dialogues, real hospital discharge notes in which we inserted a single false recommendation, and 300 physician-validated simulated vignettes. Logical fallacies-common patterns of flawed reasoning such as appeals to authority, popularity, or emotion-were used to test how rhetorical framing influences model behaviour. Each prompt was posed once in a neutral base form and ten times with a named logical fallacy. For every run we logged susceptibility (model accepts the false claim) and fallacy detection (model flags the rhetoric). FINDINGS:Across all models and corpora, LLMs were susceptible to fabricated data in 50 108 (31·7%) of 158 000 base prompts. Eight of ten fallacy framings significantly reduced or did not change that rate, led by appeal to popularity (susceptibility 11·9%; difference of -19·8 percentage points; p<0·0001); only the slippery-slope prompt (33·9%; difference of 2·2 percentage points; p<0·0001) and the appeal-to-authority prompt (34·6%; difference of 2·9 percentage points; p<0·0001) increased it. Real hospital notes (with fabricated inserted elements) produced the highest susceptibility to the base prompt (46 108 [46·1%] of 100 000), whereas social-media misinformation showed lower base prompt susceptibility (2479 [8·9%] of 28 000). Performance varied by model: GPT models were the least susceptible and most accurate at fallacy detection, whereas others, such as Gemma-3-4B-it, showed 63·6% (5023 of 7900) susceptibility. INTERPRETATION:These results show that LLMs still absorb harmful medical fabrications, especially when phrased in authoritative clinical prose, yet, counter-intuitively, become less vulnerable when the same claims are wrapped in most logical fallacy styles. Therefore, improving safety appears to depend less on model scale and more on fact-grounding and context-aware guardrails. FUNDING:Scientific Computing and Data at Icahn School of Medicine and National Institutes of Health Office of Research Infrastructure.
Accurate staging of unfavorable intermediate- or high-risk prostate cancer (PCa) is essential for treatment decisions. Conventional imaging often fails to detect lymph node, bone, and visceral metastases, and for this purpose 68Ga-prostate-specific membrane antigen (PSMA)-11 PET/CT is clinically used. This prospective, multicenter, International Atomic Energy Agency-supported trial evaluated the accuracy of 68Ga-PSMA-11 PET/CT for initial staging compared with MRI and histopathology and the impact of 68Ga-PSMA-11 PET/CT on determining surgical eligibility. Methods: In a prospective, international study supported by the International Atomic Energy Agency, 775 patients with high-risk or unfavorable intermediate-risk PCa from 12 centers across 11 countries-including low-, middle-, and high-income settings, scheduled for radical prostatectomy based on conventional imaging (including bone scanning and pelvic MRI) underwent 68Ga-PSMA-11 PET/CT before treatment. PET and MRI findings were compared with radical prostatectomy histopathology, and the impact of PET on radical prostatectomy was assessed. Results: 68Ga-PSMA-11 PET/CT detected metastatic disease (M1) in 20.4% of cases, altering management and preventing prostatectomy in 24.0%. The accuracy for seminal vesicle invasion was 90.1% for 68Ga-PSMA-11 PET/CT versus 57.3% for MRI, and for lymph node metastases it was 91.1% for 68Ga-PSMA-11 PET/CT versus 69.7% for MRI. In 13.1% of patients (78/593), there were discordant results between 68Ga-PSMA-11 PET/CT and histopathology. 68Ga-PSMA-11 PET/CT had false-negative lymph node findings in 8.6% of cases, with the most clinically significant being 4.5% of patients incorrectly staged as N0. False-positive lymph node findings at 68Ga-PSMA-11 PET/CT occurred in 4.5% of patients. Conclusion: 68Ga-PSMA-11 PET/CT significantly improves staging accuracy, reducing the indication for prostatectomy and impacting treatment decisions. These findings, from a broad international cohort including low-, middle-, and high-income countries, support the global adoption of 68Ga-PSMA-11 PET/CT into standard staging protocols for high-risk PCa.
Real-world studies based on electronic health records often require manual chart review to derive patients’ clinical phenotypes, a labor-intensive task with limited scalability. Here, we developed and compared computable phenotyping based on rules using the spaCy framework and a Large Language Model (LLM), GPT-4, for sub-phenotyping of patients with Crohn’s disease, considering age at diagnosis and disease behavior. For our rule-based approach, we leveraged the spaCy framework and for the LLM-based approach, we used the GPT-4 model. The underlying data included 49,572 clinical notes and 2204 radiology reports from 584 Crohn’s disease patients. A test set of 280 clinical texts was labeled at sentence-level, in addition to patient-level ground truth data. The algorithms were evaluated based on their recall, precision, specificity values, and F1 scores. Overall, we observe similar or better performance using GPT-4 compared to the rules. On a note-level, the F1 score is at least 0.90 for disease behavior and 0.82 for age at diagnosis, and on patient level at least 0.66 for disease behavior and 0.71 for age at diagnosis. To our knowledge, this is the first study to explore computable phenotyping algorithms based on clinical narrative text for these complex tasks, where prior inter-annotator agreements ranged from 0.54 to 0.98. There is no statistical evidence for a difference to the performance of human experts on this task. Our findings underline the potential of LLMs for computable phenotyping and may support large-scale cohort analyses from electronic health records and streamline chart review processes in the future. Doctors and researchers often need to group patients by specific medical features (called “phenotypes”) to study the disease and improve care. Much of this information is in free-text clinical notes rather than in tabular data. As an example, for patients with Crohn’s disease, the free-text notes can include important details describing the disease course over time, such as bowel narrowings (strictures), abnormal openings (fistulas), problems with the area around the anus, and age at diagnosis. Prior studies show that relying only on structured data, such as codes used to describe particular diseases, often miss these types of details. Reading notes by hand is more accurate but slow and costly. Natural language processing (NLP) is a computational method to automatically read and extract this information from clinical text. We built two NLP approaches and created new sentence-level datasets to test them. Both approaches found complications well, with combining both approaches giving the most balanced results. This method could save time in research and help clinicians flag people who may need extra care. Schmidt et al. develop sentence-level datasets for Crohn’s phenotypes and compare rule-based NLP with GPT-4 to extract disease behavior and age at diagnosis from EHR notes. Both methods achieve high recall on notes; GPT-4 perfectly identifies age at diagnosis and simple ensembles improve precision and enable chart-review prioritization.
Background: Anterior cervical discectomy and fusion (ACDF) is a widely performed surgical procedure for treating cervical spine pathologies, with pseudarthrosis remaining a significant postoperative challenge. While glucagon-like peptide-1 (GLP-1) receptor agonists have demonstrated beneficial effects on vascular health and bone metabolism, their impact on cervical fusion outcomes remains unexplored. This study investigates the relationship between perioperative GLP-1 agonist use and pseudarthrosis rates following single-level ACDF procedures. Methods: We conducted a retrospective propensity-matched cohort study using the TriNetX database. retrospective analysis using the TriNetX Research Network database, examining records from October 2010 to October 2022. The study population included patients who underwent single-level ACDF procedures. One-to-one propensity score matching (PSM) was performed to account for demographic factors, body mass index (BMI), hemoglobin A1c (HbA1c), and relevant comorbidities. Pseudarthrosis rates were evaluated at 6 months, 1 year, and 2 years postoperatively. Results: Of 28,133 patients who underwent ACDF, 555 were prescribed GLP-1 agonists within 6 months of surgery. After PSM, 546 patients were included in each cohort. The GLP-1 agonist cohort demonstrated significantly lower odds of developing pseudarthrosis at 6 months [odds ratio (OR): 0.60, 95% confidence interval (CI): 0.42-0.87], 1 year (OR: 0.65, 95% CI: 0.46-0.94), and 2 years (OR: 0.62, 95% CI: 0.44-0.86) postoperatively compared to the non-GLP-1 agonist cohort. Conclusions: GLP-1 agonist use was associated with significantly reduced pseudarthrosis rates following ACDF procedures across all measured time points. These findings suggest potential benefits of GLP-1 agonists in cervical fusion outcomes, independent of their metabolic effects. Further prospective studies are warranted to validate these results and elucidate the underlying biological mechanisms.
Mount Sinai Health System (MSHS), one of New York City's largest academic medical centers, faced a common patient-access challenge: individuals presenting with new symptoms often did not know where to turn, contributing to delayed care, unnecessary clinic and emergency department visits, and inefficient use of provider and facility resources. To improve care navigation, MSHS implemented a scalable, evidence-based, artificial intelligence (AI)-driven, digital self-triage solution. After a competitive market evaluation and request for proposal process beginning in 2021, MSHS selected Clearstep, a platform built in partnership with Dr. Barton Schmitt, the coauthor of the Schmitt-Thompson telephone triage protocols. Clearstep combined a probabilistic natural-language processing layer with a rules-based clinical expert system. The solution, branded "Check Symptoms & Get Care," was deployed across the MSHS public website and integrated within the MyMountSinai mobile app (powered by Epic) in early 2023. From over 60,000 visits, approximately 22,000 patients completed digital triage sessions (37% initiation and approximately 80% completion), with high satisfaction (a system usability scale score of 85.5 and approximately 75% of users rating greater than or equal to 8 out of 10), broad after-hours utilization (71% of use outside business hours), and zero reported safety incidents. Blinded clinician comparison and continuous post-deployment review demonstrated 88%-96% concordance with physician triage. The authors describe the team, hurdles, metrics, and a tiered road map so that organizations of varying technical readiness can adapt this model.