Health systems face a paradoxical translational gap: Despite operational and domain expertise and real-world implementation environments, they continue to face challenges in innovating with emerging technologies to improve care delivery. This gap often stems from a fundamental tension between the large-scale, centralized approach required for foundational information technology infrastructure and the nimble, decentralized methods essential for rapid, user-driven innovation, highlighting a critical need to reconcile these divergent mindsets within health systems. This case study describes how Stanford Health Care, a quaternary academic medical center, addressed this gap through a bottom-up grassroots innovation approach enabling rapid identification, iterative prototyping, and enterprise scaling of an artificial intelligence (AI)-enabled intervention that was sourced and developed internally by the frontline staff and resulted in operational impact at scale. FastFax is an automated triage system that assists the enterprise referral management team in the triage of urgent, externally faxed referrals, a previously manual process that required sorting through individual fax cover sheets. By leveraging an agile, user-centered approach, frontline staff on the referrals team identified key leverage points in their workflow that could be addressed by AI, resulting in the codevelopment of a targeted solution that shortened processing times for urgent faxed referrals from about 33 hours to about 1 hour, enabling the organization to reach its goal of same-day processing of urgent referrals. FastFax was initially piloted for 6 months from January through June 2023 and has continued post pilot as an interim enterprise-wide solution for triaging faxed referrals. FastFax has also informed the procurement of broader vendor solutions, demonstrating the value of health systems being active developers rather than passive consumers of technology. Indeed, based on the learnings and insights from the experience, FastFax - initially envisioned as a stopgap solution - is now being refined internally into FastFax 2.0 to address all faxed referrals, rather than pursuing an external vendor solution.
Repetitive laboratory testing that is unlikely to yield clinically useful information is a common practice that burdens patients and increases health care costs. Education and feedback interventions have limited success, while general test ordering restrictions and electronic alerts impede appropriate clinical care. We introduce and evaluate SmartAlert, a machine learning-driven clinical decision support (CDS) system integrated into the electronic health record that predicts stable laboratory results to reduce unnecessary repeat testing. This Case Study describes the implementation process, challenges, and lessons learned from deploying SmartAlert targeting complete blood count (CBC) utilization in a randomized controlled pilot across 9270 admissions in eight acute care units across two hospitals between August 15, 2024, and March 15, 2025. The results show a significant decrease in the number of CBC results within 52 hours of SmartAlert display (1.54 vs. 1.82; P<0.01) without adverse effect on secondary safety outcomes, representing a 15% relative reduction in repetitive testing. Implementation lessons learned include interpretation of probabilistic model predictions in clinical contexts, stakeholder engagement to define acceptable model behavior, governance processes for deploying a complex model in a clinical environment, user interface design considerations, alignment with clinical operational priorities, and the value of qualitative feedback from end users. In conclusion, a machine learning-driven CDS system backed by a deliberate implementation and governance process can provide precision guidance on inpatient laboratory testing to safely reduce unnecessary repetitive testing. (Funded by the Agency for Science, Technology, and Research and others.).
While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with "workflow friction" from manual data entry. We developed ChatEHR, a system that enables the use of LLMs with the entire patient timeline spanning several years. ChatEHR enables automations - which are static combinations of prompts and data that perform a fixed task - and interactive use in the electronic health record (EHR) via a user interface (UI). The resulting ability to sift through patient medical records for diverse use-cases such as pre-visit chart review, screening for transfer eligibility, monitoring for surgical site infections, and chart abstraction, redefines LLM use as an institutional capability. This system, accessible after user-training, enables continuous monitoring and evaluation of LLM use. In 1.5 years, we built 7 automations and 1075 users have trained to become routine users of the UI, engaging in 23,000 sessions in the first 3 months of launch. For automations, being model-agnostic and accessing multiple types of data was essential for matching specific clinical or administrative tasks with the most appropriate LLM. Benchmark-based evaluations proved insufficient for monitoring and evaluation of the UI, requiring new methods to monitor performance. Generation of summaries was the most frequent task in the UI, with an estimated 0.73 hallucinations and 1.60 inaccuracies per generation. The resulting mix of cost savings, time savings, and revenue growth required a value assessment framework to prioritize work as well as quantify the impact of using LLMs. Initial estimates are $6M savings in the first year of use, without quantifying the benefit of the better care offered. Such a "build-from-within" strategy provides an opportunity for health systems to maintain agency via a vendor-agnostic, internally governed LLM platform.
Stroke affected millions annually, yet poor symptom recognition often delayed care-seeking. To address risk recognition gap, we developed a passive surveillance system for early stroke risk detection using patient-reported symptoms among individuals with diabetes. Constructing a symptom taxonomy grounded in patients own language and a dual machine learning pipeline (heterogeneous GNN and EN/LASSO), we identified symptom patterns associated with subsequent stroke. We translated findings into a hybrid risk screening system integrating symptom relevance and temporal proximity, evaluated across 3-90 day windows through EHR-based simulations. Under conservative thresholds, intentionally designed to minimize false alerts, the screening system achieved high specificity (1.00) and prevalence-adjusted positive predictive value (1.00), with good sensitivity (0.72), an expected trade-off prioritizing precision, that was highest in 90-day window. Patient-reported language alone supported high-precision, low-burden early stroke risk detection, that could offer a valuable time window for clinical evaluation and intervention for high-risk individuals.
We examined telemedicine use across 38,883 surgical oncology visits (2021-2023) at a Northern California cancer center. At ≥20 miles from clinics, Hispanic (OR = 0.76, 95% CI [0.68,0.85]), Asian/Pacific Islander (OR = 0.75, 95% CI [0.66,0.84]), interpreter-needing (OR = 0.67, 95% CI [0.59,0.77]), and Medicaid patients (OR = 0.85, 95% CI [0.76,0.96]) had lower telemedicine use, while low-income patients showed higher utilization (OR = 1.67, 95% CI [1.46,1.91]). At <20 miles, no differences were observed for Hispanic, interpreter-needing, Medicaid, or low-income patients, but Asian/Pacific Islanders showed higher use (OR = 1.16, 95% CI [1.04-1.30]). Geographic distance modifies telemedicine access patterns.
Objectives:Electronic health record (EHR) order preference lists and order sets potentially improve efficiency but have limited utility in complex primary care settings. We assessed adoption, impact on ordering efficiency, and clinician perceptions of a comprehensive set of nested order panels (xOrders) for adult primary care. Methods:In Phase 1 (gradual implementation), 404 xOrders were released (November 29, 2020-September 25, 2021). Beginning of Phase 2 (rapid implementation), 630 xOrders were released with an additional 253 xOrders added (September 26, 2021-June 24, 2023). Three outcomes captured adoption: xOrders used per week; number of clinician users per week; and percent of xOrders of all orders. Impact of xOrders on times in orders per encounter per clinician was evaluated with mixed effects interrupted time series. t-Tests evaluated differences between low, moderate, and high utilizers. A survey captured clinicians' perceptions in November 2022. Results:xOrders were used 536 (SD, 245) times/week and by 57(15) clinicians/week in Phase 2. xOrders as a percent of all orders ranged from 0% to 31% across clinicians. Time spent in orders per encounter decreased by 14 ± 5 s (P =.01) from Phase 1 to 2 for high utilizers, decreased by 7(3) s (P=.05) for moderate utilizers, and increased by 1(3) s for low utilizers (P=.81); low and high utilizers were significantly different (P=.02). Most (77%) survey respondents agreed that xOrders improved ordering efficiency. Discussion and Conclusions:Despite yielding time savings and positive clinician feedback, the xOrder intervention showed limited adoption and impact, suggesting the need for expanded content and increased adoption to realize larger efficiency gains.
Objective:This study assessed over 2000 patient perspectives on the use of ambient AI scribes in outpatient visits. Materials and Methods:This prospective quality improvement study was conducted at Stanford Health Care between May and July 2025. Outcome measures included patient perceived helpfulness of ambient AI scribes and patient interest in future use. Results:Among 2202 survey respondents, 70.1% patients found the ambient AI scribe helpful and 73.6% preferred future use of ambient AI scribes. Small but statistically significant differences were observed across gender, age, and race. Discussion:In an evaluation of patient perceptions of an ambient AI scribe integrated into clinical practice, the majority of patients who had experienced the technology found it acceptable, and most found the tool helpful and desired use with future visits. Conclusion:These results suggest that ambient AI scribes may be acceptable to many patients in clinical practice. Further research is needed to inform patient-centered design and workflow improvements.
Despite advances in science and technology, persistent challenges in the delivery of healthcare call for care model transformations that have yet to be realized. Artificial intelligence could drive these transformations, but has yet to do so at scale. We present a four-layer framework for leveraging AI to design new care models: Knowledge (clinical content and institutional expertise), Intelligence (AI-powered synthesis and reasoning), Application (user interfaces), and Workflow (redesigned care processes). These layers are modular yet tightly interdependent, requiring cross-functional teams to design across the full stack. We illustrate this framework through an AI-enabled specialty consultation service deployed within Stanford Health Care, a quaternary academic medical center, that integrates all four layers to transform how expertise is delivered. This framework offers health system leaders a roadmap for moving beyond technology deployment toward systematic care model engineering—an organizational capability that will help shape the future of healthcare delivery.
This quality improvement study evaluates clinician perspectives on the usability and utility of generative artificial intelligence (AI)–based large language model tool to draft result comments for laboratory, imaging, and pathology results.
Importance:Limited qualitative studies exist evaluating ambient artificial intelligence (AI) scribe tools. Such studies can provide deeper insights into ambient AI implementations by capturing lived experiences. Objective:To evaluate physician perspectives on ambient AI scribes. Design, Setting, and Participants:A qualitative study using semistructured interviews guided by the Reach, Efficacy, Adoption, Implementation, Maintenance/Practical, Robust Implementation, and Sustainability Model (RE-AIM/PRISM) framework, with thematic analysis using both inductive and deductive approaches. Physicians participating in an AI scribe pilot that included community and faculty practices, across primary care and ambulatory specialties, were invited to participate in interviews. This ambient AI scribe pilot at a health care organization in California was conducted from November 2023 to January 2024. Main Outcome and Measures:Facilitators and barriers to adoption, practical effectiveness, and suggestions for improvement to enhance sustainability. Results:Twenty-two semistructured interviews were conducted with AI pilot physicians from primary care (13 [59%]) and ambulatory specialties (9 [41%]), including physicians from community practices (12 [55%]) and faculty practices (10 [45%]). Facilitators to adoption included ease of use, ease of editing, and generally positive perspectives of tool quality. Physicians expressed positive sentiments about the impact of the ambient AI scribe tool on cognitive demand (16 of 16 comments [100%]), temporal demand (28 comments [62%]), work-life integration (10 of 11 comments [91%]), and overall workload (8 of 9 comments [89%]). Physician perspectives of the impact of the ambient AI scribe tool on their engagement with patients were mostly positive (38 of 56 comments [68%]). Barriers to adoption included limited functionality with non-English speaking patients and lack of access for physicians without a specific device. Physician perspectives on accuracy and style were largely negative, particularly regarding note length and editing requirements. Several specific suggestions for tool improvement were identified, and physicians were optimistic regarding the potential for long-term use of ambient AI scribes. Conclusion and Relevance:In this qualitative study, ambient AI scribes were found to positively impact physician workload, work-life integration, and patient engagement. Key facilitators and barriers to adoption were identified, along with specific suggestions for tool improvement. These findings suggest the potential for ambient AI scribes to reduce clinician burden, with user-centered recommendations offering practical guidance on ways to improve future iterations and improve adoption.
Importance Large language model (LLM)-assisted early warning system may help overcome existing barriers to timely depression diagnosis in patients with cardiovascular disease (CVD). This novel application of LLMs to screen patient messages could be applied to other chronic diseases, facilitating automated symptom-driven diagnoses and interventions. Objective To prospectively simulate the impact (change in time to diagnosis) of population mental health screening using LLMs by screening patient portal messages, and measure LLMs accuracy to identify individuals at high risk for depression diagnosis in patients with CVD. Design Prospective cohort study Setting Electronic health records from an academic hospital (Stanford Health Care) Participants Individuals with CVD diagnosed 2014-2024, subsequently diagnosed with depression Intervention/Exposure LLMs (Llama 3.1 8B, July 2024 version, Meta LLC, and MedGemma 4B, July 2025 version, Google DeepMind, LLC) to identify individuals with depression Main outcome Accuracy of LLMs in sensitivity for completeness to capture positive cases and positive predictive value (PPV) for correctness to capture positive cases, and changes in time to depression diagnosis Results We identified 115,156 patients with CVD, and 23.1% (N = 26,578/115,156) of those had co-morbid depression. We included individuals (n=2,314) who sent at least one message between CVD and depression diagnoses. Participants were mostly 65 years and older (n = 1,718/2,314, 74.2%), or of non-Hispanic ethnicity (N = 2,078/2,314, 89.9%), or of the White race (N = 1,506/2,314, 66.1%), but sex was balanced (females, N = 1,197/2,314, 51.7%). PPV was 51.2% [95% CI: 47.5-54.5%] Llama 3.1 8B, and sensitivity was 83.6% [81.1-85.9] Llama 3.1 8B and 71.0% [67.0-75.3] MedGemma 4B. On average, the LLM (Llama 3.1 8B) detected depression 660 days earlier than the first charted diagnosis over a 1,746-day assessment period, a typical timeline from CVD to depression diagnoses in our cohort. Conclusion/Relevance LLMs identified individuals with depression significantly earlier than official diagnosis among patients with CVD, relying solely on longitudinal patient messages without additional medical information, with high sensitivity and Patient Health Questionnaire-9 comparable PPV. This novel approach is applicable to various diagnoses. ### Competing Interest Statement In the last 3 years, C.Rodriguez has served as a consultant for Biohaven Pharmaceuticals, Osmind, and Biogen; and receives research grant support from Biohaven Pharmaceuticals, a stipend from American Psychiatric Association Publishing for her role as Deputy Editor at The American Journal of Psychiatry, and book royalties from American Psychiatric Association Publishing. F. Rodriguez reports consulting fees from Novartis, NovoNordisk, Esperion Therapeutics, Movano Health, Kento Health, Inclusive Health, Edwards, Arrowhead Pharmaceuticals, HeartFlow, iRhythm, Amgen, and Cleerly Health outside the submitted work. ### Funding Statement Kim is supported by the NIH (K01MH137386). Linos is supported by the NIH (grants R01AR082109 and K24AR075060). F. Rodriguez was funded by grants from the NIH National Heart, Lung, and Blood Institute (R01HL168188; R01HL167974; R01HL169345). The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Stanford University Institutional Review Board approved this study. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data (patient messages) used in this study is not publicly sharable.
OBJECTIVES:To quantify utilization and impact on documentation time of a large language model-powered ambient artificial intelligence (AI) scribe. MATERIALS AND METHODS:This prospective quality improvement study was conducted at a large academic medical center with 45 physicians from 8 ambulatory disciplines over 3 months. Utilization and documentation times were derived from electronic health record (EHR) use measures. RESULTS:The ambient AI scribe was utilized in 9629 of 17 428 encounters (55.25%) with significant interuser heterogeneity. Compared to baseline, median time per note reduced significantly by 0.57 minutes. Median daily documentation, afterhours, and total EHR time also decreased significantly by 6.89, 5.17, and 19.95 minutes/day, respectively. DISCUSSION:An early pilot of an ambient AI scribe demonstrated robust utilization and reduced time spent on documentation and in the EHR. There was notable individual-level heterogeneity. CONCLUSION:Large language model-powered ambient AI scribes may reduce documentation burden. Further studies are needed to identify which users benefit most from current technology and how future iterations can support a broader audience.
Importance:Leveraging technology to prompt team-based care might improve ambulatory hypertension care. Objective:To assess whether an electronic medical record (EMR) high blood pressure (BP) advisory improves hypertension control. Design, Setting, and Participants:This quality improvement study assessed hypertension control in patients presenting to primary care office visits from March 2018 to February 2020. Data were included from 28 primary care clinics (8 clinics contributed data toward the primary objective and 28 contributed data toward secondary objectives) in a single academic health system in California before and after intervention and concurrent care team observations and interviews assessing implementation. Data were analyzed from November 2019 to October 2020. Intervention:An EMR high BP advisory combined with team training, audit, and feedback. EMR entry of elevated BP (systolic BP ≥140 mm Hg or diastolic BP ≥90 mm Hg) prompted an interruptive medical assistant-facing advisory to recheck BP. Persistently elevated BP prompted a second interruptive clinician-facing advisory with order panel link. Main Outcomes and Measures:The primary outcome was BP lower than 140 mm Hg systolic and lower than 90 mm Hg diastolic during an office visit within 6 months of an initial primary care visit. Secondary outcomes included BP recheck after initial elevated value, antihypertensive medication change, and new hypertension diagnoses. Qualitative outcomes focused on implementation barriers and facilitators. Results:The primary outcome assessed 2760 control patients and 3018 intervention patients with preexisting hypertension (mean [SD] age, 66.5 [14.4] years; 2847 [49.2%] women, 1746 [30.2%] Asian, 619 [10.7%] Hispanic, and 2407 [41.7%] White). The likelihood of hypertension control increased 18.3% per month on average (odds ratio [OR], 1.18; 95% CI, 1.10-1.27; P < .001) in the intervention vs control groups. Modeled rates of adjusted hypertension control over 6 months increased from 82.3% to 92.3% for the intervention cohort and decreased from 71.5% to 70.3% for the control (preintervention) cohort. BP recheck rate increased (from 37.6% to 77.9%; OR, 4.76; 95% CI, 4.45-5.10; P < .001), while ordered antihypertensive medications was unchanged. New hypertension diagnosis increased from 12.1% to 20.6% (OR, 1.34; 95% CI, 1.13-1.58; P = .01). In interviews of 34 care team members (clinicians, medical assistants, and managers) from 6 clinics, implementation barriers included competing priorities and time for BP rechecks, order panel complexity, and mixed clinician engagement; facilitators included intervention visibility, EMR integration, and team-based approach. Conclusions and Relevance:This quality improvement study of an EMR high BP advisory intervention found significantly improved primary care hypertension control and diagnosis due to the combination of team-based care and technology.
e23345 Background: Telemedicine enables patients to attend visits remotely and is associated with shorter wait times (i.e. scheduling-to-appointment duration). However, its impact on patient–clinic distance and wait times for new patients in high-volume surgical oncology centers remains unclear. In these centers, where appointment availability and travel distance significantly influence care decisions, understanding telemedicine’s impact on access is essential as federal policies supporting telemedicine continue to be under legislative review. To address this, we examined the relationship between new patient visit modality, distance, and wait time across demographics, providing insights to guide policy makers, and clinical practice. Methods: We extracted visit data from a multispecialty surgical oncology center in Northern California from 2019 to 2023 and used linear models with fixed effects for clinician and patient characteristics to assess the association of NPV modality with distance and wait times. We calculated distance using geodesic measurements between residential ZIP codes and care sites and derived median household income and education levels from the American Community Survey. Results: There were 10596 NPVs in 2019, and 13982 in 2023, conducted by 78 and 100 clinicians respectively. The proportion of NPV visits conducted by telemedicine increased from 0% to 28% during this period. From 2021–2023, NPVs conducted by telemedicine were more likely to fall outside the center’s 2019 catchment area (80th percentile of distance) with an adjusted odds ratio of 1.89 (95% CI 1.62, 2.21). Wait times for telemedicine NPVs was 3.0 days shorter (95% CI [2.0, 3.9]) than in-person visits. Compared to the reduction in wait times for patients aged 45–65 (2.8 days, 95% CI [1.8, 3.8]), those 80 and older had a smaller reduction (1.9 days, 95% CI [0.9, 2.8]), while patients aged 65–79 had a greater reduction (3.4 days, 95% CI [2.4, 4.5]). This effect did not significantly vary by sex, race, ethnicity, education, income, or insurance status. Conclusions: At this multispecialty surgical oncology center, new patients seen via telemedicine were more likely to reside outside the center’s traditional catchment area and had shorter wait times, with these differences varying by demographic characteristics. Strategic telemedicine integration tailored to patient needs could enhance access by reducing wait times and expanding specialized surgical oncology care to patients across broader geographic distances.
Telemedicine is now a sustained modality of ambulatory surgical oncology care, yet its association with workforce utilization, patient volume, and visit type at high-volume academic centers remains understudied. Characterizing these patterns is essential for guiding clinical operations and long-term integration of telemedicine into surgical oncology practice. We conducted a retrospective cohort study across nine oncology subspecialties at Stanford Medicine’s ambulatory surgical oncology clinics from January 2019 to December 2023 to compare yearly visit volumes and telemedicine use. The study included a total of 231,746 visits, including 50,667 new and 181,079 return visits. We measured overall visit volumes, telemedicine utilization, and their association with increase in unique patients served, including both new and return visits. In 2023, visit volumes increased by 44
BACKGROUND:Cardiovascular disease (CVD) remains the leading cause of death worldwide, yet many web-based sources on cardiovascular (CV) health are inaccessible. Large language models (LLMs) are increasingly used for health-related inquiries and offer an opportunity to produce accessible and scalable CV health information. However, because these models are trained on heterogeneous data, including unverified user-generated content, the quality and reliability of food and nutrition information on CVD prevention remain uncertain. Recent studies have examined LLM use in various health care applications, but their effectiveness for providing nutrition information remains understudied. Although retrieval-augmented generation (RAG) frameworks have been shown to enhance LLM consistency and accuracy, their use in delivering nutrition information for CVD prevention requires further evaluation. OBJECTIVE:To evaluate the effectiveness of off-the-shelf and RAG-enhanced LLMs in delivering guideline-adherent nutrition information for CVD prevention, we assessed 3 off-the-shelf models (ChatGPT-4o, Perplexity, and Llama 3-70B) and a Llama 3-70B+RAG model. METHODS:We curated 30 nutrition questions that comprehensively addressed CVD prevention. These were approved by a registered dietitian providing preventive cardiology services at an academic medical center and were posed 3 times to each model. We developed a 15,074-word knowledge bank incorporating the American Heart Association's 2021 dietary guidelines and related website content to enhance Meta's Llama 3-70B model using RAG. The model received this and a few-shot prompt as context, included citations in a Context Source section, and used vector similarity to align responses with guideline content, with the temperature parameter set to 0.5 to enhance consistency. Model responses were evaluated by 3 expert reviewers against benchmark CV guidelines for appropriateness, reliability, readability, harm, and guideline adherence. Mean scores were compared using ANOVA, with statistical significance set at P<.05. Interrater agreement was measured using the Cohen κ coefficient, and readability was estimated using the Flesch-Kincaid readability score. RESULTS:The Llama 3+RAG model scored higher than the Perplexity, GPT-4o, and Llama 3 models on reliability, appropriateness, guideline adherence, and readability and showed no harm. The Cohen κ coefficient (κ>70%; P<.001) indicated high reviewer agreement. CONCLUSIONS:The Llama 3+RAG model outperformed the off-the-shelf models across all measures with no evidence of harm, although the responses were less readable due to technical language. The off-the-shelf models scored lower on all measures and produced some harmful responses. These findings highlight the limitations of off-the-shelf models and demonstrate that RAG system integration can enhance LLM performance in delivering evidence-based dietary information.
Patients with diabetes are at increased risk of comorbid depression or anxiety, complicating their management. This study evaluated the performance of large language models (LLMs) in detecting these symptoms from secure patient messages. We applied multiple approaches, including engineered prompts, systemic persona, temperature adjustments, and zero-shot and few-shot learning, to identify the best-performing model and enhance performance. Three out of five LLMs demonstrated excellent performance (over 90% of F-1 and accuracy), with Llama 3.1 405B achieving 93% in both F-1 and accuracy using a zero-shot approach. While LLMs showed promise in binary classification and handling complex metrics like Patient Health Questionnaire-4, inconsistencies in challenging cases warrant further real-life assessment. The findings highlight the potential of LLMs to assist in timely screening and referrals, providing valuable empirical knowledge for real-world triage systems that could improve mental health care for patients with chronic diseases.
Patients with diabetes are at increased risk of comorbid depression or anxiety, complicating their management. This study evaluated the performance of large language models (LLMs) in detecting these symptoms from secure patient messages. We applied multiple approaches, including engineered prompts, systemic persona, temperature adjustments, and zero-shot and few-shot learning, to identify the best-performing model and enhance performance. Three out of five LLMs demonstrated excellent performance (over 90 with Llama 3.1 405B achieving 93 approach. While LLMs showed promise in binary classification and handling complex metrics like Patient Health Questionnaire-4, inconsistencies in challenging cases warrant further real-life assessment. The findings highlight the potential of LLMs to assist in timely screening and referrals, providing valuable empirical knowledge for real-world triage systems that could improve mental health care for patients with chronic diseases.