
Abstract Objective To extend the PyCox to accommodate time-varying covariates using counting process data, addressing immortal time bias in survival analysis. Materials and Methods We modified PyCox to support counting process data structures for time-varying covariates and applied it to 2,246,913 Medicare beneficiaries aged ≥65 with COVID-19 (January-September 2022; 2,785,807 longitudinal records). We compared traditional Cox regression, original PyCox, and modified PyCox in estimating associations between early antiviral treatment (nirmatrelvir or molnupiravir) and long-COVID. Performance metrics included concordance index (C-index), integrated Brier score (IBS), and integrated negative binomial log-likelihood (IBLL) and time-dependent Area Under the Receiver Operating Characteristic Curve (AUC), Brier score, and permutation importance. Results Among patients, 19.5% received nirmatrelvir, 2.6% received molnupiravir, and 14% developed long-COVID. Traditional Cox and modified PyCox produced concordant hazard ratios (HR) for nirmatrelvir (0.874 and 0.878) and molnupiravir (0.909 and 0.918); original PyCox estimated stronger associations (HR = 0.822 and 0.879). The treatment-effect difference formed a gradient: 3.5-4.0 percentage-points for time-varying models, 4.6 for time-fixed Cox, and 5.7 for time-fixed PyCox. Discrimination, calibration, and fit were comparable overall; modified PyCox showed higher time-dependent AUCs (0.536-0.609) than time-fixed PyCox (0.481-0.519) and relied on largely different top predictors. Modified and original PyCox trained in 1 minute 14 seconds and 3 minutes 18 seconds, versus 5 minutes for traditional Cox. Discussion Treatment-contrast magnitude and covariate importance varied by models, suggesting temporal covariate structure and modeling approach jointly influence treatment-effect estimates. Conclusion Extending PyCox to accommodate time-varying covariates improves computational efficiency while maintaining estimation accuracy comparable to standard time-varying methods.
Objective:To evaluate natural language processing (NLP) and machine learning (ML) approaches for identifying social needs in electronic health record (EHR) notes, using balanced versus real-world imbalanced datasets. Materials and Methods:A rule-based NLP framework using definite, non-negated keywords flagged social needs domains and "Any Social Need" in unstructured notes. Logistic regression, random forest, and XGBoost models were trained under balanced and imbalanced prevalence and evaluated using area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), F1, precision, recall, and specificity, with threshold sensitivity analyses. Results:On the imbalanced test set for any social need, the rule-based model achieved F1 0.373 (precision 0.484; recall 0.303; and specificity 0.978). On the balanced test set for any social need, XGBoost and logistic models achieved AUPRC 0.719 and 0.718, F1 0.438, and precision 0.944; random forest had AUPRC 0.714, F1 0.423, and precision 0.945. Under imbalanced evaluation, XGBoost led (AUPRC 0.319, F1 0.312, precision 0.839, recall 0.191, and specificity 0.998), followed by logistic (AUPRC 0.312 and F1 0.289) and random forest (AUPRC 0.305 and F1 0.282). Discussion:Balanced training inflated apparent performance, while under realistic class imbalance the ML models traded recall for very high precision, with XGBoost performing best but still lagging in sensitivity. Conclusion:These findings underscore the need to train and tune under real-world prevalence and to set thresholds based on use-case priorities (eg, maximizing precision for referrals versus improving recall for screening).
Objectives:Patients with substance misuse are at high risk for clinical deterioration, and pre-hospital encounters constitute important risk factors. We sought to incorporate these risk factors into a novel prediction model using linked electronic health record, emergency medical services (EMSs), and claims data. Materials and methods:Using 23 454 hospital encounters, we developed machine-learning models to predict mechanical ventilation, vasopressor initiation, or death within 12 hours of a vital sign or laboratory measurement. Models were compared to the modified early warning score. Results:Extreme gradient boosting achieved an area under the receiver operating characteristic curve of 0.93, significantly outperforming modified early warning score (MEWS 0.79). Notably, removing EMS and claims data did not reduce predictive performance. Discussion:The model demonstrated strong predictive performance of clinical deterioration in this high-risk population. Incorporation of pre-hospital risk factors did not improve performance beyond EHR data alone. Conclusion:Electronic health record-based machine learning models can support accurate short-term prediction of clinical deterioration among this population at risk for delayed recognition.
Objective:To examine how algorithmic fairness is measured, operationalized, and reported in machine learning (ML) models designed to predict or support secondary prevention of cardiovascular disease (CVD) outcomes including progression, recurrence, readmission, and post-index mortality in racialized populations. Materials and Methods:This scoping review was conducted in accordance with PRISMA-ScR guidelines and registered with the Open Science Framework (OSF registration: https://doi.org/10.17605/OSF.IO/9W67V). Systematic searches of OVID MEDLINE, EMBASE, Scopus, and CENTRAL were performed to identify studies evaluating fairness in ML models applied to secondary cardiovascular outcomes. Results:Of 2669 records screened, three retrospective cohort studies met inclusion criteria. All studies used large-scale electronic health record data from the United States and evaluated model performance across racial subgroups. Only one study implemented fairness-aware model development, reporting improvements of approximately 5%-12% in equity-related metrics, accompanied by modest trade-offs in calibration and sensitivity. The remaining studies assessed fairness post hoc and demonstrated limited ability to mitigate subgroup performance differences. Discussion:Most full-text studies excluded during screening addressed fairness in predicting primary CVD incidence rather than secondary outcomes, highlighting a substantial gap in the literature. Across included studies, observed fairness limitations appeared to be driven largely by upstream structural and data-generating factors such as representation, care patterns, and documentation rather than algorithmic design alone. Conclusion:Evidence on algorithmic fairness in ML models for secondary cardiovascular outcomes remains sparse. Improved reporting of subgroup performance, missingness, and calibration, alongside integration of fairness throughout model development, is necessary before equitable clinical deployment. Study Registration:Open Science Framework: https://doi.org/10.17605/OSF.IO/9W67V.
Introduction:Participation in cancer clinical trials is low in community oncology settings, partly because institution-specific trial information is fragmented. We evaluated feasibility of embedding curated trial content in an AI-enabled knowledge management application. Materials and Methods:At a regional community oncology network, coordinators and disease teams compiled actively recruiting trials. Core elements (title, conditions, biomarkers, stage/line, and recruiting status) were structured for point-of-care display and uploaded. AI-assisted extraction generated protocol summaries and eligibility elements, which underwent systematic human validation. Results:Fifty-three trials across 10 disease groups were embedded and validated; 91% were recruiting. Trials covered 28 cancer types; 30% were biomarker-specific and most enrolled advanced/metastatic disease. Initial configuration took 2-4 weeks per disease group using existing personnel, without added staffing or electronic health record (EHR) build. Discussion:Embedding institution-specific trial content within an AI-enabled knowledge application is feasible in community oncology using existing clinical and research infrastructure, establishing a prerequisite for future usability and implementation studies.
Objectives:Substance use disorder (SUD) care continues to be hindered by persistent gaps in health information exchange (HIE). This study examined people, organizational, regulatory, and technical barriers to SUD data sharing from the provider perspective and identified key opportunities to improve care through health record interoperability. Materials and Methods:Behavioral health providers (52% prescribers) representing 4 SUD treatment organizations in 14 US states participated in 11 focus groups (n = 31) and 5 validation interviews (n = 5) covering scenarios related to SUD Healthcare Effectiveness Data and Information Set (HEDIS) metrics. Thematic analysis, keyword searches, and workflow analysis [using Unified Modeling Language (UML) diagrams] were performed. Results:Incomplete data access at the point of care frequently necessitated manual exchange via fax, phone, and secure email. When available, HIE and Prescription Drug Monitoring Program (PDMP) data were helpful, but confusion regarding the release of SUD information requirements in the context of HIPAA (Health Insurance Portability and Accountability Act), 42 CFR (Code of Federal Regulations) Part 2 and state requirements was ubiquitous. UML modeling of 4 core SUD care scenarios rising from participants' HEDIS examples revealed 3 common data sharing subprocesses. Discussion:Fragmented data sharing workflows, and extensive use of non-electronic data sharing methods impact outcomes for individuals with SUD. Interoperable consent management, HIE and PDMP integration with EHRs, and harmonized privacy regulations are provider priorities. Conclusion:Mapping provider-identified barriers reveals technical, regulatory, and workflow breakdowns that limit SUD care coordination. Advancing electronic, consent-driven interoperability employing standards is critical to improving continuity of care while safeguarding patient privacy.
Objectives:Investigate how patients at a Federally-Qualified Health Center (FQHC) perform with, and perceive, complex telehealth tasks. Determine which complexity dimensions most affect patients' performance and perceptions, and compare this to researcher-evaluators' complexity walkthrough results. Materials and Methods:A novel complexity walkthrough inspection method was implemented by researcher-evaluators (n = 8) to identify theory-informed complexity dimensions in 6 FQHC-required telehealth tasks. Remote user testing (n = 24) where FQHC patients were observed performing tasks, then completed newly-developed complexity-focused interviews and surveys. Descriptive statistics regarding complexity dimension presence, task performance, cognitive load, and perceived difficulty were integrated with qualitative data analyzed using inductive and deductive coding. Results:Patients completed 33.9% of the required subtasks without issues. Cognitive load and perceived difficulty were high for 2 tasks. Complexity dimensions of ambiguity (unclear inputs/processes; new concepts/words) and relationship (context switching; deep navigational hierarchies) most affected patient-perceived difficulty. Patients spent twice as long as walkthrough evaluators on tasks, and encountered broader complexity dimensions: new concepts/words, and errors. Many patients ended tasks early, asserting that they would abandon them outside of a study or had previously done so. Discussion:Technology-mediated task complexity may explain some telehealth uptake inequities. The complexity dimensions that challenge patients extend known usability heuristics by enhancing their equity sensitivity. Complexity walkthroughs surface design patterns that challenge patients, but complexity-focused user testing with patients reveals additional difficulties. Findings support complexity reduction of tasks via structuring, familiar concepts/words, feature integration, and shallow/broad navigation. Conclusion:This paper's novel, theoretically-grounded complexity-focused methods and findings may inform future equitable design and evaluation of technology-mediated tasks for socioeconomically marginalized patients.
Objectives:Machine learning (ML) models are increasingly being developed to support healthcare delivery. However, concerns remain about their potential to perpetuate existing biases rooted in the data used to develop them. We aim to assess the impact of using race in predicting hospital admission probabilities for patients visiting the emergency department (ED). Materials and Methods:Data from the MIMIC-IV ED dataset were used to train2 ML models predicting hospital admission: one included race; the other did not. Differences in predicted admission probabilities were evaluated across racial groups under multiple validation conditions. Results:Including race as a model input was associated with meaningful differences in predicted admission probabilities for White (3.2%), Black (-1.5%), and Hispanic (-3.0%) patients, while minimal differences were observed for Asian (0.2%) and Other (0.5%) patients. These differences were associated with large Cohen's d effect sizes in the baseline model for White (d = 1.00), Black (d = -1.23), and Hispanic (d = -1.35) patients. After balancing racial group prevalence, the effects persisted for White (1.22%; d = 1.22) and Hispanic (-1.00%; d = -1.00) patients. Discussion:These findings suggest that race is associated with differences in ML-based hospital admission predictions for ED patients, underscoring the need for caution when incorporating race into clinical prediction models and the importance of rigorous bias assessment. Conclusion:As race was associated with predictions, there is a crucial need to address underlying social factors and the use of broader, more equitable clinical data for ML model training.
Objectives:To evaluate the comparative effectiveness of ambient documentation tools (ADTs) with distinct architectures (Tablet-Based Virtual Human-Assisted Ambient [Tool A], EHR-Integrated Ambient [Tool B], and Standalone Ambient [Tool C]) on provider efficiency, documentation burden, and productivity in primary care. Materials and Methods:We conducted a real-world comparative-effectiveness study of 163 primary care providers in a large integrated health system (January 2024-June 2025), analyzing 59 130 provider-days. Providers who did not adopt an ADT (n = 13) were excluded from comparative models. Exposure groups included Tool A (n = 65), Tool B (n = 68), and Tool C (n = 17). Primary outcomes were pajama time (hours of after-hours EHR activity per provider-day), visit closure within 2 days (proportion of encounters signed within 48 hours), proportion of manual note composition, and a composite opportunity score reflecting EHR efficiency (scaled 0-1). Outcomes were derived from daily provider-level EHR logs. Analyses used intention-to-treat and per-protocol frameworks with provider-clustered SEs and month fixed effects. Results:Compared with Tool B, Tool A was associated with increased pajama time (+0.022 hours of after-hours EHR activity per provider-day), higher manual note composition (+0.046 absolute proportion), lower timely visit closure (-0.120 [95% CI, -0.126,-0.115]), and reduced opportunity scores (-0.035[-0.044,-0.027]). Tool C reduced after-hours work versus Tool B (-0.007 hours per provider-day) and relative to both the reference tool and pre-exposure baseline levels (-0.055 hours per provider-day), but was associated with lower timely visit closure (-0.028[-0.036, -0.020]); opportunity scores were not significantly changed (-0.007[-0.018, 0.003]). Effects for Tool C correspond to approximately 30 fewer annual after-hours hours per provider. Conclusion:Tool choice materially shapes real-world outcomes, and tool-specific comparative evaluation is essential to guide evidence-based procurement and deployment decisions.
Objectives:Disease profiles and access to care differ between rural and urban populations. The aim of this study is to examine the association of rurality with electronic health record-derived disease profiles and heart failure risk. Materials and Methods:We conducted a retrospective observational study of 242 758 participants enrolled in the All of Us Research Program. Rurality status was estimated using 2010 Rural-Urban Commuting Area codes from participant home 3-digit zip code prefixes, and participants were grouped into rural, mixed, and urban arms. We compared the disease profiles with phenome-wide association studies and the risks of heart failure between arms. Results:The final cohort included 5581 participants in the rural arm, 62 287 in the mixed arm, and 174 890 in the urban arm. Compared to the urban setting, 70 phecodes were significantly enriched in the rural arm including obesity (odds ratio [OR], 1.42, 95% confidence interval [CI], 1.34-1.51) and related diagnoses. Cancer screening-, dermatological-, and eye care-related diagnoses were significantly depleted in the rural arm. The rural arm also showed a greater risk of heart failure (adjusted hazard ratio [HR], 1.20, 95% CI, 1.10-1.31). Discussion and Conclusion:With the national dataset in All of Us, we found that rural participants had significantly higher risks for obesity-related diseases and heart failure in this study. Depleted phecodes in the rural participants suggested a lack of access to cancer screening, stressing the potential importance of targeted cancer disease management in rural communities.
Background:Clinicians are increasingly using emoji in digital communication, but limited qualitative work has examined how they assess the appropriateness, risks, and benefits of this practice. Objective:To characterize clinicians' reported utilization of and attitudes toward emoji in clinical messaging. Design:Qualitative study using focus groups and a survey from August to October 2025. Participants:Twenty-nine clinicians at a large academic health system were recruited from 4 specialties and included physicians and advanced practice providers, genetic counselors, medical students, and other healthcare workers. Approach:Rapid qualitative analysis was used to identify both the benefits and drawbacks of emoji use in clinical messaging. Key results:Emoji were seen as clarifying tone, building rapport, softening directives, and reducing notification fatigue. However, participants also reported concerns about ambiguity, informality, and potential medicolegal risk due to messages' discoverability. Their assessment of appropriateness was context-dependent, including power hierarchies, personal familiarity, clinical gravity, generational differences, and the technical affordances of specific platforms. Emoji were viewed as more acceptable in informal contexts and among peers with whom they had established relationships, and as inappropriate in high-stakes and sensitive clinical contexts. Participants felt formal guidelines could be seen as condescending but suggested modifications to promote the clear and effective use of emoji in professional settings. Conclusions:Clinicians view emoji as useful but context-sensitive communication tools that should be used conscientiously. They disprefer universal, prescriptive guidelines. However, they promote explicit conversations about emoji use at onboarding and the modification of which emoji are available in local messaging platforms.
Objectives:Embedded pragmatic clinical trials (ePCTs) are conducted as part of routine clinical care and therefore use data collected from real-world data sources, such as electronic health record systems and administrative claims. A common approach for using these types of data across all phases of trial conduct is to create a computable phenotype (an explicitly defined data query data, including specified data types, data codes, and logical parameters) to capture patients with a clinical condition, exposure, symptom, characteristic, treatment, or outcome of interest. Materials and Methods:The Electronic Health Record (EHR) Core Working Group of the NIH Pragmatic Trials Collaboratory captured the experiences of investigators in developing, adapting, and applying computable phenotype definitions in their studies. Results:Four case studies describe different approaches to developing and using computable phenotypes in pragmatic trials. Discussion:We recommend: (1) developing computable phenotypes as part of a team with multiple areas of expertise; (2) appropriate validation; (3) dissemination of salient details regarding phenotype creation; and (4) capturing and reporting any modifications made to phenotypes during the conduct of the trial. Conclusion:A range of issues and decisions influence how computable phenotypes are developed and used in pragmatic clinical trials. Multidisciplinary teams that understand the context of the data, including the reason the data are collected and potential sources of bias, are best suited for the development and validation of phenotypes for ePCTs.
Backgrounds and Objectives:Childhood obesity and respiratory tract infections (RTIs) are 2 major global public health issues that frequently co-occur and are closely interrelated. Early detection of children with prior RTIs who are at high obesity risk is crucial for targeted interventions. This study integrates interpretable machine learning (ML) models and a deep learning network to develop an obesity risk prediction model in a large pediatric cohort. Methods:Cross-sectional data from 6509 children and adolescents aged 3-12 years with prior RTIs in Beijing and Tangshan were fed to 12 ML models to predict childhood obesity (versus children with normal weight). Bayesian optimization was applied to fine-tune model hyperparameters. Prediction performance was assessed using 8 metrics. Key predictive features were identified by SHapley Additive exPlanations (SHAP). The validity of the optimal ML model was verified by the sequential neural network model. Results:Of 12 ML models, LightGBM achieved the optimal performance (accuracy: 0.8844, area under the curve [AUC]: 0.9491). SHAP analysis identified 20 key predictors, including child age, paternal body mass index (BMI), maternal BMI, birth length, birthweight, gestational age, eating speed, complementary feeding initiation age, screen time, breastfeeding duration, bedtime, maternal age, nighttime sleep, outdoor activity, dental caries, sedentary time, food allergies, family history of diabetes, and sex. The deep learning sequence network model further validated the predictive value of these features (accuracy: 0.8023, AUC: 0.8117), and the SHAP-driven feature importance rankings were in close line with LightGBM. Conclusions:Our LightGBM-based model enables effective prediction of obesity risk in children aged 3-12 years with prior RTIs, and the key features identified can inform early screening and facilitate personalized interventions.
Objective:To develop a machine learning (ML) model to predict magnetic resonance imaging (MRI) appointment no-shows and test the effectiveness of a targeted phone call intervention in reducing the no-show rate. Materials and Methods:Outpatient MRI appointments from a university hospital (2015-2021) were used to train and compare Logistic Regression, Random Forest, and XGBoost models. The best-performing model underwent a 13-month prospective evaluation, followed by a 9-month intervention study where high-risk patients were randomized to a phone call reminder or a control group. Results:Our dataset included 38 141 appointments (16.6% no-show) in total. On a held-out validation dataset, XGBoost performed best (AUROC 0.65; AUPRC 0.29). In a prospective evaluation round (5882 appointments), performance decreased (AUROC 0.61; AUPRC 0.24), due in part to observed data drift in appointment types and reasons. The intervention study (274 patients called, 158 answered, 116 not reached; 460 control appointments) showed no overall change in no-show rate (no-show: 21.9% intervention vs 22.0% control; P = .48). Within the intervention group, answering the call was associated with lower no-show rates (18.4% of those reached vs 26.7% not reached). Discussion:Models identified high-risk patients, but translating predictions into attendance improvements proved challenging. Effectiveness was limited by data drift and intervention reach. Results suggest value in dual-prediction strategies (including potential intervention responsiveness) and continuous model monitoring. Conclusion:ML can identify likely no-shows, but reducing missed appointments requires robust, up-to-date models and interventions that reliably reach patients; more nuanced, targeted engagement may yield greater impact than phone calls alone.
Objective:To clarify how validation requirements should be specified for medical digital twins used in clinical decision support, particularly when such systems are intended to compare interventions, treatment timings, dosages, or sequential care strategies. Perspective:Medical digital twins are heterogeneous systems that may combine prediction, simulation, mechanistic modeling, machine learning, data assimilation, uncertainty quantification, and decision-support functions. Their evaluation should therefore be driven by their intended use rather than by a single definition of what a digital twin is. For digital twins used primarily for visualization, monitoring, or short-term forecasting, predictive accuracy, calibration, discrimination, and robustness may be the central validation targets. However, when digital twins are used to support intervention-oriented clinical decisions, retrospective accuracy under historical clinical practice is insufficient on its own. Key message:Intervention-oriented digital twins address action-conditioned questions: what is predicted to happen under specified alternative actions, assumptions, time horizons, and clinical contexts. Their validation should therefore extend beyond scalar performance metrics to include uncertainty representation, updating stability, robustness under regime change, action-regime validity, counterfactual consistency, clinically weighted error, and decision-level consequences. This requires drawing on established traditions in forecast verification, causal inference, uncertainty quantification, model verification and validation, decision theory, control theory, and post-deployment monitoring. The level of causal or mechanistic support required should match the clinical claim being made, whether at the genotype, phenotype, physiological, or care-process level. Conclusion:The scientific-instrument framing is proposed as a pragmatic validation lens for intervention-oriented digital twins, not as a universal definition of digital twins. It helps define the scope within which their outputs can support clinical reasoning. Medical digital twins should be accompanied by explicit validation statements specifying their target population, prediction horizon, supported interventions, uncertainty bounds, and known failure conditions.
Objectives:Mental disorders are very common among children and adolescents around the world. In Mongolia, the international Strengths and Difficulties Questionnaire (SDQ) is used to detect mental disorders in adolescents. In this article, different types of classification methods were compared to determine adolescent emotional and behavioral problems using the SDQ. Materials and Methods:Data were collected from teenagers, teachers, and parents in Govi-Altai Province, and the databases were created for each group. The teenager database was divided into 10 folds using cross-validation, and the models were developed using classification methods and evaluated using performance measures. The results were mainly analyzed using the Bayes model. Results:The teenagers have emotional and behavioral problems due to emotional and peer interactions, but they were at risk of developing disorders due to hyperactivity and behavioral changes. Conclusion:Future work will examine the difficulties encountered in the creation of emotional and behavioral problems among the teenagers involved in the study using a progressive classification method.
Objectives:To evaluate the real-world performance of a transformer-based natural language processing (NLP) system for extracting Social Determinants of Health (SDoH) from clinical notes, using survey-based SDoH measures a reference comparators. Materials and Methods:This study was conducted at the University of Florida Health in adults with at least 2 clinical encounters in the prior year. A research survey was completed by 1001 participants; sampling targeted 50% Black patients to support subgroup analyses. Comparative analyses were restricted to the 414 participants who also had Epic SDoH survey data and clinical notes available for NLP extraction. We compared concepts extracted by the SOcial DeterminAnts (SODA) NLP pipeline against the independently administered research survey, which served as the primary reference standard, and against the structured Epic-embedded SDoH questionnaire. Nine domains were evaluated: abuse, alcohol use, drug use, education, financial constraints, housing, physical activity, social cohesion, and transportation. Sensitivity, specificity, positive predictive value, negative predictive value, and F1 scores were calculated by domain. Results:The NLP pipeline more consistently aligned with negative survey responses than with patient-reported social needs, although performance was lower in some domains, especially alcohol use and financial constraints. Sensitivity was higher only for alcohol use (55%); the lowest values were for abuse (5%), drug use (0%), and financial constraints (16%). These results cannot be attributed to SODA extraction alone. They reflect some combination of social information not being recorded in clinical notes, content that was recorded but not extracted, and mismatches between extracted concepts and survey definitions, and the present analysis cannot separate these contributions. The 2 surveys agreed only modestly with each other, so no single instrument provides a definitive ground truth. Conclusion:The pipeline more consistently aligned with negative survey responses than with patient-reported social needs. Because the 2 surveys agreed only modestly, the reference itself is imperfect, and apparent NLP performance depends in part on which survey is used as the comparator. Apparent gaps in NLP performance reflect both how social risks are recorded in clinical notes and how patients disclose them across different survey settings, in addition to limits of the extraction pipeline. Improving documentation practices, integrating locally tuned large language models, and monitoring subgroup performance may all be needed to make SDoH identification tools reliably detect social needs across patient populations.
Objectives:We extracted a validated disease activity measure in rheumatoid arthritis (RA), the Clinical Disease Activity Index (CDAI), from a large tertiary academic medical center electronic health record (EHR) using an automated large language model (LLM)-based approach without requiring model pretraining. Materials and Methods:The New York Presbyterian/Columbia University Medical Center Clinical Data Warehouse contains EHR data for over 4.5 million patients. RA patients were identified using International Classification of Disease-9 (ICD-9) and ICD-10 codes. Expert-curated CDAI keywords were extracted from unstructured notes using an automated natural language processing (NLP) pipeline leveraging GPT-4o API, a HIPAA-compliant, institutionally approved LLM platform. Performance was evaluated against expert chart review. Results:Among 2756 RA patients with notes, 1038 (37.7%) were seropositive, 796 (28.9%) were seronegative, and 922 (33.4%) had unknown serostatus. Clinical Disease Activity Index and its components were extracted in 15.4% (160/1038) of seropositive patients indicating remission or low disease activity. Clinical Disease Activity Index documentation was more frequent among patients with multiple notes and among faculty, with high extraction accuracy (precision/recall/F1 = 0.97). Discussion:This represents the first attempt to employ a zero-shot, ChatGPT-powered LLM platform to extract RA disease activity measures from real-world EHR data. Although a low prevalence of documentation was noted, important distinctions were observed when patients were subgrouped by serostatus, level of training, and number of visits. Conclusion:An LLM-based pipeline accurately extracted CDAI from a single large academic EHR, revealing infrequent real-world documentation.
Objectives:Accurately predicting outcomes for critically ill cancer patients remains challenging. This study aimed to integrate biologically relevant iron metabolism markers into short-term mortality prediction and to develop an interpretable, externally validated model based on a tabular prior-data fitted network (TabPFN) for risk stratification in this population. Materials and Methods:We conducted a retrospective cohort study using data from Medical Information Mart for Intensive Care IV (MIMIC-IV) (training and internal validation) and eICU Collaborative Research Database (eICU-CRD) (external validation), including critically ill adult patients with cancer. Associations between iron metabolism markers (ferritin, serum iron, and total iron-binding capacity [TIBC]) and 30-day all-cause mortality were assessed using the Kaplan-Meier curves, multivariable Cox regression, and restricted cubic splines. A TabPFN-based prediction model was developed and compared with 8 other machine-learning algorithms, with the least absolute shrinkage and selection operator used for feature selection. Model performance was evaluated using the area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), calibration, and decision curve analysis. SHapley Additive exPlanation values provided model interpretability. Results:Among 1137 patients in the MIMIC-IV cohort, 293 (26.0%) died within 30 days. Elevated ferritin (hazard ratio [HR] = 1.18, 95% CI, 1.09-1.27) and decreased TIBC (HR = 0.92, 95% CI, 0.85-0.99) were independently associated with mortality and exhibited non-linear relationships. Serum iron showed no prognostic value. The TabPFN model achieved the best performance, with an AUROC of 0.865 (95% CI, 0.812-0.919) in the internal validation set and 0.772 (95% CI, 0.744-0.814) in the external validation set, along with good calibration (Brier score = 0.121) and clinical net benefit. SHapley Additive exPlanation analysis identified lymphocyte count, platelet count, lactate, and ferritin as the most influential predictors. Conclusion:Iron metabolism dysregulation-particularly altered ferritin and TIBC levels-has important prognostic value in critically ill cancer patients. By integrating these biomarkers with a TabPFN framework, this study provides an accurate, interpretable, and externally validated tool for 30-day all-cause mortality prediction that may support clinical decision-making.
Abstract Objective Drug repurposing is particularly challenging yet essential for rare diseases, where limited patient populations and scarce biomedical evidence hinder traditional discovery pipelines. This work presents a holistic machine learning approach for drug–disease link prediction, leveraging multiple heterogeneous sources including biomedical literature, structured databases, and textual descriptions of diseases. Materials and Methods Focusing on seven rare neuro-muscular disorders, we construct a biomedical knowledge graph from literature and open databases, to evaluate a suite of rule-based, graph neural network, and path-encoding models. An ensemble of the best-performing methods, further enriched with disease similarity features derived from text-based embeddings, is used to generate candidate treatments for each disorder. Results Experimental results show that established graph neural network approaches (CompGCN), and path encoding methods (Prime Adjacency Matrix framework), outperform other approaches in metrics like Mean Reciprocal Rank. The ensemble of the best-performing methods further improves those metrics, reaching MRR = 0.3145. A manual validation of top-ranked drugs from rare disease experts illustrates a high precision (>50%) for drugs that potentially treat a rare disorder or its symptoms. Discussion The lack of vast number of publications and known drug indications for rare neuro-muscular disorders sets serious challenges in identifying potential therapies and symptom-relievers. The ensemble predictor incorporates rule-based, graph neural networks and path encoding techniques, to improve drug repurposing prediction performance on a biomedical knowledge graph created from open data. Conclusion Expert evaluation indicates that an ensemble of various knowledge graph link prediction methods can produce promising repurposing hypotheses, for disorders lacking any approved therapies.