Background:Older age is widely considered a risk factor for post-acute sequelae of SARS-CoV-2 infection (PASC), typically attributed to immunosenescence and inflammaging. However, whether this association reflects intrinsic biological ageing or accumulated comorbidity burden remains unclear, with implications for clinical risk stratification. Methods:We conducted a retrospective cohort study using the Precision PASC Research Cohort (P2RC) from Mass General Brigham, comprising 133,792 COVID-19 patients from 12 hospitals and 20 community health centres in Massachusetts (March 2020-May 2024). PASC was ascertained using a validated computational phenotyping algorithm. We used generalised estimating equations with cluster-robust variance to model PASC risk, causal mediation analysis to decompose age effects through comorbidity burden and acute severity, and specification curve analysis across 768 analytical specifications to assess robustness. Findings:After adjustment for comorbidity burden, each decade of age was associated with 6% lower odds of PASC (OR 0.94; 95% CI 0.93-0.95). Causal mediation analysis revealed that comorbidities accounted for 145% of the total age effect, indicating inconsistent mediation wherein age's direct protective effect was masked by its indirect harm through chronic disease accumulation. This protection was age-dependent: adults younger than 65 years retained robust resilience independent of comorbidities (ADE: -0.0042, p<0.001), whereas adults 65 years and older showed complete loss of this protection (ADE: +0.0020, p=0.14). Interpretation:Long COVID susceptibility is driven by physiological reserve rather than chronological age until approximately age 65, beyond which age-related protective mechanisms become exhausted. Risk stratification should prioritise comorbidity burden over birth year in younger adults. Funding:National Institute of Allergy and Infectious Diseases (NIAID).
The temporal sequence of clinical events is crucial in outcomes research, yet standard machine learning (ML) approaches often overlook this aspect in electronic health records (EHRs), limiting predictive accuracy. We introduce Temporal Learning with Dynamic Range (TLDR), a time-sensitive ML framework, to identify risk factors for post-acute sequelae of SARS-CoV-2 infection (PASC). Using longitudinal EHR data from over 85,000 patients in the Precision PASC Research Cohort (P2RC) from a large integrated academic medical center, we compare TLDR against a conventional atemporal ML model. TLDR demonstrated superior predictive performance, achieving a mean AUROC of 0.791 compared to 0.668 for the benchmark, marking an 18.4% improvement. Additionally, TLDR’s mean PRAUC of 0.590 significantly outperformed the benchmark’s 0.421, a 40.14% increase. The framework exhibited improved generalizability with a lower mean overfitting index (− 0.028), highlighting its robustness. Beyond predictive gains, TLDR’s use of time-stamped features enhanced interpretability, offering a more precise characterization of individual patient records. TLDR effectively captures exposure–outcome associations and offers flexibility in time-stamping strategies to suit diverse clinical research needs. TLDR provides a simple yet effective approach for integrating dynamic temporal windows into predictive modeling. It is available within the MLHO R package to support further exploration of recurrent treatment and exposure patterns in various clinical settings.
Objective:Accurate and scalable dementia phenotyping from electronic health records (EHRs) is foundational for population-level research, risk prediction, and learning health system interventions. Traditional rule- and keyword-based approaches are limited by inconsistent documentation and inability to capture clinical nuance. We aim to develop and evaluate a framework that leverages large language models (LLMs) with retrieval-augmented generation (RAG) to overcome these limitations and improve dementia identification from real-world EHR data. Methods:Using EHR data from the Mass General Brigham health system, we first assembled a cohort of adults with potential dementia based on diagnosis codes, problem lists, dementia-related medications, and free-text note mentions. A subset of candidate cases underwent detailed manual chart review to assign gold-standard dementia status. With this labeled sample, we implemented and compared three approaches for dementia ascertainment: (1) a rule-based classifier leveraging structured EHR data, (2) large language models (LLMs) applied to keyword-filtered clinical note excerpts, and (3) a RAG-based LLM framework that integrates retrieved, context-rich note snippets. Within each approach, we evaluated multiple configurations of embedding models, retrieval methods, LLMs, structured-data inclusion, and prompts to identify the best-performing classifier. Performance was assessed using standard classification metrics, including sensitivity, specificity, positive predictive value (PPV), and F1 score, and supplemented by qualitative error analyses to characterize common sources of false positives and false negatives across methods. Results:The RAG-based classifier achieved the highest performance (F1=0.933, sensitivity=91.1%, PPV=95.5%) compared to rule-based (F1=0.823, sensitivity=81.1%, PPV=83.5%) and keyword-filtered LLM (F1=0.903, sensitivity=91.7%, PPV=88.6%). Including ICD codes alongside free text in the RAG-based LLM pipeline significantly reduced the PPV and modestly decreased F-1 score. Error analysis revealed that structured-code dependence contributed to false positives, whereas unrecognized contextual cues in notes drove false negatives. Conclusion:A RAG-based LLM pipeline without structured ICD codes improved dementia ascertainment from EHR data compared with ICD-based rules and keyword-based filtering. This approach can enhance dementia case identification and support patient care, predictive modeling and risk analysis.
The Electronic Medical Records and Genomics (eMERGE) Network developed and implemented a genome-informed risk assessment (GIRA) to communicate genomic (polygenic risk scores [PRSs], integrated risk scores [IRSs], and monogenic results), clinical, and family history-based risk for 11 chronic diseases and provide recommended healthcare recommendations. GIRA reports have now been returned to 23,840 participants and their providers in a large prospective cohort study. We present here the study design and analysis framework for assessing the attributable impact of GIRA return. Pre-specified outcomes include (1) provider/participant adoption of recommended healthcare actions, (2) new diagnosis of disease, (3) treatment initiation/intensification, and (4) clinical outcomes (surrogate markers or clinical events). We assess outcomes in high risk vs. not-high-risk participants, adjusting for covariates. We evaluate the effect of PRS/IRS at pre-established high-risk thresholds using regression discontinuity (RD), a quasi-experimental method that mimics randomization near a cutoff, enabling estimation of causal effects and controlling for unobserved confounders. Monogenic and family history-based risk stratification are analyzed using logistic regression. With 23,840 participants and 12 months of follow-up, the study is powered to detect differences of 2%-11% with 80% power (α = 0.05 in the adoption outcome). Longer follow-up will be required to enable assessment of new disease diagnosis, treatment changes, and clinical outcomes. Through innovative RD analyses and defined outcomes and comparison groups, this study will provide new insights into the real-world clinical impact of genomic risk assessment, address critical evidence gaps, advance understanding of genomic medicine outcomes, and inform future research.
Objective:Federated research networks, like Evolve to Next-Gen Accrual of patients to Clinical Trials (ENACT), aim to facilitate medical research by exchanging electronic health record (EHR) data. However, poor data quality can hinder this goal. While networks typically set guidelines and standards to address this problem, we developed an organically evolving, data-centric method using patient counts to identify data quality issues, applicable even to sites not yet in the network. Materials and Methods:We distribute high-performance patient counting scripts as part of Integrating Biology at the Bedside (i2b2), which all ENACT sites operate. They produce counts of patients associated with ENACT ontology terms for each site. At the ENACT Hub, our pipeline aggregates site-contributed counts to produce network statistics, which our self-service web application, Data Quality Explorer (DQE), ingests to help sites conduct data quality investigation relative to the network. Results:Thirteen ENACT sites have contributed their patient counts, and currently ten sites have signed up to use DQE to analyze data quality issues. We announced a call to all ENACT sites to contribute additional patient counts. Discussion:Identifying site data quality problems relative to the network is novel. Using a metric based on evolving network statistics complements rigid data quality checks. It is adaptable to any network and has low barriers of entry, with patient counting being the sole requirement. Conclusion:We implemented a metric for conducting data quality investigation in ENACT using patient counting and network statistics. Our end-to-end pipeline is privacy-preserving and the underlying design is generalizable.
Early detection of cognitive impairment is limited by traditional screening tools and resource constraints. We developed two large language model workflows for identifying cognitive concerns from clinical notes: (1) an expert-driven workflow with iterative prompt refinement across three LLMs (LLaMA 3.1 8B, LLaMA 3.2 3B, Med42 v2 8B), and (2) an autonomous agentic workflow coordinating five specialized agents for prompt optimization. Using Llama3.1, we optimized on a balanced refinement dataset and validated on an independent dataset reflecting real-world prevalence. The agentic workflow achieved comparable validation performance (F1 = 0.74 vs. 0.81) and superior refinement results (0.93 vs. 0.87) relative to the expert-driven workflow. Sensitivity decreased from 0.91 to 0.62 between datasets, demonstrating the impact of prevalence shift on generalizability. Expert re-adjudication revealed 44% of apparent false negatives reflected clinically appropriate reasoning. These findings demonstrate that autonomous agentic systems can approach expert-level performance while maintaining interpretability, offering scalable clinical decision supports.
Polygenic risk scores are widely used for predicting genetic risk across complex diseases and traits, and several pre-trained models have been developed. Few approaches leverage these pre-trained polygenic risk scores to further refine predictive performance. Here, we present Adaptive Boosting of pre-trained Polygenic Risk Socres, a fine-tuning framework that refines pre-trained polygenic risk score models through adaptive variable selection and model boosting to identify additional predictive signals that may not be fully captured by the original models. Simulations show that our framework can identify signals orthogonal to pre-trained polygenic risk scores while controlling false discovery rates. Using UK Biobank data, we fine-tune pre-trained polygenic risk scores for binary diseases and continuous traits, and validate the results across three independent datasets: All of Us, eMERGE, and Penn Medicine Biobank. Real data analyses show that Adaptive Boosting of pre-trained Polygenic Risk Scores achieves statistically significant improvements in several scenarios while maintaining competitive performance in others. Several pre-trained PRS models exist, but they are not often used to refine prediction further. Here, the authors develop a pre-train and fine-tune framework to improve polygenic risk prediction using these pre-trained models.
Background Electronic health record (EHR) phenotyping underpins observational research, cohort discovery, and clinical trial screening. Large language models (LLMs) offer new capabilities for extracting phenotypes from unstructured text, but their performance depends on pipeline design choices-including prompting, text segmentation, and aggregation. No systematic framework has previously examined how these parameters shape accuracy and reproducibility. Methods We evaluated LLM-based phenotyping pipelines using 1,388 discharge summaries across 16 clinical phenotypes. A full factorial experiment with LLaMA-3B, 8B, and 70B systematically varied three pipeline components: prompting (zero-shot, few-shot, chain-of-thought, extract-then-phenotype), chunking (none, naive, document-based), and aggregation (any-positive, two-vote, majority), yielding 24 configurations per model. To compare intrinsic model capabilities, biomedical domain-adapted, commercial frontier (LLaMA-405B, GPT-4o, Gemini Flash 2.0), and reasoning-optimized models (DeepSeek-R1) were evaluated under a fixed configuration. Performance was assessed using precision, recall, and macro-F1; secondary analyses examined prediction consistency (Shannon entropy), self-confidence calibration, and the development of a taxonomy of recurrent model errors. Results Factorial ANOVAs showed that chunking and aggregation were the dominant drivers of performance, whereas the prompting strategy contributed minimally. Configuration effects were stable across model sizes, with no significant Model x Parameter interactions. Phenotype difficulty varied substantially (macro-F1 = 0.40-0.90), yet the highest-performing configuration-whole-document inference without aggregation-was consistent across phenotypes, as confirmed by mixed-effects modeling. In cross-model comparisons, DeepSeek-R1 achieved the highest macro-F1 (0.89), while LLaMA-70B matched GPT-4o and LLaMA-405B at substantially lower cost. Prediction entropy was low overall and driven primarily by phenotype difficulty rather than prompting or temperature. Self-confidence calibration was only moderately informative: high-confidence predictions were more accurate, but larger models exhibited systematic overconfidence. Conclusions LLM performance in EHR phenotyping is governed primarily by input structure and model capacity, not prompt engineering. Simple, document-level inference yields robust performance across diverse phenotypes, providing practical design guidance for LLM-based cohort identification while underscoring the continued need for human oversight for challenging phenotypes. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by funding from the Fonds de recherche du Quebec - Sante (FRQS) and the Association des specialistes en medecins internistes du Quebec (ASMIQ). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This work uses the publicly available MIMIC-IV dataset available to any researcher upon completing CITI training on http://physionet.org/sign-dua/mimiciv/2.2/ I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The MIMIC-IV Phenotype Atlas (MIPA) dataset labels are openly available at https://github.com/open-health-data-lab/MIPA-datacard All code for pipeline processing, supervised model training, and LLM-based phenotyping is available at https://github.com/open-health-data-lab/MIPA Note that access to the underlying discharge summaries and clinical notes is conditional upon completing CITI training and obtaining authorized access to the MIMIC-IV v2.2 database through PhysioNet.
OBJECTIVE:Long COVID (LC) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients. In this study, we use summarized symptom reports by RECOVER-Adult cohort participants linked to EHR data to characterize patients and train a computable phenotype algorithm of LC. MATERIALS AND METHODS:The study included adult participants with linked FHIR-sourced EHR data. We characterized EHR diagnoses, procedures, medications, lab tests, and vital sign features associated with LC. A computable phenotyping algorithm was trained and validated against patient-reported symptoms. MAIN OUTCOME AND MEASURES:We assessed model discrimination and calibration in a held-out test set. We describe important model features and evaluate model discrimination and calibration. RESULTS:The study included 1,501 RECOVER-Adult cohort participants with linked EHR data. 376 (25%) met criteria for highly symptomatic LC based on the RECOVER Long COVID Research Index (LCRI). EHR features associated with LC included clinician diagnosis of shortness of breath, malaise and fatigue, and cardiac dysrhythmias; documented treatment with albuterol, gabapentin, or duloxetine; or elevated heart rate. The algorithm identifying patients with highly symptomatic LC had an AUROC of 0.80 (95% confidence interval (CI) 0.74-0.85), and AUPRC of 0.58 (95% CI, 0.47-0.69). CONCLUSION AND RELEVANCE:These findings demonstrate that, using EHR data, a machine-learning model can accurately select patients with sets of self-reported LC symptoms. The model could help identify patients within a health system with the highest probability of the condition and facilitate screening, recruitment for clinical trials, and etiologic studies.
BACKGROUND:The impact of antiviral therapies, including Paxlovid, on post-acute sequelae of COVID-19 (PASC) remains inconclusive. METHODS:We analyzed data from 19,413 patients (age > 18) from a validated PASC research cohort in New England who experienced at least one COVID-19 infection episode between January 1, 2022, and June 7, 2022, totaling 22,094 episodes. Multivariable logistic regression with inverse probability weights was used to infer the causal effects of Paxlovid treatment during acute infection and the risk of PASC overall (primary outcome), stratified by age group and organ system. RESULTS:Across all age groups, Paxlovid shows no statistically significant effect in lowering overall PASC risk. Stratification by organ system reveals a 37% reduction in gastrointestinal PASC (OR: 0.63; 95% CI: [0.468, 0.850]; p < 0.05) but a 97.4% increase in the risk of eye and ear-related PASC (OR: 1.974; 95% CI: [1.048, 3.718]; p < 0.05). Among patients aged 65 to 75 years who were not hospitalized, Paxlovid is associated with a 16.8% reduction in PASC risk (OR: 0.832; 95% CI: [0.7, 0.989]; p < 0.05). No statistically significant effects is observed for other organ-specific outcomes. CONCLUSIONS:Paxlovid demonstrates organ-specific effects on the risk of PASC, with a reduction in gastrointestinal symptoms and an increased risk of eye and ear-related symptoms. In older, non-hospitalized patients, Paxlovid modestly reduces overall PASC risk. These findings highlight the complexity of antiviral therapy's long-term impact and underscore the need for further research to clarify the mechanisms underlying these outcomes.
Importance:Surveillance of postacute sequelae of SARS-CoV-2 infection (PASC) depends on diagnostic coding systems that capture fewer than one-half of affected individuals, rendering millions invisible to health systems and policymakers. Objective:To quantify the gap between true PASC burden and diagnostic code-based estimates, determine the proportion representing chronic disease, and characterize organ system heterogeneity and temporal trends across diverse populations. Design, Setting, and Participants:This retrospective cohort study used electronic health record data from 58 hospitals and affiliated clinics in 4 US regions, from 2017 to 2025. Adults (aged ≥18 years) with laboratory-confirmed SARS-CoV-2 infection or a COVID-19 diagnosis code were included. A custom artificial intelligence algorithm, the Precision Phenotyping for Research Cohorts (P2RC), was implemented using federated infrastructure. Exposure:Laboratory-confirmed SARS-CoV-2 infection or COVID-19 diagnosis code. Main Outcomes and Measures:The primary outcomes were PASC prevalence, the proportion classified as chronic conditions, organ system distribution, and temporal trends from 2020 to 2024. χ2 Tests were used to assess organ system heterogeneity across regions, and negative binomial regression was used to model quarterly temporal trends, yielding incidence rate ratios (IRRs) with 95% CIs. Results:In this cohort study of 457 950 COVID-19 cases (mean age, 52.05 years; 275 107 [60.07%] female), the P2RC algorithm identified 74 560 PASC cases (16.28% overall; 28 585 [18.58%] in New England, 978 [19.55%] in Southeast Texas, 10 534 [22.69%] in Southern California, and 34 463 [13.64%] in Western Pennsylvania), more than 2-fold higher than the proportion identified by code-based surveillance (<7%). Of 883 International Statistical Classification of Diseases, Tenth Revision, Clinical Modification codes associated with PASC, 594 (67.27%) represented chronic or potentially chronic conditions. Of 74 560 patients with PASC, 66 587 (89.31%) developed chronic conditions requiring ongoing clinical management; this represents 14.54% of the total number of 457 950 patients with COVID-19. Substantial organ system heterogeneity was observed (χ2 = 2504.73; P < .001): New England demonstrated thyroid-predominant endocrine patterns, while Southeast Texas, Southern California, and Western Pennsylvania showed metabolic-predominant profiles. Negative binomial regression revealed increasing PASC prevalence through mid-2024 (IRR per quarter, 1.01 [95% CI, 1.00-1.01; P < .001] in New England; 1.00 [95% CI, 1.00-1.01; P < .001] in Southern California; and 1.02 [95% CI, 1.01-1.02; P < .001] in Western Pennsylvania), indicating an accumulating rather than resolving burden. Conclusions and Relevance:In this cohort study, approximately 1 in 6 patients with COVID-19 developed PASC, and 89.31% of these patients had at least 1 chronic condition. Current diagnostic coding captured fewer than one-half of the cases, obscuring a substantial chronic disease burden. The persistently increasing prevalence through 2024 indicated an accumulating health care burden requiring investment in surveillance infrastructure and integrated care pathways.
Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. Existing extraction approaches either rely on context-insensitive rules that generate false positives or on supervised models that require substantial fine-tuning. We present Pythia, a multi-agent system that autonomously writes and optimizes extraction prompts for clinical concepts without manual prompt engineering or fine-tuning. Running on a locally hosted open-weights model, Pythia keeps clinical notes on local infrastructure and selects prompts using development-set sensitivity and specificity. We compared Pythia with a curated lexicon across 72 signs and symptoms from 400 clinical notes representing 387 patients. Development (n=300) and validation (n=100) sets were partitioned independently for each concept. Pythia achieved mean sensitivity of 0.76 and specificity of 0.95, compared with 0.82 and 0.76 for the lexicon, and matched or exceeded the lexicon on both metrics for 20 of 62 directly comparable concepts. For 14 concepts where the lexicon labeled every note positive, Pythia recovered mean specificity of 0.97 by requiring a present-tense, patient-attributed finding rather than any textual mention of a term. Specificity transferred from development to validation with minimal degradation across prevalences, whereas sensitivity transfer weakened below 5
Objective:Although machine learning (ML) holds significant potential to transform healthcare, there has been a recent surge in research output that often lacks methodological rigor, contributing to a reproducibility crisis. Additionally, the growing reliance on electronic health records (EHR) for developing ML models has heightened concerns about patient data privacy. To tackle these challenges, we have extended the open-source i2b2 (Informatics for Integrating Biology and the Bedside) platform to allow researchers to train and run ML models without requiring manual programming or direct access to patient-level data. Materials and Methods:We have developed a proof-of-concept ML module for the i2b2 platform for creating and executing ML models. We describe the design of the module and demonstrate its use on a publicly available Kaggle dataset. Next we test its scalability on a large real-world dataset. Results:Model training with EHR of 100,000 patients randomly selected from the MIMIC-IV dataset was completed in 75.8 minutes and the developed model was applied to classify 28,985 patients in 1.61 minutes. Discussion:Implementation of the ML functionalities of the i2b2-ML module was successfully evaluated with a publicly available dataset. The developed module allows seamless training and execution of ML models without the need for manual programming and export of patient-level data, thus addressing many of the challenges associated with data privacy and reproducibility. Conclusion:In summary, the developed i2b2-ML module can reduce the technical overhead for researchers for applying ML to health data. Future work will focus on improving the i2b2 graphical interface to further simplify the use of the ML module and on streamlining the distribution of the ML module to existing i2b2 installations, so researchers can more easily analyze EHR data that exists in their current installations.
Early identification of cognitive concerns is critical but often hindered by subtle symptom presentation. This study developed and validated a fully automated, multi-agent AI workflow using LLaMA 3 8B to identify cognitive concerns in 3,338 clinical notes from Mass General Brigham. The agentic workflow, leveraging task-specific agents that dynamically collaborate to extract meaningful insights from clinical notes, was compared to an expert-driven benchmark. Both workflows achieved high classification performance, with F1-scores of 0.90 and 0.91, respectively. The agentic workflow demonstrated improved specificity (1.00) and achieved prompt refinement in fewer iterations. Although both workflows showed reduced performance on validation data, the agentic workflow maintained perfect specificity. These findings highlight the potential of fully automated multi-agent AI workflows to achieve expert-level accuracy with greater efficiency, offering a scalable and cost-effective solution for detecting cognitive concerns in clinical settings.
BACKGROUNDPrevious epidemiologic studies of autoimmune diseases in the US have included a limited number of diseases or used metaanalyses that rely on different data collection methods and analyses for each disease.METHODSTo estimate the prevalence of autoimmune diseases in the US, we used electronic health record data from 6 large medical systems in the US. We developed a software program using common methodology to compute the estimated prevalence of autoimmune diseases alone and in aggregate that can be readily used by other investigators to replicate or modify the analysis over time.RESULTSOur findings indicate that over 15 million people, or 4.6% of the US population, have been diagnosed with at least 1 autoimmune disease from January 1, 2011, to June 1, 2022, and 34% of those are diagnosed with more than 1 autoimmune disease. As expected, females (63% of those with autoimmune disease) were almost twice as likely as males to be diagnosed with an autoimmune disease. We identified the top 20 autoimmune diseases based on prevalence and according to sex and age.CONCLUSIONHere, we provide, for what we believe to be the first time, a large-scale prevalence estimate of autoimmune disease in the US by sex and age.FUNDINGAutoimmune Registry Inc., the National Heart Lung and Blood Institute, the National Center for Advancing Translational Sciences, the Intramural Research Program of the National Institute of Environmental Health Sciences.
Recent advances in computing technology and the development of data utilization environments have rapidly accelerated the application of artificial intelligence in clinical research and healthcare. This review provides a comprehensive overview of current machine learning techniques for analyzing clinical data, with illustrative examples from the field of allergic diseases. In addition to conventional methods for clinical data analysis, we discuss emerging approaches including medical image analysis and time-series modeling of electronic health record data. Recent developments such as large language models and foundation models trained on massive datasets are also discussed. Looking ahead, we explore future directions in analytical methodology, including mathematical modeling, interpretable artificial intelligence, and multimodal learning that integrates various data types. We also introduce the concept of the digital twin-a virtual representation of an individual patient that simulates disease progression and treatment response-as a promising concept for advancing precision medicine. Finally, we discuss the essential role of physicians in the development and implementation of machine learning tools and discuss emerging ethical issues such as fairness, privacy, and patient autonomy. By synthesizing recent technical advances with clinical relevance, this review aims to provide clinicians and researchers with a practical and forward-looking guide to machine learning in clinical medicine, including its growing application in the field of allergy.
BACKGROUND:Incidence estimates of post-acute sequelae of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, also known as long COVID, have varied across studies and changed over time. We estimated long COVID incidence among adult and pediatric populations in 3 nationwide research networks of electronic health records (EHRs) participating in the RECOVER (Researching COVID to Enhance Recovery) Initiative using different classification algorithms (computable phenotypes). METHODS:This EHR-based retrospective cohort study included adult and pediatric patients with documented acute SARS-CoV-2 infection and 2 control groups: contemporary coronavirus disease 2019 (COVID-19)-negative and historical patients (2019). We examined the proportion of individuals identified as having symptoms or conditions consistent with probable long COVID within 30-180 days after COVID-19 infection (incidence proportion). Each network (the National COVID Cohort Collaborative [N3C], National Patient-Centered Clinical Research Network [PCORnet], and PEDSnet) implemented its own long COVID definition. We introduced a harmonized definition for adults in a supplementary analysis. RESULTS:Overall, 4% of children and 10%-26% of adults developed long COVID, depending on computable phenotype used. Excess incidence among SARS-CoV-2 patients was 1.5% in children and ranged from 5% to 6% among adults, representing a lower-bound incidence estimation based on our control groups. Temporal patterns were consistent across networks, with peaks associated with introduction of new viral variants. CONCLUSIONS:Our findings indicate that preventing and mitigating long COVID remains a public health priority. Examining temporal patterns and risk factors for long COVID incidence informs our understanding of etiology and can improve prevention and management.
Machine learning in medicine is typically optimized for population averages. This frequency weighted training privileges common presentations and marginalizes rare yet clinically critical cases, a bias we call the average patient fallacy. In mixture models, gradients from rare cases are suppressed by prevalence, creating a direct conflict with precision medicine. Clinical vignettes in oncology, cardiology, and ophthalmology show how this yields missed rare responders, delayed recognition of atypical emergencies, and underperformance on vision-threatening variants. We propose operational fixes: Rare Case Performance Gap, Rare Case Calibration Error, a prevalence utility definition of rarity, and clinically weighted objectives that surface ethical priorities. Weight selection should follow structured deliberation. AI in medicine must detect exceptional cases because of their significance.
Background: Essential thrombocythemia (ET) is a rare, chronic myeloproliferative neoplasm characterized by sustained thrombocytosis, risk of thrombosis and bleeding, and progression to post-ET myelofibrosis (MF) or blast-phase disease. Cytoreductive therapy is recommended for thrombosis prevention in high-risk or symptomatic patients, but despite standard of care (SOC) therapies, many ET patients continue to experience disease complications, inadequate symptom control and suboptimal outcomes. Contemporary real-world studies are needed to evaluate disease burden and unmet treatment needs in ET. Aims: To characterize the treatment landscape and clinical outcomes among patients with ET treated with cytoreductive therapy in the U.S. Methods: Adult patients who met the following inclusion criteria were identified from the Mass General Brigham Research Patient Data Registry between January 1, 2012 and September 5, 2024: 1) an ET diagnosis code (by ICD10) associated with a hematologist/oncologist visit (i.e., first ET diagnosis); 2) a subsequent physician visit with an associated ET diagnosis code; 3) elevated platelet count (≥450x109/L) in the 6-months prior to first ET diagnosis; and (4) prescription for a SOC cytoreductive therapy on or after the first ET diagnosis. The prescription date of the first-observed cytoreductive agent (i.e., first-line [1L] therapy) was defined as the index date. 1L treatment discontinuation was defined as a switch to a different 2L therapy, or as a gap of >60 days following the end of 1L therapy unless a subsequent prescription of the same agent was observed after the 60 day gap (defined as treatment interruption). Clinical outcomes including thrombotic events and major bleeding events were summarized from the index date to the earliest of 1L treatment discontinuation, progression to MF or blast phase disease, or end of follow-up. Blast phase disease was defined by ≥1 acute myeloid leukemia diagnosis code. Major bleeding events were defined using modified ISTH criteria including those occurring in a critical area or organ in an inpatient setting, fatal bleeding (defined as bleeding in a critical area or organ preceding death by ≤45 days), or symptomatic bleeding (defined as bleeding leading to a drop in hemoglobin ≥2g/dL or red blood cell transfusion within 48 hours). Results: Among 673 patients who met the inclusion criteria, median (range) age at the index date was 70 (20 - 96) years, 67.3% were female, 88.0% were white. 14.1% had a history of thrombosis in the 6-months prior to 1L initiation. The mean (SD) follow-up time was 41.0 (31.3) months. The most common 1L treatment was hydroxyurea (94.1%); use of other treatments (e.g., interferon alfa [1.8%], ruxolitinib [1.5%], anagrelide [1.3%], and busulfan [0.3%]) was rare. 1L treatment modifications occurred frequently: 235 patients (34.9%) experienced treatment interruptions and 131 (19.5%) discontinued treatment. Among those who discontinued, median (interquartile range) time from 1L initiation to discontinuation was 14.1 (5.1 – 36.6) months; of these, 75 (57.3%) received no further treatment during our follow-up. Among all 1L treated patients, 56 (8.3%) switched to 2L therapy. Overall, 30.5% of patients experienced ≥1 ET-related complication after the index date, including thrombotic events (23.2%), major bleeding events (3.6%), and progression to MF (4.9%) or blast phase disease (2.1%). Sixty-six patients (9.8%) died of any cause. Among patients with available labs, 65.6% (368/561) of patients had an average platelet count ≥600 x 109/L in the first 6 months post-index, and 30.1% (137/455) during months 6-12. Conclusions: Despite SOC treatment with cytoreductive therapy, a significant proportion (30.5%) of ET patients experienced disease-related complications including thrombosis, disease progression, and major bleeding. These complications may be due in part to frequent therapy interruption or discontinuation (47.7% of patients), underscoring the limitations of SOC ET therapy. Our findings highlight ongoing challenges in ET management and the need for alternative therapeutic strategies to improve long-term outcomes in ET.