Hybrid controlled trials (HCTs) incorporate real-world data into randomized controlled trials (RCTs) by augmenting the internal control arm with patients receiving the same treatment in routine care. Beyond increasing power, HCTs may improve recruitment by supporting unequal randomization ratios that increase patient access to experimental treatments. However, HCT validity is threatened by bias from unmeasured confounding due to lack of randomization of external controls, leading to outcome non-exchangeability between internal and external control patients. To address this challenge, we developed a sensitivity analysis framework to assess the robustness of HCT results to potential unmeasured confounding. We propose a tipping point analysis that adapts the E-value framework to the HCT setting where trial participation rather than treatment assignment is subject to confounding. To aid interpretation, we also introduce a data-driven benchmark representing the strength of unmeasured confounding reflected by the observed outcome non-exchangeability. We then propose an operational decision rule and evaluate its performance through simulation studies. Finally, we illustrate the approach using an asthma trial augmented by data from electronic health records. Simulation results demonstrate that our decision rule safeguards against Type I error inflation while preserving the power gains achieved by incorporating external data. In settings where moderate unmeasured confounding led to poorer outcomes for external controls, Type I error was controlled near the nominal 5% level, and power increased by 10-20% compared with analyses using RCT data alone. Our approach provides a practical, interpretable method to assess HCT robustness, supporting rigorous inference when integrating external real-world data.
Medical image foundation models can predict clinical phenotypes from computed tomography (CT), but strong performance leaves open whether they read disease-specific findings or shortcuts that correlate with the diagnosis. We tested this in 221 electronic-health-record (EHR) phenotypes using Auditable CT phenotyping (ACT), built on report-derived radiological observations. We trained ACT on 38,317 patients, mined 376,194 observations and evaluated it in 25,183 held-out patients. ACT exceeded five vision-language baselines on zero-shot annotation, and CT-CLIP across 221 phenotypes from unseen CT pulmonary angiography, both under zero-shot scoring (0.651 versus 0.572) and under linear probing (0.709 versus 0.662). Reading each probe exposes what accuracy conceals: only 97 observations occupy the 221 rank-1 positions, and one phrase describing aortic and coronary calcification ranks first for 20 phenotypes, including osteoporosis, urinary tract infection and major depressive disorder. Restricting the bank to clinician-specified evidence redirects those probes onto phenotype-related observations in 86 phenotypes at no accuracy cost (0.751 versus 0.741). Accurate CT-based EHR phenotyping can therefore rest on observations that are not valid evidence for the coded phenotype and that ACT can identify and intervene on.
A retrospective, exploratory cross-sectional analysis exploring whether social media data is associated with cardiovascular disease (CVD) risk beyond traditional clinical models. While social media data may capture behavioral and social markers relevant to CVD, their associations with CVD risk remains uncertain.
PURPOSE:Approximately 8% of the US population speaks primary languages other than English. Limited English proficiency (LEP) contributes to under-representation of Hispanic patients in oncology clinical trials. Although certified translation services exist, they are time-consuming and costly. Artificial intelligence (AI)-generated translations of informed consent forms (ICFs) could provide low-cost alternatives, but data on accuracy and safety remain limited. We evaluated language equivalence of English-to-Spanish translations for three oncology clinical trial ICFs using two general-purpose AI translation tools (DeepL Pro and ChatGPT-4o) and a medically trained AI translation tool (Med_English2Spanish) compared with certified translations. METHODS:Translational equivalence was assessed using a five-point Likert scale on five domains: Semantic, Idiomatic, Experiential, Conceptual, and Safety. Two native Spanish-speaking bilingual board-certified physicians independently scored each translation. Weighted Cohen's kappa determined inter-rater reliability, and the two-sample t-test compared AI-generated and certified translations. RESULTS:Weighted Cohen's kappa (0.95, 95% CI 0.85 to 0.97) exhibited high inter-rater agreement. Certified translations exhibited the highest equivalence (mean = 4.99, SD = 0.02). ChatGPT-4o similarly demonstrated high equivalence (mean = 4.89, SD = 0.17). DeepL Pro scored well (mean = 4.43, SD = 0.07) but lower than certified translation (P < 0.001). Med_English2Spanish demonstrated the lowest degree of equivalence (mean = 3.32, SD = 0.40) compared with certified translations (P < 0.001). CONCLUSION:Low-cost AI translations of ICFs exhibited variable language equivalence compared with certified translations across several domains. ChatGPT-4o scored nearly equivalent across domains in translating procedural trial information. While AI-generated translations are currently not suitable for clinical deployment without human review, this exploratory study supports further analysis of AI translation tools for reducing language barriers to LEP population enrollment.
OBJECTIVES:Despite historically limited adoption of prone positioning, a potentially life-saving guideline-recommended intervention for moderate-severe acute respiratory distress syndrome, its use increased for mechanically ventilated patients during the COVID-19 pandemic. Whether implementation of this guideline-recommended intervention was sustained is unknown. Thus, we aimed to evaluate peri-pandemic trends in proning use. DESIGN:We conducted a retrospective cohort study of proning use among mechanically ventilated adults compared across pre-pandemic (from January 2018 to February 2020), pandemic (from March 2020 to February 2022), and post-pandemic (from March 2022 to December 2024) periods. SETTING:Thirty-seven North American hospitals. PATIENTS:Mechanically ventilated patients with persistent moderate-to-severe hypoxemia (Pa o2 /F io2 ≤ 150 mm Hg, F io2 ≥ 0.6, and positive end-expiratory pressure ≥ 5 cm H 2 O). INTERVENTIONS:Proning within 12 hours of meeting hypoxemia criteria. MEASUREMENTS AND MAIN RESULTS:Among 5944 proning-eligible patients, 2155 (36.2%) received proning: 11.0% pre-pandemic, 51.9% pandemic, and 25.6% post-pandemic. The adjusted odds ratio (OR) for proning during the pandemic vs. pre-pandemic periods was 7.6 (95% CI, 5.5-10.4), and during the pandemic vs. post-pandemic periods was 2.7 (95% CI, 1.8-3.9). Proning varied widely by hospital and was quantified with median ORs (median change in odds of proning for similar patients admitted at a lower vs. higher proning hospital) of 2.9 (95% credible interval [CrI], 1.9-5.3) pre-pandemic, 1.9 (95% CrI, 1.6-2.3) pandemic, and 2.3 (95% CrI, 1.8-3.3) post-pandemic. Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) positive vs. negative pandemic-period patients had increased proning use (OR, 5.1; 95% CI, 4.1-5.6), as did pandemic-period patients without SARS-CoV-2 compared with pre-pandemic patients (OR, 3.8; 95% CI, 2.7-5.2). The difference in proning use between pandemic SARS-CoV-2 negative and post-pandemic patients was smaller and not significant (OR, 1.3; 95% CI, 0.9-1.8). CONCLUSIONS:In a North American cohort of proning-eligible patients, proning increased during the pandemic and then declined. Interventions that improve and sustain implementation of this guideline-recommended intervention are needed.
RATIONALE:Patients receiving invasive mechanical ventilation (IMV) require accurate assessments of ventilator parameters. Documentation of these parameters in standard practice may fail to capture meaningful variation due to intermittent missingness. OBJECTIVES:To assess variation in continuously measured ventilator parameters and agreement with measurements documented as part of routine care in the electronic health record (EHR). METHODS:We performed a retrospective cohort study of patients receiving IMV in a medical intensive care unit from November 2024 through March 2025. We compared the observed tidal volume, minute ventilation, peak inspiratory pressure, and positive end-expiratory pressure, measured continuously from device waveforms with intermittent EHR documentation. We calculated descriptive statistics and measures of agreement between these sources. RESULTS:For 59 encounters, the median age was 65 years (IQR, 59-72 years), 33 (56%) patients were male, and 17 (29%) were Black. Thirty-four (58%) patients died or were discharged to hospice. Among 358 patient-days of data, continuous measurements captured significantly more variation than EHR-documented measurements across all parameters. The largest errors were in observed tidal volume (mean absolute error, 69 mL [95% CI, 62-77 mL]; correlation coefficient 0.540). Agreement in tidal volume was worse among patients receiving mandatory modes of ventilation (correlation coefficient 0.454). CONCLUSIONS:Intermittent measurement of ventilator parameters fails to capture large variability observed in continuous, waveform-derived measurements. Poor agreement in parameters like tidal volume, even in mandatory modes of ventilation, highlights the potential for ventilator waveform data to improve care and advance research for patients with acute respiratory distress syndrome and others receiving IMV.
Objective:. To develop a machine learning model that predicts surgical case length and benchmark its performance against an embedded electronic health record (EHR) model. Background:. Surgical care accounts for one-third of U.S. healthcare expenditure. Current case length prediction models are generally overly simplistic and inaccurate or too specialized to have a broad impact, contributing to operating room (OR) inefficiency and dissatisfaction for patients and providers. Methods:. Retrospective analysis of 55,495 surgical cases performed by 299 surgeons between January 2022 and April 2024 at a metropolitan, quaternary care hospital. The dataset was split temporally for training (46,767 cases) and holdout validation (8728 cases). Three separate machine learning models predicted preprocedure, operative, and postprocedure times using patient and provider characteristics, operation details, and hospital features available at least 1 day before surgery. Approximately 22% of cases lacked historical time averages and relied on procedural time heuristics. Results:. The machine learning model significantly outperformed the embedded EHR model, achieving lower root mean squared error (61.0 vs 91.0 minutes; P < 0.01), lower mean average error (39.6 vs 51.8 minutes; P < 0.01), and higher R2 (0.78 vs 0.50; P < 0.01). The model predicted 213 more cases within ±30 minutes of actual duration. In cases without historical time averages, the model increased cases within ±30 minutes of actual duration (35% vs 29%; P < 0.01). Conclusions:. A machine learning model leveraging comprehensive preoperative data significantly improved surgical case length prediction compared to an embedded EHR model. Future implementation has the potential to improve OR efficiency and patient and provider satisfaction.
BACKGROUND:Post-colectomy adverse events occur in up to one third of patients. We used machine learning (ML) models to predict complications and inform optimal surgeon-hospital assignment. STUDY DESIGN:Adults ≥18 undergoing colon or rectal resection in an academic system from 2018 to 2024 were included. Multiple ML algorithms trained on patient and provider features predicted postoperative complications. Model performance was compared by the scaled Brier score. Patients were simulated to their optimal surgeon-hospital dyad to estimate risk reduction. RESULTS:Among 4689 cases, 1562 (33.3%) experienced ≥1 adverse event. LightGBM outperformed alternate ML models (p < 0.05). LightGBM achieved a scaled Brier score of 0.11 (95% CI 0.08-0.16). Top predictors included pre-operative diagnosis, surgeon identifier, and ostomy. Simulations suggested a 4.6% (CI 4.2-4.9%) net absolute complication risk reduction and 6.9% (CI 6.5-7.3%) reduction with reassignment to an optimal surgeon-hospital dyad. CONCLUSIONS:ML models could enable a proactive referral strategy to promote optimal surgical outcomes.
Background:The coronavirus 2019 disease (COVID-19) pandemic prompted major disruptions in chronic disease self-management and health care delivery, yet its impact on adults with asthma remains poorly characterized. Objective:Our aim was to assess changes in asthma-related health care utilization among adults during the COVID-19 pandemic (2020) compared with before the pandemic (2017-2019) and after the pandemic (2021-2024) within a large, multihospital health system. Methods:We conducted a retrospective electronic health record-based study of 42,242 adults with asthma who were receiving care at Penn Medicine from 2017 to 2024. Weekly counts of 5 encounter types (refill, telemedicine, telephone/audio, outpatient, and emergency encounters) and prescriptions for short-acting β-agonists, inhaled corticosteroids, and oral corticosteroids were compared across years. Generalized linear models evaluated changes in encounter rates during the pandemic and postpandemic periods relative to prepandemic levels, stratified by key transition intervals in 2020. Results:From 2017 to 2019, adults averaged 397 weekly asthma-related visits; in 2020, this number increased to 481. During the lockdown weeks, refill and telemedicine encounters rose by 123% and 36,445%, respectively, whereas outpatient visits declined by 65%. Prescriptions for short-acting β-agonists and inhaled corticosteroids increased by 73% and 43%, respectively, whereas oral corticosteroid prescriptions decreased by 5%. Primary care visits spiked during the lockdown, whereas allergy/immunology and pulmonary encounters remained stable throughout the year. Conclusion:The COVID-19 pandemic was associated with major shifts in adult asthma care, characterized by short-term surges in primary care visits and medication refills and reductions in in-person encounters. These patterns illustrate the capacity of asthma care systems to rapidly adapt, and they highlight the need to tailor future crisis response strategies to adult patients.
Background:Social determinants of health (SDOH) and environmental triggers contribute to risk of asthma exacerbations. Electronic health record data rarely capture complete SDOH and asthma trigger information, thereby limiting comprehensive risk factor assessments absent additional data sources. Objective:We collected detailed SDOH and trigger data from patients with asthma who were identified from electronic health records to better understand modifiable risk factors contributing to risk of asthma emergency department (ED) visits. Methods:We invited patients with asthma identified through the Penn Medicine electronic health record to complete an online questionnaire covering SDOH, asthma triggers, and asthma history. Zero-inflated Poisson regression models were fit to identify factors associated with ED visits for asthma, and hierarchical agglomerative clustering was used to explore underlying patterns in the data. Results:Among 974 survey respondents, 206 reported one or more asthma-related ED visit in the last year. Nearly all SDOH and asthma trigger variables were significantly associated with asthma ED visits in univariable models, but most of these effects were strongly attenuated in multivariable regression models. The effect of race, however, remained strong in both models. Post hoc analyses showed that additionally adjusting for SDOH variables attenuated the relationships between asthma triggers and ED visits. Cluster analysis revealed that the subgroup of participants reporting highest social vulnerability and most asthma triggers also had the most asthma-related ED visits, hospitalizations, and days missed of work or school (all P < 10-4). Conclusion:Self-reported SDOH and asthma trigger data are useful to identify subpopulations at highest risk for adverse asthma outcomes.
OBJECTIVE:The aim of this study was to develop and externally validate a machine-learning model that retrospectively identifies patients with acute respiratory distress syndrome (acute respiratory distress syndrome [ARDS]) using electronic health record (EHR) data.DESIGN:In this retrospective cohort study, ARDS was identified via physician-adjudication in three cohorts of patients with hypoxemic respiratory failure (training, internal validation, and external validation). Machine-learning models were trained to classify ARDS using vital signs, respiratory support, laboratory data, medications, chest radiology reports, and clinical notes. The best-performing models were assessed and internally and externally validated using the area under receiver-operating curve (AUROC), area under precision-recall curve, integrated calibration index (ICI), sensitivity, specificity, positive predictive value (PPV), and ARDS timing.PATIENTS:Patients with hypoxemic respiratory failure undergoing mechanical ventilation within two distinct health systemsINTERVENTIONS:None.MEASUREMENTS AND MAIN RESULTS:There were 1,845 patients in the training cohort, 556 in the internal validation cohort, and 199 in the external validation cohort. ARDS prevalence was 19%, 17%, and 31%, respectively. Regularized logistic regression models analyzing structured data (EHR model) and structured data and radiology reports (EHR-radiology model) had the best performance. During internal and external validation, the EHR-radiology model had AUROC of 0.91 (95% CI, 0.88-0.93) and 0.88 (95% CI, 0.87-0.93), respectively. Externally, the ICI was 0.13 (95% CI, 0.08-0.18). At a specified model threshold, sensitivity and specificity were 80% (95% CI, 75%-98%), PPV was 64% (95% CI, 58%-71%), and the model identified patients with a median of 2.2 hours (interquartile range 0.2-18.6) after meeting Berlin ARDS criteria.CONCLUSIONS:Machine-learning models analyzing EHR data can retrospectively identify patients with ARDS across different institutions.
RATIONALE: Rapidly improving acute respiratory distress syndrome (RIARDS) is a subgroup of ARDS in which hypoxemia significantly improves within 24 hours after initiation of mechanical ventilation. RIARDS has been described in up to 20% of all-cause ARDS cases in an observational cohort, and has been associated with the hypoinflammatory ARDS phenotype. The purpose of this study was to define the clinical characteristics and outcomes of RIARDS among patients with sepsis-associated ARDS, and to validate its association with inflammatory subphenotypes within this cohort. METHODS: We analyzed data from 787 endotracheally intubated patients who met Berlin criteria for ARDS within three days of ICU admission and were enrolled in a prospective observational cohort of critically ill patients with sepsis. RIARDS was defined according to previous studies as improvement of hypoxemia determined by either (i) PaO2:FIO2 > 300 or SpO2:FIO2 > 315 on the day following diagnosis of ARDS (day 2); or (ii) unassisted breathing by day 2 and for the next 48 hours (defined as absence of endotracheal intubation on day 2 through day 4). Participants who did not meet RIARDS criteria were categorized as persistent ARDS. Plasma biomarkers were measured on samples collected on the day of ICU admission, and ARDS subphenotypes were determined using a published parsimonious algorithm using IL-8, sTNFR1, and serum bicarbonate. Two-group comparisons between RIARDS and persistent ARDS were done using the Mann-Whitney U test or Fisher's exact test. RESULTS: Of the 787 enrolled patients, 70 (9%) met criteria for RIARDS. Participants with RIARDS had a lower prevalence of vasopressors use (17% vs 36%, p=0.003), less severe ARDS (p=0.004), and lower 28-day mortality (50% vs 64%; p=0.028). Compared to persistent ARDS, patients with RIARDS less commonly had sepsis from pneumonia (57% vs 70%, p=0.042) and had lower plateau pressures on the day of ARDS diagnosis (21cmH2O vs 23 cmH2O, p=0.005). Plasma levels of inflammatory biomarkers did not differ between RIARDS and persistent disease, and the hyperinflammatory ARDS phenotype was present in over two-thirds of both groups (86% and 71%, p=0.009, respectively). CONCLUSION: In patients with sepsis-associated ARDS, we detected a lower prevalence of RIARDS than previously reported in all-cause ARDS. Consistent with previous studies, RIARDS was associated with less severe clinical disease and lower mortality. The hyperinflammatory phenotype was equally highly enriched in both RIARDS and persistent ARDS, suggesting that the degree of systematic inflammation may not explain clinical differences between these groups among sepsis-associated ARDS.
Rationale: Many research studies rely on the accurate identification of people with asthma using algorithms applied to Electronic Health Record (EHR) data. Although machine learning and rules-based approaches applied to codified variables and extracted notes have reported high classification accuracy, few reports describe in detail the creation of gold standard classifications that precede the creation of predictive models or involve asthma specialists in creating such labels. We sought to measure the consistency of asthma classification according to data in the EHR by asthma specialists, and whether such labels were consistent with commonly applied rules-based definitions of asthma. Methods: We obtained EHR data for 600 adults who had encounters at Penn Medicine with International Classification of Diseases, Tenth Revision (ICD-10) code for asthma (i.e., J45[asterisk]) between January 2017 and August 2023, of these: 200 had no prescription for short-acting beta agonist (SABA) or inhaled corticosteroid (ICS); 200 had a SABA prescription, and 200 had an ICS prescription. Asthma specialists (1 pulmonologist, 3 allergists/immunologists) iteratively created a classification guide with options “Definite/Highly Probable”, “Probable”, “Probably Not/No”, or “Unknown” for having asthma. Two specialists independently labeled records for each adult and disagreements were addressed to reach consensus. We used inter-rater reliability to report classification consistency across specialists. The gold standard classification was compared to four commonly used EHR rules-based schema. Results: After initial chart review, 465 of 600 records had consistent classifications, indicating moderate inter-rater reliability (κ-coefficient=0.66). Following attempted consensus, 593 of 600 records had consistent classifications (κ-coefficient=0.98). The final classification for the 7 remaining disagreements was adjudicated by a third physician. Presence of an asthma ICD-10 code and SABA prescription yielded the greatest consistency with asthma specialist chart review, with 68% of those records having the “Definite/Highly Probable” classification. When aggregating “Definite/Highly Probable” and “Probable” labels together, 89% of records were consistent. Using the asthma ICD-10 code plus SABA and ICS prescriptions yielded similarly high consistency. However, the number of records available when requiring presence of SABA or ICS prescription decreased to a third, suggesting that for some studies, a less strict definition may be preferable. Conclusion: Studies that rely on rules to classify asthma status based on EHR data are inherently limited, partly due to the difficulty of classifying adults with asthma according to information available. However, reasonable consistency between our gold standard classifications and rules-based algorithms provides supporting evidence of using EHR data to understand asthma outcomes with real-world data.
Background:ICU readmissions are associated with increased morbidity, mortality, and healthcare costs. As ICU patient complexity increases and care practices evolve, the contemporary epidemiology of ICU readmissions remains unclear. We aimed to examine ICU readmission rates and timing across multiple health systems, focusing on unplanned readmissions occurring within 24, 48, and 72 hours after ICU discharge. Methods:We performed a retrospective cohort study using federated data from the Common Longitudinal ICU data Format (CLIF) Consortium, comprising nine healthcare systems between January 2020 and December 2021 and the MIMIC-IV database. The cohort included adult patients (≥18 years) discharged alive from the ICU. Readmissions following planned surgeries or interventional procedures were excluded. Data were analyzed locally at each site without centralizing patient-level data, and analyses focused on patient demographics, discharge disposition, readmission timing, and clinical interventions during ICU stays and readmissions. Statistical comparisons were performed using two-proportion z-tests and chi-squared tests. Results:Among 185,241 hospital admissions across 19 hospitals, 8.6% of ICU discharges were readmitted during the same hospitalization. Unplanned readmissions occurred within 24 hours in 1.9% of cases, 3.4% within 48 hours, and 4.5% within 72 hours. Readmitted patients experienced higher in-hospital mortality (20.6% vs. 2.1%, p<0.001). Compared to the initial ICU stay, ICU readmissions were associated with significantly increased respiratory (42.3% vs. 35.3%, p<0.001) and vasopressor support (26.1% vs. 23.1%, p<0.001). Conclusions:ICU readmissions remain common and are linked to worse outcomes. Readmissions require more respiratory and vasopressor support. Future work should focus on characterizing these subphenotypes and improving ICU discharge processes to reduce preventable readmissions.
Rationale Global Initiative for Chronic Obstructive Lung Disease (GOLD) guidelines recommend the diagnosis of chronic obstructive pulmonary disease (COPD) only in patients with a post-bronchodilator forced expiratory volume in 1 second to forced vital capacity ratio (FEV1/FVC) less than 0.7. However the impact of this recommendation on clinical practice is unknown. We hypothesized that if physicians applied GOLD guidelines to spirometry to diagnose COPD, this would yield a substantial discontinuity in the probability of a COPD diagnosis at the 0.7 cutoff; COPD would not be diagnosed in any patient with an FEV1/FVC just above 0.7, and would be diagnosed in most patients with an FEV1/FVC just below 0.7. As GOLD guidelines do not recommend that spirometry directly inform treatment decisions, though treatment decisions likely follow the establishment of a COPD diagnosis, we hypothesized that a post-bronchodilator FEV1/FVC < 0.7 would affect COPD treatment but to a lesser extent than it would affect diagnosis. Methods This retrospective cohort study used the Optum Labs Data Warehouse, a database of electronic health record data collected from across the United States. We included patients who were 18 years of age and older and had a clinical encounter between 2007 and 2022 in which a post-bronchodilator FEV1/FVC value was documented. A clinical encounter was associated with a COPD diagnosis if an international classification of disease code for COPD was assigned, and was associated with COPD treatment if a prescription for a medication commonly used to treat COPD was filled within 90 days. We used a regression discontinuity design to measure the effect of a post-bronchodilator FEV1/FVC < 0.7 on COPD diagnosis and treatment. Results The cohort included 27,817 clinical encounters involving 18,991 different patients and more than 2,083 different physicians. The presence of a documented post-bronchodilator FEV1/FVC < 0.7 increased the probability of a COPD diagnosis by 0.06 (95% confidence interval [CI] 0.01 to 0.11) from 0.38 just above the 0.7 cutoff to 0.44 just below this cutoff (Figure). The presence of a documented post bronchodilator FEV1/FVC < 0.7 had no effect on the probability of COPD treatment (−0.02, 95% CI −0.07 to 0.03). Conclusions The presence of a documented post-bronchodilator FEV1/FVC < 0.7 had only a small effect on the diagnosis of COPD and had no effect on COPD treatment. Physicians do not use spirometry to diagnose or treat COPD in the manner recommended by GOLD guidelines.
Background:Though a normal forced vital capacity (FVC) is typically thought to imply the absence of restriction, recent data suggest that restriction may in fact be common among patients with normal spirometry. However, the clinical significance of restriction with normal spirometry is unknown. Research Question:What clinical characteristics and outcomes are associated with restriction with normal spirometry? Study Design and Methods:We interpreted pulmonary function tests (PFTs) with both static and dynamic lung volume measurements performed between 2012 and 2025 at four pulmonary diagnostic labs. We used multivariable logistic regression to identify clinical characteristics associated with restriction among patients with normal spirometry and used a Cox proportional hazards model to assess the association of restriction with survival, adjusting for age, sex, forced expiratory volume in 1 second (FEV1) z-score, FVC z-score, and FEV1/FVC z-score. Results:We interpreted 83,886 PFTs from 47,597 patients (mean age 58.8 years, 59.8% female, 63.6% White). The prevalence of restriction among patients with normal spirometry was 25.7% Restriction with normal spirometry was more likely in older patients (adjusted odds ratio [aOR] 1.01 per year, 95% CI 1.01-1.01), in non-White patients (aOR 1.33, 95% CI 1.26-1.41), and in patients with a diagnosis (aOR 3.65, 95% CI 3.43-3.88) or radiographic evidence (aOR 3.02, 95% CI 2.79-3.28) of interstitial lung disease (ILD). Restriction with normal spirometry was less likely among female patients (aOR 0.64, 95% CI 0.60-0.67), and patients with a diagnosis (0.74, 95% CI 0.67-0.82) or radiographic evidence (aOR 0.81, 95% CI 0.73-0.89) of chronic obstructive pulmonary disease. Restriction with normal spirometry was associated with increased all-cause mortality (adjusted hazard ratio 1.45, 95% CI 1.34-1.57) as compared to normal spirometry without restriction. Interpretation:Restriction with normal spirometry is associated with ILD and with decreased survival. Clinically significant ventilatory impairments are common in patients with normal spirometry.
Artificial intelligence (AI) methods were first developed nearly seven decades ago. Only in recent years have they demonstrated their potential to improve clinical care at the bedside. AI systems are now capable of interpreting, predicting, and even generating important medical information. AI medical devices share many similarities with traditional medical devices but also diverge from them in important ways. Despite widespread optimism and enthusiasm surrounding the use of such devices to improve care processes, patient outcomes, and the healthcare experience for patients, caregivers, and clinicians alike, little evidence exists so far for their effectiveness in practice. Even less is known about the safety or equity of AI medical devices. As with any new technology, this exciting time is accompanied by appropriate questions regarding if, how much, when, and who such AI systems really help. Different stakeholders, ranging from patients to clinicians to industry device developers, may have divergent preferences or assessments of risk and benefits, warranting an informed public discussion to guide emerging regulatory efforts. This review summarizes the rapidly evolving recent efforts and evidence related to the regulation and evaluation of AI medical devices and highlights opportunities for future work to ensure their effectiveness, safety, and equity.