Regression analysis of correlated data, where multiple correlated responses are recorded on the same unit, is ubiquitous in many scientific areas. With the advent of new technologies, in particular high-throughput omics profiling assays, such correlated data increasingly consist of a large number of variables compared with the available sample size. Motivated by recent longitudinal proteomics studies of COVID-19, we propose a novel inference procedure for linear functionals of high-dimensional regression coefficients in generalized estimating equations, which are widely used to analyze correlated data. Our estimator for this more general inferential target, obtained via constructing projected estimating equations, is shown to be asymptotically normally distributed under mild regularity conditions. We also introduce a data-driven cross-validation procedure to select the tuning parameter for estimating the projection direction, which is not addressed in the existing procedures. We illustrate the utility of the proposed procedure in providing confidence intervals for associations of individual proteins and severe COVID risk scores obtained based on high-dimensional proteomics data, and demonstrate its robust finite-sample performance, especially in estimation bias and confidence interval coverage, via extensive simulations.
Background The significant clinical and molecular heterogeneity of pulmonary arterial hypertension (PAH) poses challenges in identifying effective therapies. Advanced multidimensional profiling offers an opportunity to capture molecular responses and assess biomarker stability, yet its application in randomised trials remains limited. Methods We evaluated the multi-omic profiles of participants with PAH in a randomised, placebo-controlled trial of famotidine. Plasma metabolomic and proteomic profiling was performed at enrolment and 24 weeks. Baseline profiles were compared between treatment arms to assess randomisation balance. Intraclass correlation coefficients quantified within-subject stability over time. Linear regression models adjusting for age, sex, body mass index and PAH aetiology evaluated famotidine's molecular effects. False discovery rate was controlled for multiple comparisons. Findings For the 79 participants, baseline multi-omic profiles were similar between groups. At 24 weeks, 34 and 37 participants remained in the famotidine and placebo groups respectively. The placebo group showed high molecular stability, while greater variability was observed in the famotidine group. Famotidine treatment was associated with significant changes across 191 proteomic pathways (q-value <0.05), but no metabolomic changes remained significant after multiple-testing correction. Interpretation Integrating multi-omics into a prospective clinical trial is feasible and yields stable longitudinal profiles in the absence of intervention. While famotidine did not yield clinical benefit, associated proteomic changes illustrate how molecular profiling can reveal treatment-related biology and inform future trial design. These findings highlight the broader utility of multi-omics for evaluating drug responses and identifying molecular endotypes in PAH and beyond. Funding US National Institutes of Health.
RATIONALE Pulmonary arterial hypertension (PAH) is a progressive disease characterized by worsening vascular remodeling and eventual right heart failure. Identifying biological signatures that reflect or predict clinical decline could offer new opportunities to improve patient outcomes and guide therapeutic interventions. METHODS Plasma metabolomic and proteomic profiling was performed on samples from participants in a negative clinical trial of PAH, assessing 1118 metabolites and 6386 unique proteins at baseline and 24-week follow-up. Changes in right ventricular (RV) dilation, function, 6-minute walk distance (6MWD), brain natriuretic peptide (BNP), New York Heart Association (NYHA) Functional Class, and EmPHasis-10 scores were evaluated as measures of disease progression. Associations between clinical decline and changes in omic signatures over 6 months were analyzed. To identify common molecular pathways associated with clinical outcomes, weighted correlation network analysis (for metabolomics) and gene set enrichment analysis (for proteomics) were performed. Models were adjusted for age, sex, body mass index, and PAH etiology, with control for multiple comparisons. RESULTS A total of 71 participants with World Health Organization Group 1 PAH completed the study. Metabolomic pathway analyses revealed that acylcarnitine related metabolites were upregulated in participants who had worsening RV dilation and function, higher levels of BNP, worse NYHA functional class, higher EMPHASIS10 scores and shorter 6MWD (see Figure). Proteomic pathway analyses showed that extracellular matrix (ECM) related pathways were upregulated in participants with shorter 6MWD, and downregulated in those who had improving EmPHasis-10 scores, NYHA Functional Class and lower levels of BNP during follow-up. CONCLUSIONS Distinct plasma metabolomic and proteomic profiles are associated with progressive decline in PAH. Acylcarnitine and ECM pathways represent promising candidates for identifying patients at high risk for poor outcomes and for understanding the molecular mechanisms underlying disease progression and RV adaptation. Figure: Weighted Correlation Network Analysis of changes in plasma metabolomic profiles and metrics of clinical decline over 6 months in participants with pulmonary arterial hypertension. The red module[asterisk], enriched with acylcarnitine metabolites, shows significant upregulation across all metrics of clinical worseningΤ. Τ All outcomes were coded to reflect clinical decline. This included increases in brain natriuretic peptide levels, EmPHasis-10 scores, New York Heart Association Functional Class, and right ventricular basal diameter (indicating dilation), as well as decreases in 6-minute walk distance and tricuspid annular plane systolic excursion (indicating impaired function).
Background:Community-acquired pneumonia (CAP) is a major public health threat globally but is understudied in regions with the highest burden. The host immune response during infection may differ based on the site of infection. We hypothesised that analysis of the plasma metabolome in patients hospitalised with suspected infection could identify host response pathways specific to CAP. Methods:We analysed the plasma metabolomes of adults admitted to a tertiary care hospital in northeastern Thailand with suspected community-acquired infection. Multivariable linear regression was performed for differential metabolite analyses and the global test was used for pathway analysis comparing patients with CAP versus non-CAP infections and uninfected controls. The least absolute shrinkage and selection operator (LASSO) was used to identify a parsimonious metabolite prognostic signature that was tested on an internal validation set to predict mortality. Results:841 metabolites from 107 CAP patients and 152 non-CAP infected patients were analysed. 52 metabolites were differentially abundant between the CAP and non-CAP groups. CAP was characterised by increased metabolites involved in polyamine metabolism and decreased metabolites involved in lipid pathways. 13 pathways were differentially enriched between the CAP and non-CAP groups, consistent with individual metabolite analyses. 40 metabolites and four pathways were associated with CAP-specific mortality. A four-metabolite signature predicted 28-day mortality in CAP (area under the curve 0.79, 95% CI 0.62-0.97). Conclusion:In a rural tropical setting, CAP induced a distinct metabolomic state compared to non-CAP presentations of infection that may reflect the activation of select host immune responses.
RATIONALE Pulmonary arterial hypertension (PAH) is a devastating disease with high mortality and morbidity. The significant clinical and molecular heterogeneity of PAH presents challenges in identifying effective therapeutic interventions. Targeting molecular changes through advanced omics profiling may offer new insights into treatment effects. METHODS In an NIH-sponsored, single-center, randomized, placebo-controlled trial of famotidine (an H2 receptor antagonist), 79 adults with PAH received treatment and clinical follow-up over 24 weeks. The primary end-point, 6-minute walk distance at 24 weeks, was not statistically different between the two groups. We collected plasma metabolomic and proteomic profiling on 1118 metabolites and 6386 unique proteins at time of enrollment and 24-week follow-up. Baseline metabolomic and proteomic profiles were compared between treatment arms, and paired analyses were performed to assess the impact of famotidine. Significant metabolites, proteins and pathways were identified, controlling for multiple comparisons. All analyses were adjusted for age, sex, body mass index and PAH etiology. RESULTS Of the 79 participants with PAH, 40 were assigned to in the famotidine arm and 39 to the placebo arm at baseline. No significant differences in metabolomic or proteomic profiles were detected between the two groups at baseline. By the 24-week follow-up, 34 participants remained in the famotidine group and 37 in the placebo group. Although metabolomic changes were not observed, famotidine treatment was associated with significant changes in 704 proteins and 20 proteomic pathways (adjusted p-value < 0.05) (see Figure). CONCLUSIONS This study demonstrates that robust randomization effectively balances clinical and molecular profiles between treatment groups. Our findings show that plasma metabolites and proteins are valuable tools for assessing molecular changes in response to therapy, even in the absence of significant clinical differences. Famotidine treatment induced notable proteomic alterations over the 24-week period, highlighting its potential molecular effects, while metabolomic changes were minimal. These results emphasize the need for future biomarker research to identify subgroups of patients who may have greater molecular responsiveness to famotidine, supporting the development of more personalized therapeutic approaches in PAH management. Figure: Summary plots of the 704 proteins significantly associated with treatment effect of famotidine for participants with pulmonary arterial hypertension. Each line represents the average fold change in protein expression over 6 months for each treatment group, highlighting the differences in response between groups.
Rationale: The global burden of sepsis is greatest in low-resource settings. Melioidosis, infection with the gram-negative bacterium Burkholderia pseudomallei, is a frequent cause of fatal sepsis in endemic tropical regions such as Southeast Asia. Objectives: To investigate whether plasma metabolomics would identify biological pathways specific to melioidosis and yield clinically meaningful biomarkers. Methods: Using a comprehensive approach, differential enrichment of plasma metabolites and pathways was systematically evaluated in individuals selected from a prospective cohort of patients hospitalized in rural Thailand with infection. Statistical and bioinformatics methods were used to distinguish metabolomic features and processes specific to patients with melioidosis and between fatal and nonfatal cases. Measurements and Main Results: Metabolomic profiling and pathway enrichment analysis of plasma samples from patients with melioidosis (n = 175) and nonmelioidosis infections (n = 75) revealed a distinct immuno-metabolic state among patients with melioidosis, as suggested by excessive tryptophan catabolism in the kynurenine pathway and significantly increased levels of sphingomyelins and ceramide species. We derived a 12-metabolite classifier to distinguish melioidosis from other infections, yielding an area under the receiver operating characteristic curve of 0.87 in a second validation set of patients. Melioidosis nonsurvivors (n = 94) had a significantly disturbed metabolome compared with survivors (n = 81), with increased leucine, isoleucine, and valine metabolism, and elevated circulating free fatty acids and acylcarnitines. A limited eight-metabolite panel showed promise as an early prognosticator of mortality in melioidosis. Conclusions: Melioidosis induces a distinct metabolomic state that can be examined to distinguish underlying pathophysiological mechanisms associated with death. A 12-metabolite signature accurately differentiates melioidosis from other infections and may have diagnostic applications.
BACKGROUND:Pulmonary arterial hypertension (PAH) is a disease of progressive right ventricular (RV) failure with high morbidity and mortality. Our goal is to investigate proteomic features and pathways associated with RV-focused outcomes including mortality, RV dilation, and NT-proBNP (N-terminal pro-B-type natriuretic peptide) in PAH.METHODS:Participants in a single-institution cohort with 3 years of follow-up underwent proteomic profiling of their plasma using 7288 aptamers (targeting 6467 unique human proteins). Partial least squares discriminant analysis was performed to assess global protein variation associated with mortality, RV dilation, and NT-proBNP levels. Differentially abundant proteins and enriched pathways associated with outcomes were identified following baseline adjustments. RV vulnerability models estimated associations for individuals with similar afterload following adjustment for pulmonary vascular resistance.RESULTS:A total of 117 participants with PAH were included. Partial least squares discriminant analysis of the proteome showed clear separation between survivors and nonsurvivors, participants with dilated versus nondilated RVs, and across NT-proBNP levels. Proteins and pathways involving the ECM (extracellular matrix) were upregulated in participants who died during follow-up, those with severe RV dilation, and those with higher levels of NT-proBNP. Pulmonary vascular resistance adjustment reinforced the importance of ECM proteins in the association with RV vulnerability, independent of afterload. These findings were confirmed in independent PAH cohorts with available plasma proteomics and RV tissue gene and protein expression.CONCLUSIONS:Distinct plasma proteomic profiles are associated with mortality, RV dilation, and NT-proBNP in PAH. Proteins and pathways governing tissue remodeling are strongly associated with poor outcomes, may mediate RV vulnerability to right heart failure, and represent promising candidates as biomarkers and potential therapeutic targets.
Melioidosis, a neglected tropical infection caused by Burkholderia pseudomallei, , commonly presents as pneumonia or sepsis with mortality rates up to 50% despite appropriate treatment. A better understanding of the early host immune response to melioidosis may lead to new therapeutic interventions and prognostication strategies to reduce disease burden. Whole blood transcriptomic signatures in 164 patients with melioidosis and in 70 patients with other infections hospitalized in northeastern Thailand enrolled within 24 hours following hospital admission were studied. Key findings were validated in an independent melioidosis cohort. Melioidosis was characterized by upregulation of interferon (IFN) signaling responses compared with other infections. Mortality in melioidosis was associated with excessive inflammation, enrichment of type 2 immune responses, and a dramatic decrease in T cell-mediated immunity compared with survivors. We identified and independently confirmed a 5-gene predictive set classifying fatal melioidosis (validation cohort area under the receiver operating characteristic curve 0.83; 95% CI, 0.67-0.99). This study highlights the intricate balance between innate and adaptive immunity during fatal melioidosis and can inform future precision medicine strategies for targeted therapies and prognostication in this severe infection.
Background: Pulmonary arterial hypertension (PAH) is a complex disease characterized by progressive right ventricular (RV) failure leading to significant morbidity and mortality. Investigating metabolic features and pathways associated with RV dilation, mortality, and measures of disease severity can provide insight into molecular mechanisms, identify subphenotypes, and suggest potential therapeutic targets. Methods: We collected data from a prospective cohort of PAH participants and performed untargeted metabolomic profiling on 1045 metabolites from circulating blood. Analyses were intended to identify metabolomic differences across a range of common metrics in PAH (eg, dilated versus nondilated RV). Partial least squares discriminant analysis was first applied to assess the distinguishability of relevant outcomes. Significantly altered metabolites were then identified using linear regression, and Cox regression models (as appropriate for the specific outcome) with adjustments for age, sex, body mass index, and PAH cause. Models exploring RV maladaptation were further adjusted for pulmonary vascular resistance. Pathway enrichment analysis was performed to identify significantly dysregulated processes. Results: A total of 117 participants with PAH were included. Partial least squares discriminant analysis showed cluster differentiation between participants with dilated versus nondilated RVs, survivors versus nonsurvivors, and across a range of NT-proBNP (N-terminal pro-B-type natriuretic peptide) levels, REVEAL 2.0 composite scores, and 6-minute-walk distances. Polyamine and histidine pathways were associated with differences in RV dilation, mortality, NT-proBNP, REVEAL score, and 6-minute walk distance. Acylcarnitine pathways were associated with NT-proBNP, REVEAL score, and 6-minute walk distance. Sphingomyelin pathways were associated with RV dilation and NT-proBNP after adjustment for pulmonary vascular resistance. Conclusions: Distinct plasma metabolomic profiles are associated with RV dilation, mortality, and measures of disease severity in PAH. Polyamine, histidine, and sphingomyelin metabolic pathways represent promising candidates for identifying patients at high risk for poor outcomes and investigation into their roles as markers or mediators of disease progression and RV adaptation.
The Scientific Registry of Transplant Recipients (SRTR) system has become a rich resource for understanding the complex mechanisms of graft failure after kidney transplant, a crucial step for allocating organs effectively and implementing appropriate care. As transplant centers that treated patients might strongly confound graft failures, Cox models stratified by centers can eliminate their confounding effects. Also, since recipient age is a proven non-modifiable risk factor, a common practice is to fit models separately by recipient age groups. The moderate sample sizes, relative to the number of covariates, in some age groups may lead to biased maximum stratified partial likelihood estimates and unreliable confidence intervals even when samples still outnumber covariates. To draw reliable inference on a comprehensive list of risk factors measured from both donors and recipients in SRTR, we propose a de-biased lasso approach via quadratic programming for fitting stratified Cox models. We establish asymptotic properties and verify via simulations that our method produces consistent estimates and confidence intervals with nominal coverage probabilities. Accounting for nearly 100 confounders in SRTR, the de-biased method detects that the graft failure hazard nonlinearly increases with donor's age among all recipient age groups, and that organs from older donors more adversely impact the younger recipients. Our method also delineates the associations between graft failure and many risk factors such as recipients' primary diagnoses (e.g. polycystic disease, glomerular disease, and diabetes) and donor-recipient mismatches for human leukocyte antigen loci across recipient age groups. These results may inform the refinement of donor-recipient matching criteria for stakeholders.
For statistical inference on regression models with a diverging number of covariates, the existing literature typically makes sparsity assumptions on the inverse of the Fisher information matrix. Such assumptions, however, are often violated under Cox proportion hazards models, leading to biased estimates with under-coverage confidence intervals. We propose a modified debiased lasso method, which solves a series of quadratic programming problems to approximate the inverse information matrix without posing sparse matrix assumptions. We establish asymptotic results for the estimated regression coefficients when the dimension of covariates diverges with the sample size. As demonstrated by extensive simulations, our proposed method provides consistent estimates and confidence intervals with nominal coverage probabilities. The utility of the method is further demonstrated by assessing the effects of genetic markers on patients' overall survival with the Boston Lung Cancer Survival Cohort, a large-scale epidemiology study investigating mechanisms underlying the lung cancer.
Regression analysis of correlated data, where multiple correlated responses are recorded on the same unit, is ubiquitous in many scientific areas. With the advent of new technologies, in particular high-throughput omics profiling assays, such correlated data increasingly consist of large number of variables compared with the available sample size. Motivated by recent longitudinal proteomics studies of COVID-19, we propose a novel inference procedure for linear functionals of high-dimensional regression coefficients in generalized estimating equations, which are widely used to analyze correlated data. Our estimator for this more general inferential target, obtained via constructing projected estimating equations, is shown to be asymptotically normally distributed under mild regularity conditions. We also introduce a data-driven cross-validation procedure to select the tuning parameter for estimating the projection direction, which is not addressed in the existing procedures. We illustrate the utility of the proposed procedure in providing confidence intervals for associations of individual proteins and severe COVID risk scores obtained based on high-dimensional proteomics data, and demonstrate its robust finite-sample performance, especially in estimation bias and confidence interval coverage, via extensive simulations.
Monitoring outcomes of health care providers, such as patient deaths, hospitalizations and hospital readmissions, helps in assessing the quality of health care. We consider a large database on patients being treated at dialysis facilities in the United States, and the problem of identifying facilities with outcomes that are better than or worse than expected. Analyses of such data have been commonly based on random or fixed facility effects, which have shortcomings that can lead to unfair assessments. A primary issue is that they do not appropriately account for variation between providers that is outside the providers' control due, for example, to unobserved patient characteristics that vary between providers. In this article, we propose a smoothed empirical null approach that accounts for the total variation and adapts to different provider sizes. The linear model provides an illustration that extends easily to other nonlinear models for survival or binary outcomes, for example. The empirical null method is generalized to allow for some variation being due to quality of care. These methods are examined with numerical simulations and applied to the monitoring of survival in the dialysis facility data.
Modeling and drawing inference on the joint associations between single-nucleotide polymorphisms and a disease has sparked interest in genome-wide associations studies. In the motivating Boston Lung Cancer Survival Cohort (BLCSC) data, the presence of a large number of single nucleotide polymorphisms of interest, though smaller than the sample size, challenges inference on their joint associations with the disease outcome. In similar settings, we find that neither the debiased lasso approach (van de Geer et al., 2014), which assumes sparsity on the inverse information matrix, nor the standard maximum likelihood method can yield confidence intervals with satisfactory coverage probabilities for generalized linear models. Under this "large n, diverging p" scenario, we propose an alternative debiased lasso approach by directly inverting the Hessian matrix without imposing the matrix sparsity assumption, which further reduces bias compared to the original debiased lasso and ensures valid confidence intervals with nominal coverage probabilities. We establish the asymptotic distributions of any linear combinations of the parameter estimates, which lays the theoretical ground for drawing inference. Simulations show that the proposed refined debiased estimating method performs well in removing bias and yields honest confidence interval coverage. We use the proposed method to analyze the aforementioned BLCSC data, a large-scale hospital-based epidemiology cohort study investigating the joint effects of genetic variants on lung cancer risks.
De-biased lasso has emerged as a popular tool to draw statistical inference for high-dimensional regression models. However, simulations indicate that for generalized linear models (GLMs), de-biased lasso inadequately removes biases and yields unreliable confidence intervals. This motivates us to scrutinize the application of de-biased lasso in high-dimensional GLMs. When $p >n$, we detect that a key sparsity condition on the inverse information matrix generally does not hold in a GLM setting, which likely explains the subpar performance of de-biased lasso. Even in a less challenging "large $n$, diverging $p$" scenario, we find that de-biased lasso and the maximum likelihood method often yield confidence intervals with unsatisfactory coverage probabilities. In this scenario, we examine an alternative approach for further bias correction by directly inverting the Hessian matrix without imposing the matrix sparsity assumption. We establish the asymptotic distributions of any linear combinations of the resulting estimates, which lay the theoretical groundwork for drawing inference. Simulations show that this refined de-biased estimator performs well in removing biases and yields an honest confidence interval coverage. We illustrate the method by analyzing a prospective hospital-based Boston Lung Cancer Study, a large scale epidemiology cohort investigating the joint effects of genetic variants on lung cancer risk.
To assess the quality of health care, patient outcomes associated with medical providers (eg, dialysis facilities) are routinely monitored in order to identify poor (or excellent) provider performance. Given the high stakes of such evaluations for payment as well as public reporting of quality, it is important to assess the reliability of quality measures. A commonly used metric is the inter-unit reliability (IUR), which is the proportion of variation in the measure that comes from inter-provider differences. Despite its wide use, however, the size of the IUR has little to do with the usefulness of the measure for profiling extreme outcomes. A large IUR can signal the need for further risk adjustment to account for differences between patients treated by different providers, while even measures with an IUR close to zero can be useful for identifying extreme providers. To address these limitations, we propose an alternative measure of reliability, which assesses more directly the value of a quality measure in identifying (or profiling) providers with extreme outcomes. The resulting metric reflects the extent to which the profiling status is consistent over repeated measurements. We use national dialysis data to examine this approach on various measures of dialysis facilities.
OBJECTIVES:Classical methods for combining summary data from genome-wide association studies only use marginal genetic effects, and power can be compromised in the presence of heterogeneity. We aim to enhance the discovery of novel associated loci in the presence of heterogeneity of genetic effects in subgroups defined by an environmental factor. METHODS:We present a pvalue-assisted subset testing for associations (pASTA) framework that generalizes the previously proposed association analysis based on subsets (ASSET) method by incorporating gene-environment (G-E) interactions into the testing procedure. We conduct simulation studies and provide two data examples. RESULTS:Simulation studies show that our proposal is more powerful than methods based on marginal associations in the presence of G-E interactions and maintains comparable power even in their absence. Both data examples demonstrate that our method can increase power to detect overall genetic associations and identify novel studies/phenotypes that contribute to the association. CONCLUSIONS:Our proposed method can be a useful screening tool to identify candidate single nucleotide polymorphisms that are potentially associated with the trait(s) of interest for further validation. It also allows researchers to determine the most probable subset of traits that exhibit genetic associations in addition to the enhancement of power.
In monitoring health care providers, various outcomes are used to assess the performance and quality of care given. We consider a measure that is normally distributed across the majority of providers with both a within and between component contributing to the overall variability of the measure. In such cases, the inter-unit reliability (IUR) is commonly used to assess the usefulness of a measure for identifying extreme providers. In this article, we define and discuss the IUR and note its role under various assumptions about the source of the between-provider variance. This variability may be due primarily to differences in the patients that are not accounted for in the measured covariates, or differences in the quality of care provided, or a combination of these two. The IUR is a simple population characteristic specifying the proportion of the variation in the measure that is related to between-provider differences without regard to the source of that variation. On the other hand, the IUR does not characterize the suitability of a measure for profiling or identifying outliers except in very special circumstances where all the variation is due to quality of care and there are no outliers. In assessing the reliability of a measure for profiling, the key question is whether outlying providers with poor (or excellent) quality of care give rise to extreme values of the measure.