
Graph-based denoising is a critical preprocessing step for analyzing noisy data, particularly in genomic applications where gene regulatory networks exhibit inherent directional dependencies. This paper introduces a directed acyclic graph trend filtering (GTF) framework that leverages novel higher-order Bayesian networks and graphical shrinkage processes to enhance local adaptivity in signal smoothing along the directed edges of a graph. Unlike traditional GTF, which is based on undirected graphs, the proposed method explicitly respects the directional structure of graphs, improving interpretability and accuracy in capturing dependencies. We employ a Hamiltonian Monte Carlo algorithm for efficient posterior inference. Through simulations and genomic applications, the proposed method outperforms a state-of-the-art GTF algorithm in terms of mean squared error reduction and signal-to-noise ratio improvement, demonstrating its utility in recovering true signals while accounting for meaningful structural information.
In this paper, we propose a Bayesian multivariate network meta-regression model to compare multiple treatments used to treat cardiovascular and diabetes diseases, where the multivariate aggregate outcomes include low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, and triglycerides. We assume a log-linear regression model for the standard deviation of the treatment random effects to overcome the difficulty that some treatments may present only in a single study. As the within-study sample covariance matrix $\boldsymbol {S}$ is partially observed or completely missing and the within-study sample correlations are not observed at all, we postulate a hierarchical structure on the unknown within-study covariance matrices. We further develop a Markov chain Monte Carlo sampling algorithm to sample from the posterior distribution and a Monte Carlo procedure to rank the treatment effects for the multivariate outcomes. DIC is used for model comparison. Two variations of DIC are further developed to quantify (i) the overall improvement in the fit and (ii) the gain in the fit of each outcome due to the multivariate model versus the univariate model alone. A detailed analysis of the aggregate data from real randomized controlled trials is carried out to further demonstrate the proposed methodology.
Clustering is a fundamental tool for uncovering heterogeneity in data, but two challenges remain: determining whether observed clusters reflect genuine structure rather than sampling variability, and identifying the variables that drive any significant clustering pattern. Statistical significance of clustering (SigClust) addresses the first problem by assessing clustering significance using the cluster index under a Gaussian null model, with the null distribution estimated by Monte Carlo simulation in high dimensions. We propose SigClust-DE, a method that improves null covariance estimation in SigClust and extends the framework to variable-level inference for identifying features associated with cluster separation. In this way, SigClust-DE provides a joint framework for clustering significance testing and differential expression analysis, a central task in RNA-seq studies. Through extensive simulations and an application to RNA-seq data, we show that SigClust-DE controls Type I error in clustering significance testing, controls the false discovery rate in variable-level inference, and achieves strong power for detecting differentially expressed features.
Random-effects models are central to meta-analysis, yet the between-study variance is often underestimated when the number of studies is small. In such settings, confidence intervals become unduly narrow and fail to attain the nominal coverage probability. Although several small-sample corrections, including the Bartlett correction, have been developed under the normal-normal model, corresponding methodology for generalized linear mixed models (GLMMs) remains limited. This study proposes a unified framework for random-effects meta-analysis within the GLMM framework that relies exclusively on aggregate data and accommodates outcomes following any distribution in the exponential family, including the binomial, Poisson, and gamma distributions. To improve interval estimation with few studies, we develop a profile likelihood method with a simplified Bartlett correction (PLSBC), which refines the chi-squared approximation of the profile likelihood ratio statistic without requiring higher-order derivatives. We show theoretically that the proposed estimators preserve the consistency and asymptotic normality of the maximum likelihood estimators. Simulation studies demonstrate that the PLSBC yields nearly unbiased estimates and maintains nominal coverage across a variety of outcome types. Applications to three published meta-analyses with binomial, Poisson, and gamma outcomes indicate that the proposed approach provides robust and interpretable inference with few studies. The PLSBC therefore offers a practical and broadly applicable framework for random-effects meta-analysis when the number of studies is limited.
Neyman-Pearson (NP) classifiers, which aim to maximize the clinical benefit while adhering to risk constraints, are crucial in many practical fields, including early cancer detection. However, applying these classifiers can be challenging due to discrepancies between the data distributions of the source and target populations. The potential impact can be disproportionately severe for under-represented groups. We propose a semi-parametric model-based approach for adapting NP classifier decision rules to different populations while equitably controlling classification errors specific to clinical applications. Our method involves a shift-adjustment strategy that leverages a small unlabeled sample from the target population, along with minimal auxiliary information and the labeled source data. This approach enhances the applicability of the learned decision rules and ensures they are consistently tailored for the target population. We demonstrate the performance through theoretical studies and simulations and illustrate the approach with an example of a prostate cancer study.
For the analysis of interval-censored data, we propose a deep generalized accelerated hazards model. This model is designed to facilitate a detailed exploration of the relationship between various risk factors and the hazard associated with failure time. We develop a sieve maximum likelihood estimation procedure that combines deep neural networks and monotonic splines. By employing deep neural networks, we can effectively capture nonparametric effects, enabling a flexible and adaptive modeling approach for complex relationships. Under certain regularity conditions, we derive a nonasymptotic error bound for the resulting estimator and show that the finite-dimensional estimator is asymptotically normal and achieves the semiparametric efficiency. We conduct simulation studies to evaluate the finite-sample performance of the proposed approach. Furthermore, the proposed method is applied to the Atherosclerosis Risk in Communities study for practical illustration.
Serial dilution assays are essential tools in biomedical research. Current analytical methods work well when measurements are in the steep part of the calibration curve but are not so efficient when measurements are near the upper and lower bounds. Moreover, sample contamination is common in serial dilution assays and can result in problematic concentration estimation. To address these challenges, we propose a Bayesian contamination model that simultaneously flags contaminated samples and estimates the concentrations of unknown samples, while quantifying estimation uncertainty. In extensive simulation studies under various contamination scenarios, our method consistently outperforms commonly used approaches. We further demonstrate a step-by-step application of a Bayesian workflow for modeling contamination using raw immunoassay data from the New York City Neighborhood Asthma and Allergy Study. This work advances previous research by introducing a practical method for detecting contaminated samples and estimating unknown concentrations in immunoassays and by presenting a general framework for incorporating unknown contamination processes into the Bayesian modeling workflow.
Composite endpoints, which combine multiple events of interest, are commonly used in medical research. Compared to evaluating a single outcome, a composite endpoint analysis captures the full clinical impact of treatment and leverages a greater number of events, potentially reducing the required sample size for the study. While composite endpoints have gained much attention in clinical trials, studying them in group sequential designs remains challenging due to the correlated nature of event data collected from the same individual, as the sequence of test statistics may not possess an independent increments structure. In this paper, we propose both one-sample and two-sample group sequential designs grounded in the mean frequency function, defined as the cumulative count of all recurrent and terminal events over time. Recognizing the lack of independent increments, our proposed method leverages the asymptotic covariance structure of the test statistics to construct group sequential boundaries that control the Type I error rate. Extensive simulation studies show the proposed design controls Type I error well and achieves the desired power. We illustrate the utility of our method through a reanalysis of BMT CTN 1703, a phase III randomized controlled trial that evaluated an experimental therapy for the prevention of adverse outcomes after allogeneic stem cell transplant.
We propose one-at-a-time knockoffs (OATK), a new methodology for detecting important explanatory variables in linear regression problems (including ridge and lasso) while controlling the false discovery rate (FDR). For each explanatory variable, OATK generates a knockoff design matrix that preserves the Gram matrix by replacing one-at-a-time only the single corresponding column of the original design matrix. OATK is a substantial relaxation and simplification of the knockoff filter, which simultaneously generates all columns of the knockoff design matrix to satisfy a much larger set of constraints. To test each variable's importance, statistics are then constructed by comparing the original versus knockoff coefficients. Under a mild correlation assumption on the original design matrix, we prove that OATK asymptotically controls the FDR at any desired level. Moreover, numerical results across a variety of simulation examples and a real genetics data set demonstrate that OATK provides good FDR control and regularly achieves (often substantially) higher power than existing approaches. In particular, in a genome-wide association study, OATK yielded more uniquely discovered genetic mutations associated with virological response to HIV therapeutics that were corroborated in clinical studies. Generating knockoffs one-at-a-time also has substantial computational advantages and facilitates additional enhancements, such as conditional calibration or derandomization, to further improve power and consistency of FDR control.
Spatial clustering is crucial in disease mapping by identifying subregions with different patterns of disease incidence or mortality. This study proposes a novel Bayesian spatial clustering method for multivariate spatial disease data, which allows for understanding geographic variations of multivariate disease patterns while accounting for both spatial information and dependence among multiple disease measurements. We develop a new random tele-connected graph partition model with an unknown number of clusters, which is capable of encouraging locally contiguous clusters and allowing for remote subregions to be clustered together. We use this prior in a Bayesian hierarchical model to detect spatial clusters and estimate cluster-specific disease patterns and dependence across the multivariate disease variables. We develop a tailored Markov chain Monte Carlo (MCMC) algorithm for posterior inference, utilizing efficient doubly split-merge samplers taking advantage of graph algorithms. We illustrate our method with simulation studies and apply it to investigate the clustering patterns of county-level prostate cancer mortality rate decline across six southern U.S. states from 1985 to 2014.
The analysis of contemporary longitudinal data problems often involves high-dimensional measurements of time-course data collected on a small number of observations. The estimation of such models with limited sample sizes can pose substantial challenges and lead to highly variable estimates, unstable predictions, and low power. In such instances, it is natural to borrow information from additional datasets with similar, although not necessarily identical, covariate-outcome relations to improve inference in the target population. We develop a novel Bayesian transfer learning model for longitudinal data (BTLL). BTLL leverages mixture models for the discrepancies between the pivotal parameters of the outcome models of the target and source studies to enhance estimation accuracy and enable data-adaptive information borrowing. BTLL aims to minimize the negative transfer of information from source studies that would otherwise introduce large bias into inference in the target population. Through extensive simulations and real data applications, we show that BTLL improves the precision of parameter estimates in the target study substantially compared to alternative methods. Moreover, our BTLL reduces the estimation bias in heterogeneous settings when the outcome distributions of the target and source datasets deviate from each other. Motivated by the progressive nature of dementia and the importance of early detection, we develop a BTLL prediction model using datasets for individuals with mild cognitive impairment. We show that BTLL can leverage non-invasive biomarkers to identify subjects at high risk for dementia and consistently achieves the lowest prediction error.
This commentary has three aims. First, we review existing forms of exposure mappings and their associated positivity assumptions. Second, we examine some limitations of the proposed method. Third, we reinterpret results and provide two potentially useful methodological applications of the work.