High-dimensional mediation analysis (HMA) seeks to uncover complex causal mechanisms involving numerous mediators and plays a crucial role in scientific and social sciences. In this work, we introduce the Generative Adversarial High-dimensional Mediation Network (GAHMN), a novel, scalable structured generative framework designed for causal analysis in high-dimensional settings. GAHMN formulates mediation analysis as dual conditional generative blocks, explicitly capturing mediators' dual roles as outcomes influenced by treatments and as predictors affecting outcomes. Each block integrates a high-dimensional partially linear structure with multi-channel convolutional layers, promoting effective parameter sharing and enhanced representation learning. To induce sparsity and accurate mediator selection, GAHMN employs customized min-max optimization problems with L1 penalties on generator parameters, alongside specially designed optimization algorithms for efficient computation. Unlike existing benchmark methods relying on restrictive parametric assumptions or random-effect specifications, GAHMN flexibly captures heterogeneity, complex distributions, and inter-mediator correlations. With careful design, the computational complexity of GAHMN scales linearly with the number of mediators p, rather than quadratically as in conventional approaches. Theoretical results rigorously ensures estimation consistency, convergence rate, and accurate sparse recovery. GAHMN also serves as a structured generative causal modeling framework, extending to causal decomposition, structural equation modeling, and counterfactual policy evaluation. Extensive experiments confirm GAHMN's superior performance and robustness in synthetic and real-world scenarios.
This study proposes a maximum penalized likelihood procedure for simultaneous estimation and variable selection in the context of Cox proportional hazards models with informative right-censored data.A copula function is adopted to model the dependence between censoring and event times.Moreover, two penalty functions are introduced to accommodate the sparsity of regression coefficients and smooth the baseline hazard estimate.Since the baseline hazard function is nonnegative, we propose a specific algorithm comprising a modified Newton algorithm for updating regression coefficients and a multiplicative iterative algorithm for updating baseline hazard at each iteration.Furthermore, we establish the asymptotic properties of the proposed estimators.Simulation studies show that the proposed method performs satisfactorily.Finally, we apply the proposed method to investigate the potential risk factors of AIDS for HIV-1-infected patients from the AIDS Clinical Trials Group Protocol 175 study.
This study investigates the assessment of causal treatment effects on interval-censored failure time outcomes. Available techniques for this problem primarily rely on full-likelihood-based approaches, which can be computationally burdensome and prone to model misspecification for the compliance group. Motivated by a breast cancer screening study, we propose a weighted likelihood estimator for causal treatment effects under the proportional hazards model. The proposed approach employs the inverse weighting scheme, offering simplicity and enhancing computational efficiency compared to the existing methods. Also it can be implemented by using standard software. We establish the consistency and asymptotic normality of the resulting estimator for the regression parameter. Extensive simulation studies are conducted to evaluate the finite-sample performance of the proposed approach, which suggests that it performs well. Finally, we apply the proposed method to the breast cancer screening study mentioned above and obtain some new insights about the effect of periodic screening on reducing the risk of breast cancer-related mortality.
This study introduces a joint modeling framework that integrates structural equation model (SEM) with multi-state semi-competing risks model to investigate the complex causal mechanism among exposure, latent mediators, and multiple hierarchically correlated survival outcomes. The framework first employs a linear SEM to characterize latent mediators underlying highly correlated observable surrogates. These latent mediators are then incorporated into semi-parametric proportional hazards models with frailty to assess their mediating roles in the relationships between exposure and hierarchically correlated survival outcomes. Estimation and inference are conducted within an integrated Bayesian framework leveraging Monte Carlo methods, Gibbs sampling, and the Metropolis-Hastings algorithm. Extensive simulation studies demonstrate the robustness and accuracy of the proposed method. Application to the Framingham Heart Study uncovers causal pathways linking smoking and metabolic health factors and evaluates their combined impact on long-term cardiovascular outcomes, offering new insights into the complex interplay between lifestyle factors and health risks.
Deep neural networks (DNNs) have become powerful tools for modeling complex data structures through sequentially integrating simple functions in each hidden layer. In survival analysis, recent advances of DNNs primarily focus on enhancing model capabilities, especially in exploring nonlinear covariate effects under right censoring. However, deep learning methods for interval-censored data, where the unobservable failure time is only known to lie in an interval, remain underexplored and limited to specific data type or model. This work proposes a general regression framework for interval-censored data with a broad class of partially linear transformation models, where key covariate effects are modeled parametrically while nonlinear effects of nuisance multi-modal covariates are approximated via DNNs, balancing interpretability and flexibility. We employ sieve maximum likelihood estimation by leveraging monotone splines to approximate the cumulative baseline hazard function. To ensure reliable and tractable estimation, we develop an EM algorithm incorporating stochastic gradient descent. We establish the asymptotic properties of parameter estimators and show that the DNN estimator achieves minimax-optimal convergence. Extensive simulations demonstrate superior estimation and prediction accuracy over state-of-the-art methods. Applying our method to the Alzheimer's Disease Neuroimaging Initiative dataset yields novel insights and improved predictive performance compared to traditional approaches.
Extracting underlying signals from imaging data is crucial for medical applications. Medical imaging data can be contaminated by outliers or heavy-tailed noise, and the irregular domains of such data pose additional challenges. Ordinary least squares (OLS) methods are highly sensitive to outliers or heavy-tailed noise, while existing robust estimation approaches encounter difficulties with medical imaging data defined on irregular domains. To this end, we develop a novel robust estimation method by integrating M-estimation with bivariate penalized splines over triangulations. Under mild regularity conditions, we establish the L2 convergence and asymptotic normality of the proposed M-estimator. Simulation studies demonstrate that the proposed method significantly outperforms OLS when the errors do not follow a normal distribution, while maintaining comparable computational efficiency. Applications to imaging data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) validate the practical value of the proposed method.
High-frequency intraday data is crucial in financial modeling, but its high dimensionality and noise present significant challenges for pattern recognition and effective use in machine learning. In this paper, motivated by the stylized characteristics of intraday trading, we propose a novel dimension reduction method called price-volume balanced functional principal component analysis (PVB-FPCA) in terms of cumulative intraday return (CIDR) curves. Through PVB-FPCA, we make a key discovery: when crucial volume-price characteristics are incorporated, high-frequency intraday data reveals a remarkably robust and significant intraday trend pattern, described by the first functional principal component. By exploiting this pattern, we develop a network module that generates efficient low-dimensional representations of high-frequency intraday data. This approach addresses the challenges of high dimensionality and noise, enhancing the utility of high-frequency data. Building upon the module, we develop a novel deep temporal forecasting model, termed PVB-IntraFusionNet. We design a dual-channel temporal convolutional network (TCN) to extract low-frequency temporal price-volume features. These features are subsequently fused with high-frequency features via a cross multi-head attention mechanism to capture multi-scale dependencies. Through a carefully designed projection module, the fused features are utilized to generate dynamic forecasting results. Compared to state-of-the-arts, PVB-IntraFusionNet can efficiently leverage both low-and high-frequency patterns simultaneously, resulting in significantly improved forecasting performance. Through PVB-IntraFusionNet, this paper explores an effective approach for developing more accurate deep temporal learning models with high-frequency intraday data, demonstrating the potential of our method in addressing various financial learning problems. The numerical results on prominent NYSE/NASDAQ stocks further illustrate the efficiency and advantages of our approach.
In functional data analysis (FDA), the functional linear regression model (FLRM) is a popular method to describe the relationship between a scalar response and a functional predictor. The conventional approach for estimating FLRM relies on the normality assumption of the error terms. However, such analyzes may not yield robust inference when the linearity and normality assumptions are violated. In this paper, we develop a partially functional linear additive regression model for handling right- or left-censored data, where the error term follows scale mixtures of normal distributions. We use the B-spline method to approximate the additive nonparametric functions and the functional principal component analysis to estimate the slope function of the functional predictor. A Bayesian approach coupled with an efficient Markov chain Monte Carlo (MCMC) algorithm is developed to conduct statistical inference. The performance of the proposed method is evaluated through simulation studies and applied to a laryngeal squamous cell carcinoma study.
Spatial domain identification requires jointly modeling molecular signatures and physical coordinates, yet current tools frequently over-smooth biological boundaries, require user-specified cluster numbers, and lack principled multimodal integration. We introduce BaySC, an integrative Bayesian spatial clustering framework for spatial domain identification. BaySC inherently learns the true number of spatial domains from the data by employing a Mixture of Finite Mixtures (MFM) prior. Tissue topology is modeled via a Markov Random Field (MRF) applied to discrete cellular assignments, a strategy that enforces local spatial coherence without distorting the underlying gene expression features. This enables BaySC to accurately map contiguous tissue layers as well as geographically scattered, transcriptionally identical cell populations. Furthermore, BaySC handles spatial multi-omics data through a weighted log-likelihood fusion mechanism executed via Gibbs sampling. This approach assigns interpretable weights to each modality, allowing users to quantify the biological relevance of different data layers to the final tissue map. Validated across ten single-modal spatial transcriptomics and two spatial multi-omics datasets, BaySC yields highly interpretable probabilistic outputs. It demonstrates competitive accuracy on standard clustering metrics and consistently outperforms existing tools in preserving spatial topography, as measured by spatially-aware Adjusted Rand Index (spARI).
One of the most conspicuous developments in the unprecedented worldwide epidemic of COVID-19 is the pressing demand for reliable diagnostic tools. Utilizing artificial intelligence (AI) and image processing algorithms, this work proposes a novel 19-layer Convolutional Neural Network (CNN) for accurate COVID-19 detection from chest X-ray images. This CNN architecture supports structure with single/multiple labels for three classes (i.e., for classification between layers like viral pneumonia, normal, and COVID-19) and four classes (i.e., lung opacity, normal, COVID-19, and pneumonia). Across the accuracy, specificity, precision, sensitivity, confusion matrix, F1-score, and other metrics, our model was compared to periutils net from the literature such as popular pre-trained networks(Inception, AlexNet, ResNet50, SqueezeNet, VGG19). Experimental results show that the proposed CNN outperforms existing methods, providing an effective diagnostic tool with the potential for clinical usage. This means AI algorithms as advanced as Cogito could start making decisions about COVID-19 and inform clinicians about how to handle the case.
Mixed membership models are frequently utilized to capture complex individual heterogeneity in multivariate and longitudinal data. A key aspect of mixed membership modeling involves determining the number of extreme profiles (classes), a task traditionally managed through inefficient criterion-based methods. This task is particularly challenging when the predictors within the models are latent and derived from multiple observed variables using exploratory factor analysis. In this paper, we consider an innovative mixed membership latent variable model, which consists of an exploratory factor model to identify latent factors and a mixed membership model with latent predictors. We develop an efficient approach that integrates parameter estimation and model selection for the number of factors, extreme profiles, and the structure of the factor loading matrix. Our approach comprises a modified stochastic search item selection algorithm to automatically determine the number of latent factors and their associated manifest variables and a Bayesian penalized method to select the number of extreme profiles. We validate our methodology through extensive simulation studies, demonstrating its accuracy and efficiency in both parameter estimation and model selection. Applying this method to data from the Parkinson's Progression Markers Initiative, we identify clinically important latent traits and distinct disease profiles. The results underscore our model's enhanced ability to depict the intricate individual heterogeneity present in Parkinson's disease patients.
Mediation analysis with high-dimensional DNA methylation markers is critical for uncovering epigenetic pathways linking environmental exposures to health outcomes. While methodological advances have enabled mediation analysis with high-dimensional mediators for time-to-event data, existing approaches fail to account for joint mediator effects adequately. This limitation hinders the identification of mediators that are jointly dependent but marginally independent of the outcome. We propose a novel high-dimensional mediation analysis framework within a causal Cox proportional hazards model to address this gap. The proposed model integrates a non-marginal sure screening technique, a minimum approximated information criterion penalty for variable selection, and an adaptive bootstrap testing procedure to assess indirect effects. Simulation studies show that the proposed method performs satisfactorily. Applied to The Cancer Genome Atlas (TCGA) lung cancer cohort study, the proposed approach identifies candidate methylation markers with potential mediation roles in the association between tobacco smoking and reduced survival in lung cancer patients.
There exists a substantial body of literature that discusses regression analysis of interval-censored failure time data and also many methods have been proposed for handling the presence of a cured subgroup. However, only limited research exists on the problems incorporating change points, with or without a cured subgroup, which can occur in various contexts such as clinical trials where disease risks may shift dramatically when certain biological indicators exceed specific thresholds. To fill this gap, we consider a class of partly linear transformation models within the mixture cure model framework and propose a sieve maximum likelihood estimation approach using Bernstein polynomials and piecewise linear functions for inference. Additionally, we provide a data-driven adaptive procedure to identify the number and locations of change points and establish the asymptotic properties of the proposed method. Extensive simulation studies demonstrate the effectiveness and practical utility of the proposed methods, which are applied to the real data from a breast cancer study that motivated this work.
Estimating heterogeneous treatment effects has drawn increasing attention in medical studies, considering that patients with divergent features can undergo a different progression of disease even with identical treatment. Such heterogeneity can co-occur with a cured fraction for biomedical studies with a time-to-event outcome and further complicates the quantification of treatment effects. This study considers a joint framework of Bayesian causal forest and accelerated failure time cure model to capture the cured proportion and treatment effect heterogeneity through three separate Bayesian additive regression trees. Under the potential outcomes framework, conditional and sample average treatment effects within the uncured subgroup are derived on the scale of log survival time subject to right-censoring, and treatment effects on the scale of survival probability are derived for each individual. Bayesian backfitting Markov chain Monte Carlo algorithm with the Gibbs sampler is conducted to estimate the causal effects. Simulation studies show the satisfactory performance of the proposed method. The proposed model is then applied to a breast cancer dataset extracted from the SEER database to demonstrate its usage in detecting heterogeneous treatment effects and cured subgroups. Combined with popular mitigation strategies, the proposed method can also alleviate confounding induced by immortal time bias.
The envelope model has gained significant attention since its proposal, offering a fresh perspective on dimension reduction in multivariate regression models and improving estimation efficiency. One of its appealing features is its adaptability to diverse regression contexts. This article introduces the integration of envelope methods into the factor analysis model. In contrast to previous research primarily focused on the frequentist approach, the study proposes a Bayesian approach for estimation and envelope dimension selection. A Metropolis-within-Gibbs sampling algorithm is developed to draw posterior samples for Bayesian inference. A simulation study is conducted to illustrate the effectiveness of the proposed method. Additionally, the proposed methodology is applied to the ADNI dataset to explore the relationship between cognitive decline and the changes occurring in various brain regions. This empirical application further highlights the practical utility of the proposed model in real-world scenarios.
The widely adopted dimension reduction technique, functional principal component analysis (FPCA), typically represents functional data as a linear combination of functional principal components (FPCs) and their corresponding scores. However, this linear formulation is too restrictive to reflect reality because it fails to capture the nonlinear dependence of functional data when nonlinear features are present in the data. This study develops a novel FPCA model to uncover the nonlinear structures of functional data. The proposed method can accommodate multivariate functional data observed on different domains, and multidimensional functional data with gaps and holes. To navigate the complexities of spatial structure in multidimensional functional variables, tensor product smoothing and spline smoothing over triangulation are employed, providing precise tools for approximating nonparametric function. Furthermore, an efficient estimation approach and theory are developed when the number of FPCs diverges to infinity. To assess its performance comprehensively, extensive simulations are conducted, and the proposed method is applied to real data from the Alzheimer's Disease Neuroimaging Initiative study, affirming its practical efficacy in uncovering and interpreting nonlinear structures inherent in functional data.
The information extracted from imaging data has become increasingly important in disease diagnosis as it uncovers associations between imaging features and diseases of interest. This study proposes a partial functional Tobit censored quantile regression (PFTCQR) model to investigate the quantile-specific relationships between the time of incidence of laryngeal cancer and a set of imaging and clinical predictors based on the data collected from a laryngeal cancer study in the Otolaryngology Department of a tertiary hospital in Jilin Province, China. The functional principal component analysis and moment method are employed to estimate the slope and covariance functions of the functional predictors. An efficient Markov chain Monte Carlo (MCMC) algorithm is developed, leveraging the location-scale mixture representation of the asymmetric Laplace distribution (ALD) to perform the estimation. Furthermore, we extend the PFTCQR model to the composite quantile regression framework and incorporate variable selection for scalar covariates, further enhancing the robustness and efficiency of parameter estimation and improving model fitting. The proposed method is demonstrated through simulation studies and applied to the laryngeal carcinoma data. Results provide new insights into potential risk factors for laryngeal carcinoma and their effects varying across quantiles. Specific laryngeal regions are identified as significantly associated with the progression of the disease.
This article introduces a unified approach to estimating the mutual density ratio, defined as the ratio between the joint density function and the product of the individual marginal density functions of two random vectors. It serves as a fundamental measure for quantifying the relationship between two random vectors. Our method uses the Bregman divergence to construct the objective function and leverages deep neural networks to approximate the logarithm of the mutual density ratio. We establish a non-asymptotic error bound for our estimator, achieving the optimal minimax rate of convergence under a bounded support condition. Additionally, our estimator mitigates the curse of dimensionality when the distribution is supported on a lower-dimensional manifold. We extend our results to overparameterized neural networks and the case with unbounded support. Applications of our method include conditional probability density estimation, mutual information estimation, and independence testing. Simulation studies and real data examples demonstrate the effectiveness of our approach. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.