Missing data, measurement error, and population heterogeneity are pervasive challenges in analyzing data arising from modern observational studies and machine learning applications. Although these problems frequently coexist and interact, they are often treated separately in existing works. We propose a unified probabilistic framework that jointly addresses these issues utilizing deep latent variable representation. The proposed method integrates a novel hierarchical tree-routed variational autoencoder with pattern-aware latent representations and calibration-based denoising. The framework accommodates missing data mechanisms, including MCAR, MAR, and MNAR, while simultaneously learning subgroup-specific and globally shared latent structure. The introduced reconvergent routing mechanism enables selective parameters to be shared across related subpopulations, which offers flexibility as well as improved statistical efficiency. Simulation studies demonstrate substantial improvements over existing deep generative imputation approaches under complex heterogeneous missingness and measurement-error settings. The proposed framework provides a principled approach for learning from noisy and incomplete data in modern healthcare and other high-dimensional applications.
Graphical models are powerful tools for characterizing conditional dependence structures among variables with complex relationships. Although many methods have been developed under the graphical modeling framework, their validity often hinges on the quality of the data. A fundamental assumption in most existing approaches is that all variables are measured precisely, an assumption frequently violated in practice. In many applications, mismeasurement of mixed discrete and continuous variables is a common challenge. In this paper, we address error-contaminated data involving both continuous and discrete variables by proposing a mixed latent Gaussian copula graphical measurement error model. To perform inference, we develop a simulation-based expectation-maximization procedure that explicitly accounts for mismeasurement effects. We further introduce a computationally efficient refinement to reduce the computational burden. Asymptotic properties of the proposed estimator are established, and its finite-sample performance is evaluated through numerical studies.
While variable selection has received extensive attention in the literature, its exploration in the presence of response measurement error remains underexplored. In this paper, we investigate this important problem within the context of binary classification with error-prone responses. We present valid variable selection procedures to address the complexities of response errors. Leveraging validation data, we introduce both parametric and semiparametric methodologies to accommodate the mismeasurement effects. By rigorously establishing theoretical results, we offer insights and justifications of the validity of the proposed methods. By properly choosing the penalty function and regularization parameter, we demonstrate that the resulting estimators possess the oracle property. To assess the finite sample properties of the proposed methods, we conduct numerical studies that confirm the effectiveness of our proposed methods.
Function-on-scalar linear regression has been widely used to model the relationship between a functional response and multiple scalar covariates. Its utility is, however, challenged by the presence of measurement error, a ubiquitous feature in applications. Naively applying usual function-on-scalar linear regression to error-contaminated data often yields biased inference results. Further, estimation of the model parameters is complicated by the presence of inactive variables, especially when handling data with a large dimension. Building parsimonious and interpretable function-on-scalar linear regression models is in urgent demand to handle error-contaminated functional data. In this paper, we study this important problem and investigate the measurement error effects. We propose a debiased loss function, combined with a sparsity-inducing penalty function, to simultaneously estimate functional coefficients and select salient predictors. An efficient computing algorithm is developed with tuning parameters determined by data-driven methods. Under mild conditions, the asymptotic properties of the proposed estimator are rigorously established, including estimation consistency, selection consistency, and the limiting distributions. The finite sample performance of the proposed method is assessed through extensive simulation studies, and the usage of the proposed method is illustrated by a real data application.
Analyzing time-to-event data, such as cancer patient survival time, is a central task in survival analysis. Numerous modeling methods and inference strategies have been developed for various application settings, where the primary goal is to assess the relationship between survival time and covariates. However, the applicability of existing approaches is often hindered by two major challenges. First, covariates (e.g., gene expression levels) typically exhibit complex network structures. Second, they are prone to measurement error, which can substantially bias inference if ignored. To address these challenges, we developed the R package SurvGME (Survival analysis with Graphical and Measurement Error models). The package provides a comprehensive framework for survival analysis in the presence of both graphical dependence structures and measurement error. It supports a range of commonly used survival models, and its utility and performance are illustrated using a breast cancer dataset.
Partially linear single-index models prove to be flexible in facilitating various types of relationships between the outcome and covariates. However, their validity is hampered by the presence of measurement error in covariates, a feature commonly encountered in applications. In this article, we explore the use of such models to handle data subject to measurement error. In addition, with multivariate covariates, often a few of them are informative while most of them are not. In this article, we propose the three stage procedure to eliminate measurement error effects and select important variables for both the linear predictor term and the single-index part. To implement the proposed method efficiently, we develop a boosting algorithm to select variables and estimate the parameters without handling non-differentiable penalty functions. Theoretical results, including consistency and asymptotic normality of the estimator, are established to justify the validity of the proposed method. Numerical studies, including simulation and data analysis, are conducted to assess the finite sample performance of the proposed method. Supplementary materials for this article are available online.
Causal inference has gained extensive attention in various fields, including healthcare, epidemiology, and social sciences. While many methods have been developed, most research has been directed to handle data with a univariate response variable. This paper highlights the complication of handling causal inference with bivariate responses in the context of longitudinal studies, which can be further challenged by the presence of missingness and censoring. We begin by defining the overall treatment effect on bivariate responses and conceptualizing the treatment as having two components, each operating through different causal pathways. By using the decomposed treatment framework, we break down the overall treatment effect into the separable treatment effects on each response, which offers us a transparent interpretation with the sum of separable treatment effects equals twice the overall treatment effects. We establish that the separable treatment effect on each response can be identified using the observed data, provided our identification conditions. Subsequently, we employ the likelihood method to estimate the separable treatment effects and derive a hypothesis testing procedure to compare them. Finally, we conduct real data analysis and simulation studies to demonstrate the effectiveness of the proposed methods.
Multivariate regression models are commonly used to examine associations in multivariate data, and various methods have been proposed to characterise distinct features of such data across different settings. The validity of those methods, however, is compromised by the presence of measurement error. Despite extensive research on measurement error in univariate data, the impact of measurement error on the analysis of multivariate data remains an interesting topic to explore. This paper rigorously examines the measurement error effects on multivariate regression models and quantifies the asymptotic bias and covariance matrix for the na & iuml;ve method that ignores measurement error. We further develop three estimation methods to correct the measurement error effects under different scenarios, including the case with instrument variables. The asymptotic properties of these methods are established accordingly. Lastly, extensions that apply nonparametric techniques to investigate the relationship between responses and covariates contaminated by measurement error are discussed.
Precision medicine is an innovative approach that aims to customize medical treatments and interventions to patients based on their individual characteristics. Several estimation techniques, including Q-learning, have been developed to determine optimal treatment rules. However, the applicability of these methods depends on the availability of precisely measured variables. This study extends the scope of Q-learning to incorporate compound outcomes, deviating from the commonly assumed univariate outcomes, and further accommodates data with mismeasurement in both binary and continuous covariates. Two methods are described to mitigate the impact of mismeasurement. Numerical studies reveal that mismeasurement in covariates leads to notable estimation bias in parameters indexing the optimal treatment, yet the methods addressing the mismeasured effects yield improved results.
Dynamic treatment regimes (DTRs) are sequences of functions that formalize the process of precision medicine. DTRs take as input patient information and output treatment recommendations. A major focus of the DTR literature has been on the estimation of optimal DTRs, the sequences of decision rules that result in the best outcome in expectation, across the complete population if they were to be applied. While there is a rich literature on optimal DTR estimation, to date, there has been minimal consideration of the impacts of nonadherence on these estimation procedures. Nonadherence refers to any process through which an individual's prescribed treatment does not match their true treatment. We explore the impacts of nonadherence and demonstrate that, generally, when nonadherence is ignored, suboptimal regimes will be estimated. In light of these findings, we propose a method for estimating optimal DTRs in the presence of nonadherence. The resulting estimators are consistent and asymptotically normal, with a double robustness property. Using simulations, we demonstrate the reliability of these results, and illustrate comparable performance between the proposed estimation procedure adjusting for the impacts of nonadherence and estimators that are computed on data without nonadherence.
Boosting techniques have gained increasing interest in both machine learning and statistical research. However, many of these methods are primarily designed for complete datasets, which limits their applicability to handle incomplete data such as missing observations. In this paper, we propose the pseudo-outcome strategy to account for missingness effects and describe a functional gradient descent algorithm. Numerical studies demonstrate the satisfactory performance of the proposed method in finite sample settings. Furthermore, we apply the proposed method to analyze the KLIPS Data.