Motivated by the analysis of data from a clinical trial on patients with early breast cancer, we propose in this paper a new joint model that uses a Tobit partly linear mixed model for longitudinal measurements which are bounded in an interval and have a nonlinear relationship with the observation times and a semiparametric mixture cure model that incorporates a B-spline baseline hazard for survival times with cure proportion. A procedure is developed for estimating parameters in the proposed model using the partial likelihood and Laplace approximation. Additionally, a method of random weighting is proposed to compute the variances of the parameter estimators. The performance of the proposed model and the inference procedures is evaluated through simulation studies and data from the clinical trial that motivated this study.
Extrinsic Gaussian process regression methods, such as wrapped Gaussian process, have been developed to analyze manifold data. However, there is a lack of intrinsic Gaussian process methods for studying complex data with manifold-valued response variables. In this paper, we first apply the parallel transport operator on Riemannian manifold to propose an intrinsic covariance structure that addresses a critical aspect of constructing a well-defined Gaussian process regression model. We then propose a novel intrinsic Gaussian process regression model for manifold-valued data, which can be applied to data situated not only on Euclidean submanifolds but also on manifolds without a natural ambient space. We establish the asymptotic properties of the proposed models, including information consistency and posterior consistency, and we also show that the posterior distribution of the regression function is invariant to the choice of orthonormal frames for the coordinate representations of the covariance function. Numerical studies, including simulation and real examples, indicate that the proposed methods work well.
Handling massive datasets poses a significant challenge in modern data analysis, particularly within epidemiology and medicine. In this study, we introduce a novel approach using sequential ensemble learning to effectively analyze extensive datasets. Our method prioritizes efficiency from both statistical and computational perspectives, addressing challenges such as data communication and privacy, as discussed in federated learning literature. To demonstrate the efficacy of our approach, we present compelling real-world examples using COVID-19 data alongside simulation studies.
Kriging with interpolation is widely used in various noise-free areas, such as computer experiments. However, owing to its Gaussian assumption, it is susceptible to outliers, which affects statistical inference, and the resulting conclusions could be misleading. Little work has explored outlier detection for kriging. Therefore, we propose a novel kriging method for simultaneous outlier detection and prediction by introducing a normal-gamma prior, which results in an unbounded penalty on the biases to distinguish outliers from normal data points. We develop a simple and efficient method, avoiding the expensive computation of the Markov chain Monte Carlo algorithm, to simultaneously detect outliers and make a prediction. We establish the true identification property for outlier detection and the consistency of the estimated hyperparameters in kriging under the increasing domain framework as if the number and locations of the outliers were known in advance. Under appropriate regularity conditions, we demonstrate information consistency for prediction in the presence of outliers. Numerical studies and real data examples show that the proposed method generally provides robust analyses in the presence of outliers. Supplementary materials for this article are available online.
Gaussian processes (GPs) are powerful tools to nonparameterically model stochastic functions. With regard to computation load, especially for massive data with large sample size, recent researches have applied spectral representations with stationary kernels to build low-rank approaches to approximate GPs. However, these low-rank methods may struggle to capture the spatially varying properties of the data in reality due to their finite representations. In this paper, we propose a novel framework for the low-rank GPs, where a normal-gamma prior is applied instead of the standard Gaussian prior for the coefficients in the spectral representation. The normal-gamma prior promotes flexibility, robustness and sparsity in the low-rank representations of GPs. An efficient hierarchical likelihood approach is also developed for estimation, which maintains the computational benefits of the low-rank approximation. To handle the multidimensional covariates, we extend this approach to additive GPs. Numerical studies and real data examples demonstrate the performance and applicability of the proposed approaches.
In this paper, we study a high-dimensional quantile working model under a semi-supervised setting, where a relatively small-size labeled data and a large-size of unlabeled data are available. We propose a semi-supervised debiased least absolute shrinkage and selection operator (LASSO) method to estimate the target parameter in a model-free framework. Under suitable conditions, we establish the limiting properties of the proposed semi-supervised estimator. Furthermore, we demonstrate that the proposed semi-supervised estimator is more efficient than the supervised estimator under model misspecification and achieves equivalent efficiency when the working model is correctly specified. Simulation studies are conducted to examine the finite-sample performance of the proposed method. An application is illustrated with an analysis of an electronic health record dataset.
Longitudinal proportional data, which are restricted in a closed interval, are frequently observed together with survival data in clinical trials and other medical studies. In this paper, we propose a new model for the joint analysis of longitudinal proportional and survival data. This model uses generalized mean-variance mixed model for longitudinal outcomes, which avoids parametric assumption for their distribution, and Cox proportional hazards model for survival times. A procedure is developed to estimate the parameters in the proposed model based on quasi-likelihood for longitudinal data and partial likelihood for survival data with a Laplace approximation for the joint likelihood. A random weighting method is proposed to calculate the variance of these parameter estimators. The performance of the proposed model and estimation procedures are assessed through simulation studies and the application to the analysis of data from a randomized clinical trial on early breast cancer.
Many methods have been developed to analyze complex data, such as non-Euclidean shape, network, and manifold data. However, there is a lack of methods for studying interactions among complex data. In this article, we first propose a novel kernel function for a metric space and construct its associated reproducing kernel Hilbert space. The new nonstationary kernel function provides a flexible and powerful tool for learning complex structures in non-Euclidean data. We then construct an analysis of variance (ANOVA) decomposition of the nonparametric regression function defined on metric space, which provides a hierarchical structure for investigating the main effects and interactions. We develop estimation and computational methods for a semi-parametric model with a multivariate function on a product of metric spaces modeled by the ANOVA decomposition. We establish the convergence rates of parameter and nonparametric function estimates. The application of the proposed methods to the Alzheimer's Disease Neuroimaging Initiative hippocampus shape data confirms some existing and suggests some new interactions among hippocampal regions. Simulations indicate that the proposed methods work well. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
A crucial step in developing personalized strategies in precision medicine or precision marketing is to identify the latent subgroups of patients or customers of a heterogeneous population. In this article, we consider a general class of heterogeneous transformation models for subgroup identification, under which an unknown monotonic transformation of the response is linearly related to the covariates via subject-specific regression coefficients with unknown error distribution. This class of models is broad enough to cover many popular models, including the heterogeneous linear model, the heterogeneous Cox's proportional hazard model and the heterogeneous proportional odds model. Without any priori grouping information, we propose a robust method based on the maximum rank correlation and a concave fusion to automatically determine the number of subgroups, identify the latent subgroup structure, and estimate the subgroup-specific covariate effects simultaneously. We establish the theoretical properties of our proposed estimate under regularity conditions. A random weighting resampling scheme is used for variance estimation. The proposed procedure can be easily extended to handle censored data. Numerical studies including simulations and two real data analysis demonstrate that the proposed method performs reasonably well in practical situations.
Data collected from distributed sources or sites commonly have different distributions or contaminated observations. Active learning procedures allow us to assess data when recruiting new data into model building. Thus, combining several active learning procedures together is a promising idea, even when the collected data set is contaminated. Here, we study how to conduct and integrate several adaptive sequential procedures at a time to produce a valid result via several machines or a parallel‐computing framework. To avoid distraction by complicated modelling processes, we use confidence set estimation for linear models to illustrate the proposed method and discuss the approach's statistical properties. We then evaluate its performance using both synthetic and real data. We have implemented our method using Python and made it available through Github at https://github.com/zhuojianc/dsep .
Regression models with functional responses and covariates have attracted extensive research. Nevertheless, there is no existing method for the situation where the functional covariates are bivariate functions with one of the variables in common with the response function. In this article, we propose a nonparametric function-on-function regression method. We construct model spaces using a Gaussian kernel function and smoothing spline ANOVA decomposition. We estimate the nonparametric function using penalized likelihood and study properties of the Gaussian kernel function and the convergence rate of the proposed estimation method. We evaluate the proposed methods using simulations and illustrate them using two real data examples.
Multivariate regression models have been extensively studied in the literature and applied in practice. It is not unusual that some predictors may make the same nonnull contributions to all the elements of the response vector, especially when the number of predictors is very large. For convenience, we call the set of such predictors as the homogeneity set. In this paper, we consider a sparse high-dimensional multivariate generalized linear models with coexisting homogeneity and heterogeneity sets of predictors, which is very important to facilitate the understanding of the effects of different types of predictors as well as improvement on the estimation efficiency. We propose a novel adaptive regularized method by which we can easily identify the homogeneity set of predictors and investigate the asymptotic properties of the parameter estimation. More importantly, the proposed method yields a smaller variance for parameter estimation compared to the ones that do not consider the existence of a homogeneity set of predictors. We also provide a computational algorithm and present its theoretical justification. In addition, we perform extensive simulation studies and present real data examples to demonstrate the proposed method.
Robust estimation approaches are of fundamental importance for statistical modelling. To reduce susceptibility to outliers, we propose a robust estimation procedure with t-process under functional ANOVA model. Besides common mean structure of the studied subjects, their personal characters are also informative, especially for prediction. We develop a prediction method to predict the individual effect. Statistical properties, such as robustness and information consistency, are studied. Numerical studies including simulation and real data examples show that the proposed method performs well.
在实际应用数据中,我们所收集到的纵向数据可能受到测量上下限的影响,出现双侧删失的情形.此外,纵向数据的观测时间可能会携带响应变量部分的信息量.因此,在统计分析中,结合观测时间信息来分析双删失纵向数据十分必要.本文采用半参数回归模型来分析此类观测时间带信息的双删失纵向数据,通过构造鞅差估计方程来估计模型中的回归系数.同时,也指出参数估计量的渐近分布含有未知冗余参数.为进一步研究所得估计量的统计推断,本文应用随机加权方法计算估计量的渐近分布,并说明在给定样本下,随机加权估计的条件渐近分布与参数估计的渐近分布相同.为了计算参数的估计值,受交替乘子算法思想的启迪,本文建立一种新的参数计算算法,并通过数据模拟验证所提随机加权方法的可行性.最后,利用该模型分析临床试验中晚期大肠癌患者的生活质量数据,指出观测次数和治疗方法对患者生活质量有显著的影响.
Many methods have been developed to study nonparametric function-on-function regression models. Nevertheless, there is a lack of model selection approach to the regression function as a functional function with functional covariate inputs. To study interaction effects among these functional covariates, in this article, we first construct a tensor product space of reproducing kernel Hilbert spaces and build an analysis of variance (ANOVA) decomposition of the tensor product space. We then use a model selection method with the L1 criterion to estimate the functional function with functional covariate inputs and detect interaction effects among the functional covariates. The proposed method is evaluated using simulations and stroke rehabilitation data.
Identification of a subgroup of patients who may be sensitive to a specific treatment is an important problem in precision medicine. This article considers the case where the treatment effect is assessed by longitudinal measurements, such as quality of life scores assessed over the duration of a clinical trial, and the subset is determined by a continuous baseline covariate, such as age and expression level of a biomarker. Recently, a linear mixed threshold regression model has been proposed but it assumes the longitudinal measurements are normally distributed. In many applications, longitudinal measurements, such as quality of life data obtained from answers to questions on a Likert scale, may be restricted in a fixed interval because of the floor and ceiling effects and, therefore, may be skewed. In this article, a threshold longitudinal Tobit quantile regression model is proposed and a computational approach based on alternating direction method of multipliers algorithm is developed for the estimation of parameters in the model. In addition, a random weighting method is employed to estimate the variances of the parameter estimators. The proposed procedures are evaluated through simulation studies and applications to the data from clinical trials.
The analysis of data stored in multiple sites has become more popular, raising new concerns about the security of data storage and communication. Federated learning, which does not require centralizing data, is a common approach to preventing heavy data transportation, securing valued data, and protecting personal information protection. Therefore, determining how to aggregate the information obtained from the analysis of data in separate local sites has become an important statistical issue. The commonly used averaging methods may not be suitable due to data nonhomogeneity and incomparable results among individual sites, and applying them may result in the loss of information obtained from the individual analyses. Using a sequential method in federated learning with distributed computing can facilitate the integration and accelerate the analysis process. We develop a data-driven method for efficiently and effectively aggregating valued information by analyzing local data without encountering potential issues such as information security and heavy transportation due to data communication. In addition, the proposed method can preserve the properties of classical sequential adaptive design, such as data-driven sample size and estimation precision when applied to generalized linear models. We use numerical studies of simulated data and an application to COVID-19 data collected from 32 hospitals in Mexico, to illustrate the proposed method.
Motivated by comparing the distribution of longitudinal quality of life (QoL) data among different treatment groups from a cancer clinical trial, we propose a semiparametric test statistic for the homogeneity of the distributions of multigroup longitudinal measurements, which are bounded in a closed interval with excess observations taking the boundary values. Our procedure is based on a three-component mixed density ratio model and a composite empirical likelihood for the longitudinal data taking values inside the interval. A nonparametric bootstrap method is applied to calculate the p-value of the proposed test. Simulation studies are conducted to evaluate the proposed procedure, which show that the proposed test is effective in controlling type I errors and more powerful than the procedure which ignores the values on the boundaries. It is also robust to the model mispecification than the parametric test. The proposed procedure is also applied to compare the distributions of the scores of Physical Function subscale and Global Heath Status between the patients randomized to two treatment groups in a cancer clinical trial.