In this article, we propose a testing procedure for detecting the inequality of covariance functions between two-samples of large-scale functional data. The asymptotic null distribution of the test statistic is established under mild conditions, and the power of the test is shown to be consistent under quite general alternatives. The proposed test benefits from several advantages that distinguish it from the conventional framework of functional data analysis. The first advantage is named as “eigenvalue-decay-free” since none of the conditions are imposed on the decaying pattern of the eigenvalues of each functional data. The second advantage is regarded as “square-integrable-free” since we do not require any functional data to be square-integrable. Other advantages include, but are not limited to, the permission of ultra-high-dimensionality, fairly different sample sizes, and non-Gaussian functional observations. We evaluate the numerical performance of the proposed test by a simulation study as well as a real data application.
For covariance test in functional data analysis, existing methods are developed only for fully observed curves, whereas in practice, trajectories are typically observed discretely and with noise. To bridge this gap, we employ a pool-smoothing strategy to construct an FPC-based test statistic, allowing the number of estimated eigenfunctions to grow with the sample size. This yields a consistently nonparametric test, while the challenge arises from the concurrence of diverging truncation and discretized observations. Facilitated by advancing perturbation bounds of estimated eigenfunctions, we establish that the asymptotic null distribution remains valid across permissable truncation levels. Moreover, when the sampling frequency (i.e., the number of measurements per subject) reaches certain magnitude of sample size, the test behaves as if the functions were fully observed. This phase transition phenomenon differs from the well-known result of the pooling mean/covariance estimation, reflecting the elevated difficulty in covariance test due to eigen-decomposition. The numerical studies, including simulations and real data examples, yield favorable performance compared to existing methods.
Reliable Bayesian predictive inference has long been an open problem under unidentified transformation models, since the Markov Chain Monte Carlo (MCMC) chains of posterior predictive distribution (PPD) values are generally poorly mixed. We address the poorly mixed PPD value chains under unidentified transformation models through an adaptive scheme for prior adjustment. Specifically, we originate a conception of sufficient informativeness, which explicitly quantifies the information level provided by nonparametric priors, and assesses MCMC mixing by comparison with the within-chain MCMC variance. We formulate the prior information level by a set of hyperparameters induced from the nonparametric prior elicitation with an analytic expression, which is guaranteed by asymptotic theory for the posterior variance under unidentified transformation models. The analytic prior information level consequently drives a hyperparameter tuning procedure to achieve MCMC mixing. The proposed method is general enough to cover various data domains through a multiplicative error working model. Comprehensive simulations and real-world data analysis demonstrate that our method successfully achieves MCMC mixing and outperforms state-of-the-art competitors in predictive capability.
Most of existing methods of functional data classification deal with one or a few processes. In this work we tackle classification of high-dimensional functional data, in which each observation is potentially associated with a large number of functional processes, p, which is comparable to or even much larger than the sample size n. The challenge arises from the complex inter-correlation structures among multiple functional processes, instead of a diagonal correlation for a single process. Since truncation is often needed for approximation in functional data, another difficulty stems from the fact that the discriminant set of the infinite-dimensional optimal classifier may be different from that of the truncated optimal classifier, when multiple (especially a large number of) processes are involved. We bridge the gap by proposing a penalized classifier that achieves both near-perfect classification that is unique to functional data, and discriminant set inclusion consistency in the sense that the classification-responsible functional predictors include those of the underlying optimal classifier. Simulation study and real data application are carried out to demonstrate its favorable performance. for this article are available online.
In a functional partial linear regression (FPLR) model, where the response variable is scalar while the explanatory variables involve both infinite‐dimensional functional predictors and finite‐dimensional scalar covariates, the relationships between the response and the explanatory variables are often assumed to be the same for all subjects. This article relaxes this assumption and considers a subgroup analysis for the FPLR model, which allows the intercepts to vary for different subgroups from a heterogeneous population. By projecting the functional predictors onto the corresponding eigenspace, the subgroup analysis based on the FPLR model can be simplified to a framework that is similar to the classical subgroup analysis problem. To automatically identify subgroups among observations and estimate the regression parameters of interest, we combine the functional principal component analysis with the concave pairwise penalized approach and develop an ADMM algorithm for functional subgroup analysis. We also establish the consistency of the proposed estimators under mild conditions. Simulation experiments demonstrate that the concave penalized subgroup approach could potentially achieve substantial gains over the ordinary FPLR model. The analysis of data from a creative achievement study is used to illustrate the practical performance of the subgroup analysis for the FPLR model.
High-dimensional covariance matrices have attracted much attention of statisticians and econometricians during the past decades. Vast literature is devoted to the research in high-dimensional covariance matrices. However, most of them are for constant covariance matrices. In many applications, constant covariance matrices are not appropriate, e.g., in portfolio allocation, dynamic covariance matrices would make much more sense. Simply assuming each entry of a covariance matrix is a function of time to introduce a dynamic structure would not work. In this paper, we are going to introduce a class of high-dimensional dynamic covariance matrices in which a kind of additive structure is embedded. We will show the proposed high-dimensional dynamic covariance matrices have many advantages in applications. An estimation procedure is also proposed to estimate the proposed high-dimensional dynamic covariance matrices. Asymptotic properties are built to justify the proposed estimation procedure. Intensive simulation studies show the proposed estimation procedure works very well when sample size is finite. Finally, we apply the proposed high-dimensional dynamic covariance matrices, together with the proposed estimation procedure, to portfolio allocation. The results look very interesting.
Portfolio allocation is an important topic in financial data analysis. In this article, based on the mean-variance optimization principle, we propose a synthetic regression model for construction of portfolio allocation, and an easy to implement approach to generate the synthetic sample for the model. Compared with the regression approach in existing literature for portfolio allocation, the proposed method of generating the synthetic sample provides more accurate approximation for the synthetic response variable when the number of assets under consideration is large. Due to the embedded leave-one-out idea, the synthetic sample generated by the proposed method has weaker within sample correlation, which makes the resulting portfolio allocation more close to the optimal one. This intuitive conclusion is theoretically confirmed to be true by the asymptotic properties established in this article. We have also conducted intensive simulation studies in this article to compare the proposed method with the existing ones, and found the proposed method works better. Finally, we apply the proposed method to real datasets. The yielded returns look very encouraging.
Nonparametric transformation models (NTMs) have sparked much interest in survival prediction owing to their flexibility with both transformations and error distributions unspecified. However, fitting these models has been hampered because they are unidentified. Existing approaches typically constrain the parameter space to en-sure identifiablity, but they incur intractable computation and cannot scale up to complex data; other approaches address the identifiablity issue by making strong a priori assumptions on either of the nonparametric components, and thus are subject to misspecifications. Utilizing a Bayesian workflow, we address the challenge by constructing new weakly informative nonparametric priors for infinite-dimensional parameters so as to remedy flat likelihoods associated with unidentified models. To facilitate applicability of these new priors, we subtly impose an exponential transformation on top of NTMs, which compresses the space of infinite-dimensional parameters to positive quadrants while maintaining interpretability. We further develop a cutting-edge posterior modification technique for estimating the fully identified parametric component. Simulations reveal that our method is robust and outperforms the competing methods, and an application to a Veterans’ lung cancer dataset suggests that our method can predict survival time well and help develop clinically meaningful risk scores, based on patients’ demographic and clinical predictors.
This article tackles the old problem of prediction via a nonparametric transformation model (NTM) in a new Bayesian way. Estimation of NTMs is known challenging due to model unidentifiability though appealing because of its robust prediction capability in survival analysis. Inspired by the uniqueness of the posterior predictive distribution, we achieve efficient prediction via the NTM aforementioned under the Bayesian paradigm. Our strategy is to assign weakly informative priors to nonparametric components rather than identify the model by adding complicated constraints in the existing literature. The Bayesian success pays tribute to i) a subtle cast of NTMs by an exponential transformation for the purpose of compressing spaces of infinite-dimensional parameters to positive quadrants considering non-negativity of the failure time; ii) a newly constructed weakly informative quantile-knots I-splines prior for the recast transformation function together with the Dirichlet process mixture model assigned to the error distribution. In addition, we provide a convenient and precise estimator for the identified parameter component subject to the general unit-norm restriction through posterior modification, enabling effective relative risks. Simulations and applications on real datasets reveal that our method is robust and outperforms the competing methods. An R package BuLTM is available to predict survival curves, estimate relative risks, and facilitate posterior checking.
Combination of multiple biomarkers to improve diagnostic accuracy is meaningful for practitioners and clinicians, and are attractive to lots of researchers. Nowadays, with development of modern techniques, functional markers such as curves or images, play an important role in diagnosis. There exists rich literature developing combination methods for continuous scalar markers. Unfortunately, only sporadic works have studied how functional markers affect diagnosis in the literature. Moreover, no publication can be found to do combination of multiple functional markers to improve the diagnostic accuracy. It is impossible to apply scalar combination methods to the multiple functional markers directly because of infinite dimensionality of functional markers. In this article, we propose a one-dimension scalar feature motivated by square loss distance, as an alternative of the original functional curve in the sense that, it can retain information to the most extent. The square loss distance is defined as the function of projection scores generated from functional principal component decomposition. Then existing variety of scalar combination methods can be applied to scalar features of functional markers after dimension reduction to improve the diagnostic accuracy. Area under the receiver operating characteristic curve and Youden index are used to assess performances of various methods in numerical studies. We also analyzed the high- or low- hospital admissions due to respiratory diseases between 2010 and 2017 in Hong Kong by combining weather conditions and media information, which are regarded as functional markers. Finally, we provide an R function for convenient application.
Multivariate functional data has received considerable attention but testing for equality of mean surfaces and its profile has limited progress. The existing literature has tested equality of either mean curves of univariate functional samples directly, or mean surfaces of bivariate functional data samples but turn into functional curves comparison again. In this paper, we aim to develop both the profile and globe tests of mean surfaces for two-sample bivariate functional data. We present valid approaches of tests by employing the idea of pooled projection and by developing a novel profile functional principal component analysis tool. The proposed methodology enjoys the merit of readily interpretability and implementation. Under mild conditions, we derive the asymptotic behaviors of test statistics under null and alternative hypotheses. Simulations show that the proposed tests have a good control of the type I error by the size and can detect difference in mean surfaces and its profile effectively in terms of power in finite samples. Finally, we apply the testing procedures to two real data sets associated with the precipitation change affected jointly by time and locations in the Midwest of USA, and the trends in human mortality from European period life tables.
As an extension of partially linear models and additive models, partially linear additive model is useful in statistical modelling. This paper proposes an empirical likelihood based approach for testing serial correlation in this semiparametric model. The proposed test method can test not only zero first-order serial correlation, but also higher-order serial correlation. Under the null hypothesis of no serial correlation, the test statistic is shown to follow asymptotically a chi-square distribution. Furthermore, a simulation study is conducted to illustrate the performance of the proposed method.
This paper studies Bayesian inference on longitudinal mixed effects models with non-normal AR(1) errors. We model the nonparametric zero-mean noise in the autoregression residual with a Dirichlet process (DP) mixture model. Applying the empirical likelihood tool, an adjusted sampler based on the Pólya urn representation of DP is proposed to incorporate information of the moment constraints of the mixing distribution. A Gibbs sampling algorithm based on the adjusted sampler is proposed to approximate the posterior distributions under DP priors. The proposed method can easily be extended to address other moment constraints owing to the wide application background of the empirical likelihood. Simulation studies are used to evaluate the performance of the proposed method. Our method is illustrated via the analysis of a longitudinal dataset from a psychiatric study.
This article considers testing serial correlation in partially linear additive errors-in-variables model. Based on the empirical likelihood based approach, a test statistic was proposed, and it was shown to follow asymptotically a chi-square distribution under the null hypothesis of no serial correlation. Finally, some simulation studies are conducted to illustrate the performance of the proposed method.
As a generalization of additive model and partially linear model, partially linear additive model has been paid considerably attention in recent years. This paper considers estimation of the parametric component of the semiparametric model when the covariates in the linear part are measured with additive error and some additional stochastic linear restrictions on the parametric component are available. Based on the corrected profile least-squares approach and mixed regression estimation method, we propose a corrected profile mixed estimator for the parametric component, and derive its asymptotic distribution. Finally, some simulation studies are conducted to illustrate the proposed procedure and the results are satisfactory.
In this paper, we characterize approximate solutions of vector optimization problems with set-valued maps. We gives several characterizations of generalized subconvexlike set-valued functions(see [10), which is a generalization of nearly subconvexlike functions introduced in [34]. We present alternative theorem and derived scalarization theorems for approximate solutions with generalized subconvexlike set-valued maps. And then, Lagrange multiplier theorems under generalized Slater constraint qualification are established.
In this paper, we point out some deficiencies in a recent paper (Lee and Kim in J. Nonlinear Convex Anal. 13:599–614, 2012 ), and we establish strong duality and converse duality theorems for two types of nondifferentiable higher-order symmetric duals multiobjective programming involving cones.
In this paper, we establish a strong duality theorem for Mond-Weir type multiobjective higher order nondifferentiable symmetric dual programs. Our works correct some deficiencies in recent papers [higher-order symmetric duality in nondifferentiable multiobjective programming problems, J. Math. Anal. Appl. 290(2004)423-435] and [A note on higher-order nondifferentiable symmetric duality in multiobjective programming, Appl. Math. Letters 24(2011) 1308-1311].
In this work, we established a converse duality theorem for higher-order Mond-Weir type multiobjective programming involving cones. This fills some gap in recently work of Kim et al. [Kim D S, Kang H S, Lee Y J, et al. Higher order duality in multiobjective programming with cone constraints. Optimization, 2010, 59: 29–43].
In this paper, we obtain two new characterizations of preinvex functions. More specifically, a real valued function is preinvex function if and only if it is intermediate-point preinvex and semi-strictly quasi-preinvex; and a real valued function is preinvex function if and only if it is intermediate-point preinvex and semi-locally semi-strictly quasi-preinvex. Copyright © 2012 Watam Press.