Functional principal component analysis (FPCA) is an important technique for dimension reduction in functional data analysis (FDA). Classical FPCA method is based on the Karhunen-Loève expansion, which assumes a linear structure of the observed functional data. However, the assumption may not always be satisfied, and the FPCA method can become inefficient when the data deviates from the linear assumption. In this paper, we propose a novel FPCA method that is suitable for data with a nonlinear structure by neural network approach. We construct networks that can be applied to functional data and explore the corresponding universal approximation property. The main use of our proposed nonlinear FPCA method is curve reconstruction. We conduct a simulation study to evaluate the performance of our method. The proposed method is also applied to two real-world data sets to further demonstrate its superiority.
Background:Sepsis biomarker research over the past 30 years has been plagued by the use of wrong animal models and inappropriate patient selections, leading to the failure of translating findings into precision medicine. Thousands of sepsis-related gene biomarkers have been published, but this excess hinders medical advancement because (1) an overwhelming number of genes make targeted drug development and precision medicine unfeasible; (2) many biomarkers lack cross-cohort validation, rendering them clinically unhelpful. Our goal is to identify a highly informative, single-digit set of sepsis biomarkers to advance precision medicine. Methods:We conducted large-scale research on heterogeneous populations, including patients with sepsis, severe sepsis, and septic shocks, and collected plasma samples from 32 sepsis patients and 18 healthy controls at Renmin Hospital of Wuhan University, China. RNA was isolated using the HYCEZMBIO Serum/Plasma RNA Kit, and RT-qPCR was performed on the Roche Light Cycler 480 platform. An AI-based max-logistic competing classifier was applied across 11 cohorts with thousands of samples, using both self-designed and public datasets to identify the most critical sepsis biomarkers. Results:Our analysis highlights CKAP4, FCAR, and RNF4 as key genetic drivers in sepsis-related variations. In whole blood, NONO is crucial for immune response, while in plasma, PLEKHO1 and BMP6 reveal further genetic heterogeneities. Pediatric patients also exhibit significant contributions from RNASE2 and OGFOD3. These genes form the most effective miniature set of biomarkers. Conclusion:Achieving 99.42% accuracy across cohorts, this miniature set outperforms larger published gene sets. These findings provide critical insights for personalized risk assessment, targeted drug development, and tailored treatments for both adult and pediatric sepsis patients.
Deep neural networks have a wide range of applications in data science. This paper reviews neural network modeling algorithms and their applications in both supervised and unsupervised learning. Key examples include: (i) binary classification and (ii) nonparametric regression function estimation, both implemented with feedforward neural networks ($\mathrm{FNN}$); (iii) sequential data prediction using long short-term memory ($\mathrm{LSTM}$) networks; and (iv) image classification using convolutional neural networks ($\mathrm{CNN}$). All implementations are provided in $\mathrm{MATLAB}$, making these methods accessible to statisticians and data scientists to support learning and practical application.
This article introduces the "autoregressive conditional generalized beta distribution of the second kind" (AcGB2 ) to model the dynamic cross-sectional maxima of multivariate financial times series. The temporal dependence of the resulting univariate maxima time series is characterized by the parameter dynamics of the standard GB2 distribution, which offer versatility in approximating various distributions, including heavy-tailed distributions, through parameter adjustments. Consequently, the newly proposed AcGB2 modeling enhances flexibility in fitting real data, especially in scenarios where extreme value theory conditions are not met, as demonstrated through simulations. Real data analysis is conducted on three datasets: two with medium-sized cross-sectional dimensions (30 or fewer) and one with high dimension (100 or higher). These datasets are based on the daily negative simple-returns of 30 stocks in the Dow Jones Industrial Average, stocks in S&P 100, and stocks from 22 primary dealers, respectively. The article establishes stationary and ergodic solutions for this new time series model under mild parameter conditions and derives the consistency, asymptotic normality, and uniqueness of the statistical conditional maximum likelihood estimators.
In longitudinal studies, it is common that the response and the covariateare not measured at the same time, which complicates the subsequent analysis.In this study, we consider the estimation of a generalized varying coefficientmodel with such asynchronous observations. We construct a penalized kernel-weighted estimating equation using the kernel technique in a functional data analysisframework. Moreover, we consider local sparsity in the estimating equation toimprove the interpretability of the estimate. We extend the iteratively reweightedleast squares algorithm in our computation, and establish the theoretical propertiesof the proposed method, including the consistency, sparsistency, and asymptoticdistribution. Lastly, we use simulation studies to verify the performance of ourmethod, and demonstrate the method by applying it to data from a study onwomen's health
Motivated by inferring causal relationships among neurons using ensemble spike train data, this paper introduces a new technique for learning the structure of a directed acyclic graph (DAG) within a large network of events, applicable to diverse multi-dimensional temporal point process (MuTPP) data. At the core of MuTPP lie the conditional intensity functions, for which we construct a generative model parameterized by the graph parameters of a DAG and develop an equality-constrained estimator, departing from exhaustive search-based methods. We present a novel, flexible augmented Lagrangian (Flex-AL) optimization scheme that ensures provable global convergence and computational efficiency gains over the classical AL algorithm. Additionally, we explore causal structure learning by integrating acyclicity-constraints and sparsity-regularization. We demonstrate: (i) in cases without regularization, the incorporation of the acyclicity constraint is essential for ensuring DAG recovery consistency; (ii) with suitable regularization, the DAG-constrained estimator achieves both parameter estimation and DAG reconstruction consistencies similar to the unconstrained counterpart, but significantly enhances empirical performance. Furthermore, simulation studies indicate that our proposed DAG-constrained estimator, when appropriately penalized, yields more accurate graphs compared to unconstrained or unregularized estimators. Finally, we apply the proposed method to two real MuTPP datasets.
Learning the sparse network-structured dependence among nodes from multivariate point process data { T i } i∈V has wide applications in information transmission, social science, and computational neuroscience. This paper develops new continuous-time stochastic models of the conditional intensity functions {λ i ( t | Ft ): t ≥ 0} i∈V , dependent on past event counts of parent nodes, to uncover the network structure within an array of non-stationary multivariate counting processes { N ( t ): t ≥ 0} for { T i } i∈V . The stochastic mechanism is crucial for statistical inference of graph parameters relevant to structure recovery but does not satisfy the key assumptions of commonly used processes like the Poisson process, Cox process, Hawkes process, queuing model, and piecewise deterministic Markov process. We introduce a new marked point process for intensity discontinuities , derive compact representations of their conditional distributions, and demonstrate the cyclicity property of N ( t ) driven by recurrence time points. These new theoretical properties enable us to establish statistical consistency and convergence properties of the proposed penalized M -estimators for graph parameters under mild regularity conditions. Simulation evaluations demonstrate computational simplicity and increased estimation accuracy compared to existing methods. Real multiple neuron spike train recordings are analyzed to infer connectivity in neuronal networks.
Temporal dependence is frequently encountered in large-scale structured noisy data, arising from scientific studies in neuroscience and meteorology, among others. This challenging characteristic may not align with existing theoretical frameworks or data analysis tools. Motivated by multi-session fMRI time series data, this paper introduces a novel semi-parametric inference procedure suitable for a broad class of “non-stationary, non-Gaussian, temporally dependent” noise processes in time-course data. It develops a new test statistic based on a tapering-type estimator of the large-dimensional noise auto-covariance matrix and establishes its asymptotic chi-squared distribution. Our method not only relaxes the consistency requirement for the noise covariance matrix estimator but also avoids direct matrix inversion without sacrificing detection power. It adapts well to both stationary and a wider range of temporal noise processes, making it particularly effective for handling challenging scenarios involving very large scales of data and large dimensions of noise covariance matrices. We demonstrate the efficacy of the proposed procedure through simulation evaluations and real fMRI data analysis.
Summary Finding a suitable representation of multivariate data is fundamental in many scientific disciplines. Projection pursuit ( ) aims to extract interesting ‘non‐Gaussian’ features from multivariate data, and tends to be computationally intensive even when applied to data of low dimension. In high‐dimensional settings, a recent work (Bickel et al., 2018) on addresses asymptotic characterization and conjectures of the feasible projections as the dimension grows with sample size. To gain practical utility of and learn theoretical insights into in an integral way, data analytic tools needed to evaluate the behaviour of in high dimensions become increasingly desirable but are less explored in the literature. This paper focuses on developing computationally fast and effective approaches central to finite sample studies for (i) visualizing the feasibility of in extracting features from high‐dimensional data, as compared with alternative methods like and , and (ii) assessing the plausibility of in cases where asymptotic studies are lacking or unavailable, with the goal of better understanding the practicality, limitation and challenge of in the analysis of large data sets.
As a prominent dimension reduction method for multivariate linear regression, the envelope model has received increased attention over the past decade due to its modeling flexibility and success in enhancing estimation and prediction efficiencies. Several enveloping approaches have been proposed in the literature; among these, the partial response envelope model [57] that focuses on only enveloping the coefficients for predictors of interest, and the simultaneous envelope model [14] that combines the predictor and the response envelope models within a unified modeling framework, are noteworthy. In this article we incorporate these two approaches within a Bayesian framework, and propose a novel Bayesian simultaneous partial envelope model that generalizes and addresses some limitations of the two approaches. Our method offers the flexibility of incorporating prior information if available, and aids coherent quantification of all modeling uncertainty through the posterior distribution of model parameters. A block Metropolis-within-Gibbs algorithm for Markov chain Monte Carlo (MCMC) sampling from the posterior is developed. The utility of our model is corroborated by theoretical results, comprehensive simulations, and a real imaging genetics data application for the Alzheimer’s Disease Neuroimaging Initiative (ADNI) study.
Analyzing “large p small n” data is becoming increasingly paramount in a wide range of application fields. As a projection pursuit index, the Penalized Discriminant Analysis ($\mathrm{PDA}$) index, built upon the Linear Discriminant Analysis ($\mathrm{LDA}$) index, is devised in Lee and Cook (2010) to classify high-dimensional data with promising results. Yet, there is little information available about its performance compared with the popular Support Vector Machine ($\mathrm{SVM}$). This paper conducts extensive numerical studies to compare the performance of the $\mathrm{PDA}$ index with the $\mathrm{LDA}$ index and $\mathrm{SVM}$, demonstrating that the $\mathrm{PDA}$ index is robust to outliers and able to handle high-dimensional datasets with extremely small sample sizes, few important variables, and multiple classes. Analyses of several motivating real-world datasets reveal the practical advantages and limitations of individual methods, suggesting that the $\mathrm{PDA}$ index provides a useful alternative tool for classifying complex high-dimensional data. These new insights, along with the hands-on implementation of the $\mathrm{PDA}$ index functions in the R package classPP, facilitate statisticians and data scientists to make effective use of both sets of classification tools.
Statistical data analysis and machine learning heavily rely on error measures for regression, classification, and forecasting. Bregman divergence ( BD ) is a widely used family of error measures, but it is not robust to outlying observations or high leverage points in large- and high-dimensional datasets. In this paper, we propose a new family of robust Bregman divergences called “ robust - BD ” that are less sensitive to data outliers. We explore their suitability for sparse large-dimensional regression models with incompletely specified response variable distributions and propose a new estimate called the “ penalized robust - BD estimate ” that achieves the same oracle property as ordinary non-robust penalized least-squares and penalized-likelihood estimates. We conduct extensive numerical experiments to evaluate the performance of the proposed penalized robust- BD estimate and compare it with classical approaches, and show that our proposed method improves on existing approaches. Finally, we analyze a real dataset to illustrate the practicality of our proposed method. Our findings suggest that the proposed method can be a useful tool for robust statistical data analysis and machine learning in the presence of outliers and large-dimensional data.
本文针对带有组结构的广义线性稀疏模型,引入布雷格曼散度作为一般性的损失函数,进行参数估计和变量选择,使得该方法不局限于特定模型或特定的损失函数.本文比较研究了Ridge,SACD,Lasso,自适应Lasso,组Lasso,分层Lasso,自适应分层Lasso和稀疏组Lasso共8种惩罚函数的特点和引入模型后参数估计和变量选择的方法,并给出了分层Lasso的坐标轴下降算法和稀疏组Lasso的加速全梯度更新算法.模拟研究验证了组Lasso,分层Lasso,自适应分层Lasso和稀疏组Lasso能更好的利用数据的组结构信息,自适应分层Lasso和稀疏组Lasso在变量选择准确性,参数估计精度方面优于其它方法,稀疏组Lasso在模型预测精度上达到最优.作为实证研究,本文将带有稀疏组Lasso惩罚的逻辑斯蒂模型应用于骨关节炎患者的外周血单核细胞基因表达水平的分析,选出了9个基因集中共136个基因与骨关节炎有关,以期对后续生物医学研究有一定指导价值.
In longitudinal study, it is common that response and covariate are not measured at the same time, which complicates the analysis to a large extent. In this paper, we take into account the estimation of generalized varying coefficient model with such asynchronous observations. A penalized kernel-weighted estimating equation is constructed through kernel technique in the framework of functional data analysis. Moreover, local sparsity is also considered in the estimating equation to improve the interpretability of the estimate. We extend the iteratively reweighted least squares (IRLS) algorithm in our computation. The theoretical properties are established in terms of both consistency and sparsistency, and the simulation studies further verify the satisfying performance of our method when compared with existing approaches. The method is applied to an AIDS study to reveal its practical merits.
This paper develops the empirical likelihood (EL) inference procedure for parameters in autoregressive models with the error variances scaled by an unknown nonparametric time-varying function. Compared with existing methods based on non-parametric and semi-parametric estimation, the proposed test statistic avoids estimating the variance function, while maintaining the asymptotic chi-square distribution under the null. Simulation studies demonstrate that the proposed E L procedure (a) is more stable, i.e., depending less on the change points in the error variances, and (b) gets closer to the desired confidence level, than the traditional test statistic.
Improving estimation efficiency for regression coefficients is an important issue in the analysis of longitudinal data, which involves estimating the covariance matrix of the within-subject errors. In the balanced or nearly balanced setting, we can also regard the covariance matrix of the dependent errors as the bivariate covariance function evaluated at specific time points. In this paper, we compare the performance of the proposed regularized-covariance-function-based estimator and the conventional high-dimensional covariance matrix estimator of the within-subject errors. It shows that when the number p of the time points in each subject is large enough compared to the number n of the subjects, i.e., p≫n1/4logn, the estimation errors of the high-dimensional covariance matrix will be accumulated, therefore the error bound of the proposed regularized-covariance-function-based estimator will be smaller than that of the high-dimensional covariance matrix estimator in Frobenius norm. We also assess the performance of these two estimators for the incomplete longitudinal data. All the comparisons and theoretical results are illustrated using both simulated and real data.
The exact distribution is typically unavailable for a two-sample t-statistic in a single test for equal population means if we have nonGaussian samples, unequal population variances, or unequal sample sizes n(1) and n(2). In this case, a calibration method using a reference distribution offers a practically feasible substitute. This study simultaneously calibrates a diverging number m of two-sample t-statistics for inferences of significance in high-dimensional data from a small sample. For the Gaussian calibration method, we demonstrate the following. First, the simultaneous "general" two-sample t-statistics achieve the overall significance level, as long as log(m) increases at a strictly slower rate than (n(1) + n(2))(1/3) as n(1) + n(2) diverges. Second, directly applying the same calibration method to simultaneous "pooled" two-sample t-statistics may substantially lose the overall level accuracy. The proposed "adaptively pooled" two-sample t-statistics overcome such incoherence, while operating as simply and performing as well as the "general" two-sample t-statistics. Third, we propose a "two-stage" t-test procedure to effectively alleviate the skewness commonly encountered in various two-sample t-statistics in practice, thus increasing the calibration accuracy. Lastly, we discuss the implications of these results using simulation studies and real-data applications.