This paper is concerned with Spearman's correlation matrices under large dimensional regime, in which the data dimension diverges to infinity proportionally with the sample size. We establish the central limit theorem for the linear spectral statistics of Spearman's correlation matrices, which extends the results of [Ann. Statist. 43(2015) 2588–2623]. We also study the improved Spearman's correlation matrices [Ann. Math. Statist 19(1948) 293–325] which is a standard U-statistic of order 3. As applications, we propose three new test statistics for large dimensional independent test and numerical studies demonstrate the applicability of our proposed methods.
Single-cell and spatial transcriptomics provide high-resolution cellular characterization, yet standard analytical approaches remain theoretically misaligned with the probabilistic nature of the data. After UMI normalization, current pipelines rely on Euclidean or log-transformed Euclidean distance for similarity measurement. Both are fundamentally ill-suited to model the multinomial count data. Euclidean distance in normalized space overemphasizes high-variance genes, while log-transformation inverts this bias but at the cost of distorting subtle, continuous expression modulations. Neither approach naturally captures the dual nature of gene expression: both discrete presence/absence transitions and continuous quantitative variation. To overcome these limitations, we introduce GAIA (Geometric Analysis from an Information Aspect), an information-geometric framework for cell representation learning and inter-cell similarity measurement. By anchoring analysis in the true probabilistic model, treating cells as multinomial distributions over genes and projecting cells to a statistical manifold, GAIA organically reconciles both the presence/absence effect and the more continuous expression modulations. Mathematically, GAIA exploits the equivalence between Fisher-Rao distance in multinomial space and geodesic distance on the unit hypersphere, a property that enables both theoretical guarantees and computational efficiency. Experiments in synthetic and real scRNA-seq and spatial transcriptomic datasets demonstrate that GAIA preserves robust and consistent cell-to-cell relationships, delineates biologically nuanced sub-types, mitigates batch effects arising from sequencing depth variation, and eliminates the dependence on knowledge-restricted gene selection for learning meaningful cell representations. Overall, GAIA offers a knowledge-lean, variance-stabilizing framework for analyzing single-cell and spatial transcriptomic data, enhancing discrimination between nuanced cell sub-type and -states.
In the field of group intention prediction, faced with increasingly complex, agile and adversarial battlefield environment, recognition methods based on manual experience and cognitive ability lack timeliness, accuracy and objectivity. In this paper, the Lightweight Random Forest (LWRF) model is used to realize the technical task autonomy of a battle group based on target state, target threat and group target components. Target clustering, group formation recognition, group threat calculation and the LWRF classification model are used to determine the tactical mission of the enemy battle group target. Through data simulation of the relevant parameters, the overall accuracy of intention prediction can reach 99.73%, demonstrating effective prediction performance.
Datasets containing both categorical and continuous variables are frequently encountered in many areas, and with the rapid development of modern measurement technologies, the dimensions of these variables can be very high. Despite the recent progress made in modelling high-dimensional data for continuous variables, there is a scarcity of methods that can deal with a mixed set of variables. To fill this gap, this paper develops a novel approach for classifying high-dimensional observations with mixed variables. Our framework builds on a location model, in which the distributions of the continuous variables conditional on categorical ones are assumed Gaussian. We overcome the challenge of having to split data into exponentially many cells, or combinations of the categorical variables, by kernel smoothing, and provide new perspectives for its bandwidth choice to ensure an analogue of Bochner's Lemma, which is different to the usual bias-variance tradeoff. We show that the two sets of parameters in our model can be separately estimated and provide penalized likelihood for their estimation. Results on the estimation accuracy and the misclassification rates are established, and the competitive performance of the proposed classifier is illustrated by extensive simulation and real data studies.
The proliferation of rumors on social networks undermines information credibility. While their dissemination forms complex networks, current detection methods struggle to capture these intricate propagation patterns. Representing each node solely through its textual embeddings neglects the textual coherence across the entire rumor propagation path, which compromises the accuracy of rumor identification on social platforms. We propose a novel framework that leverages Large Language Models (LLMs) to address these limitations. Our approach captures subtle rumor signals by employing LLMs to analyze information subchains, assign rumor probabilities and intelligently construct connections to virtual nodes. This enables the modification of the original graph structure, which is a critical advancement for capturing subtle rumor signals. Given the inherent limitations of LLMs in rumor identification, we develop a structured prompt framework to mitigate model biases and ensure robust graph learning performance. Additionally, the proposed framework is model-agnostic, meaning it is not constrained to any specific graph learning algorithm or LLMs. Its plug-and-play nature allows for seamless integration with further fine-tuned LLMs and graph techniques in the future, potentially enhancing predictive performance without modifying original algorithms.
Consider Schott's statistic (Schott, 2005) defined as the squared Frobenius norm of the sample correlation matrix for data from α-regularly varying populations. We investigate its asymptotic distribution in a general framework characterized by data dimension p, sample size n, and regularly varying coefficients α. In particular, we identify a phase transition phenomenon in the asymptotic behavior. For light-tailed populations (α> 3), we revisit the α-free asymptotic distribution but relax the constraint on the ratio of p/n. For heavy-tailed populations (α< 3), we derive a new asymptotic normal distribution whose variance explicitly depends on α. We also propose a consistent estimator for the asymptotic variance such that the standardized Schott's test statistic remains applicable for unknown location parameters and all α> 0.
This paper investigates the spectral properties of spatial-sign covariance matrices, a self-normalized version of sample covariance matrices, for data from alpha-regularly varying populations with general covariance structures. By exploiting the elegant properties of self-normalized random variables, we establish the limiting spectral distribution and a central limit theorem for linear spectral statistics. We demonstrate that the Marcenko-Pastur equation holds under the condition alpha >= 2, while the central limit theorem for linear spectral statistics is valid for alpha > 4, which are shown to be nearly the weakest possible conditions for spatial-sign covariance matrices from heavy-tailed data in the presence of dependence.
This paper investigates the spectral properties of spatial-sign covariance matrices, a self-normalized version of sample covariance matrices, for data from α-regularly varying populations with general covariance structures. By exploiting the elegant properties of self-normalized random variables, we establish the limiting spectral distribution and a central limit theorem for linear spectral statistics. We demonstrate that the Marc̆enko-Pastur equation holds under the condition α≥ 2, while the central limit theorem for linear spectral statistics is valid for α>4, which are shown to be nearly the weakest possible conditions for spatial-sign covariance matrices from heavy-tailed data in the presence of dependence.
DC capacitors (DCCs) are the key components in power electronic transformers (PETs), which maintain the PETs operating normally. It is necessary to monitor the healthy condition of DCCs. Changes in the healthy condition of DCCs result in varied values of capacitance (C) and equivalent series resistance (ESR). The existing data-driven condition monitoring methods for DCCs can only work accurately for a particular working condition of PETs but fail to work with variable working conditions. Moreover, PETs are generally required to operate uninterruptedly, and it is impossible to measure the values of C and ESR, resulting in the data of DCCs' voltage without labels. To address the above problems, a novel transferrable data-driven method is proposed. There are three parts in this method: the source-domain double-scale convolutional autoencoder (SDCAE), adversarial learning network (ALN), and extreme learning machine (ELM). First, the feature of source-domain data is extracted by the SDCAE and employed as the prior distribution of the target-domain data. Then, ALN is employed to minimize the distribution divergence between the source-domain data and target-domain data. After that, ELM is trained by feature and label of source-domain data and employed to estimate the C and ESR of target-domain data. Finally, the proposed transferrable data-driven method is verified by the simulation and experimental data of a three-phase AC-DC PET.
Fréchet regression has received considerable attention to model metric-space valued responses that are complex and non-Euclidean data, such as probability distributions and vectors on the unit sphere. However, existing Fréchet regression literature focuses on the classical setting where the predictor dimension is fixed, and the sample size goes to infinity. This paper proposes sparse Fréchet sufficient dimension reduction with graphical structure among high-dimensional Euclidean predictors. In particular, we propose a convex optimization problem that leverages the graphical information among predictors and avoids inverting the high-dimensional covariance matrix. We also provide the Alternating Direction Method of Multipliers (ADMM) algorithm to solve the optimization problem. Theoretically, the proposed method achieves subspace estimation and variable selection consistency under suitable conditions. Extensive simulations and a real data analysis are carried out to illustrate the finite-sample performance of the proposed method.
Differential network analysis plays a crucial role in capturing nuanced changes in conditional correlations between two samples. Under the high-dimensional setting, the differential network, that is, the difference between the two precision matrices are usually stylized with sparse signals and some low-rank latent factors. Recognizing the distinctions inherent in the precision matrices of such networks, we introduce a novel approach, termed 'SR-Network' for the estimation of sparse and reduced-rank differential networks. This method directly assesses the differential network by formulating a convex empirical loss function with & ell;1$$ {\ell}_1 $$-norm and nuclear norm penalties. The study establishes finite-sample error bounds for parameter estimation and highlights the superior performance of the proposed method through extensive simulations and real data studies. This research significantly contributes to the advancement of methodologies for accurate analysis of differential networks, particularly in the context of structures characterized by sparsity and low-rank features.
This paper investigates the efficient solution of penalized quadratic regressions in high-dimensional settings. A novel and efficient algorithm for ridge-penalized quadratic regression is proposed, leveraging the matrix structures of the regression with interactions. Additionally, an alternating direction method of multipliers (ADMM) framework is developed for penalized quadratic regression with general penalties, including both single and hybrid penalty functions. The approach simplifies the calculations to basic matrix-based operations, making it appealing in terms of both memory storage and computational complexity for solving penalized quadratic regressions in high-dimensional settings.
Independent component model (ICM) is a widely-used population distribution in high-dimensional data analysis and random matrix theory. In this work, we study the kurtosis of the ICM, which is an important parameter in asymptotic distributions of many commonly used statistics and is also an important criterion for measuring the heavy-tailed nature of the data. Based on U-statistics, we develop an estimation method. Theoretically, we show that the proposed estimator is consistent under regular conditions, especially we relax the restriction that the data dimension and the sample size are of the same order. Furthermore, we derive the asymptotic normality of the estimator, which allows us to construct confidence intervals and hypothesis testing statistics. Computationally, we provide a fast and efficient algorithm by leveraging the matrix structure, where the computational complexity is essentially equivalent to computing the sample covariance matrix and the Gram matrix. Finally, the effectiveness of the proposed method is demonstrated through simulated data and real-world data.
Multivariate elliptically-contoured distributions are widely used for modeling correlated and non-Gaussian data. In this work, we study the kurtosis of the elliptical model, which is an important parameter in many statistical analysis. Based on U-statistics, we develop an estimation method. Theoretically, we show that the proposed estimator is consistent under regular conditions, especially we relax a moment condition and the restriction that the data dimension and the sample size are of the same order. Furthermore, we derive the asymptotic normality of the estimator and evaluate the asymptotic variance through several examples, which allows us to construct a confidence interval. The performance of our method is validated by extensive simulations and real data analysis.
High-dimensional matrix-valued data is common in scientific and engineering studies and its classification is a significant topic in current statistics. In practice, the discriminative signals of the matrix covariates are oftentimes low rank and sparse. Motivated by this, we propose a sparse and reduced-rank matrix linear discriminant analysis called "Sr-LDA" for binary classification of high-dimensional matrix-valued data. Specifically, based on the Bayes' linear discriminant rule, we derive the theoretically optimal discriminative matrix-valued covariates under the matrix normal assumptions, and constructed a convex empirical loss function for the estimation of the optimal discriminative matrix-valued covariates under the l(1)-norm and nuclear norm penalties. Finite sample error bounds for parameter estimation and the misclassification rate are established. The superior performance of the proposed Sr-LDA is illustrated via extensive simulation and real data studies with comparison to other state-of-the-art classifiers.
In this work, we consider the discriminant analysis under the weak sparsity where many entries of the parameters are nearly zero. We develop a unified LASSO-typed framework to estimate the parameters for sparse discriminant analysis and derive the general non-asymptotic error bound under the weak sparsity condition. As applications, we revisit the sparse linear discriminant analysis and sparse quadratic discriminant analysis. We establish the consistency of the estimators and also the misclassification error rate. These results extend the sparse discriminant analysis methods to weak sparsity setting and refine the existing theoretical results.
In this paper, we investigate the limiting spectral distribution of a high-dimensional Kendall’s rank correlation matrix. The underlying population is allowed to have a general dependence structure. The result no longer follows the generalized Marc̆enko-Pastur law, which is brand new. It is the first result on rank correlation matrices with dependence. As applications, we study Kendall’s rank correlation matrix for multivariate normal distributions with a general covariance matrix. From these results, we further gain insights into Kendall’s rank correlation matrix and its connections with the sample covariance/correlation matrix.
Random matrix theory provides new insights into multiple scattering in random media. In a recent study, we demonstrated the statistical separation of single- and multiple-scattering components based on a Wishart random matrix. The first- and second-order moments were estimated through a Wishart random matrix constructed using dynamically-backscattered speckle images. In this study, this new strategy was applied to laser speckle contrast imaging (LSCI) of in-vivo blood flow. The random matrix-based method was adapted and parameterized using electric field Monte Carlo simulations and in-vitro blood flow phantom experiments. The new method was further applied in in-vivo experiments, demonstrating the benefits of separating the single- and multiple-scattering components, and was compared with the traditional temporal LASCA method. More specifically, the new method captures stimulus-induced functional changes in blood flow and tissue perfusion in the superficial and deeper layers. The new method extends the ability of LSCI to image functional and pathological changes.
ImmunoPET imaging, which combining the specificity of monoclonal antibody and high sensitivity of PET imaging, blazes new trails of molecular imaging modality. In recent years, clinical translation and application of immunoPET imaging strategies have been flourishing, as a result, facilitating early and non-invasive diagnosis of numerous human tumors, patient stratification before monoclonal antibody therapy and radiation dose estimation prior to radioimmunotherapy. Here, we summarize the most recent clinical evidence of immunoPET in tumor diagnosis and therapy.