
In this paper, we consider a random matrix of the form D-n circle dot X-n, a variance profile random matrix with variance profile D-n circle dot D-n where circle dot denotes the Hadamard product of the two matrices, D-n is a deterministic matrix and X-n is a random matrix. We call D-n circle dot X-n a missing data matrix of X-n when the entries of D-n are either 0 or 1. This framework is commonly used in various applied fields, such as biology, neuroscience and network data analysis. We study the convergence and asymptotic freeness of missing data matrices of iid, elliptic and covariance random matrices. Specifically, it is known that independent iid, elliptic and covariance matrices converge to freely independent circular, elliptic and Mar & ccaron;enko-Pastur variables, respectively. In this paper, we provide the necessary and sufficient conditions on deterministic matrices D-n for which these results to hold true for independent missing data matrices of these three types of random matrices.
This paper investigates the spectral properties of spatial-sign covariance matrices, a self-normalized version of sample covariance matrices, for data from alpha-regularly varying populations with general covariance structures. By exploiting the elegant properties of self-normalized random variables, we establish the limiting spectral distribution and a central limit theorem for linear spectral statistics. We demonstrate that the Marcenko-Pastur equation holds under the condition alpha >= 2, while the central limit theorem for linear spectral statistics is valid for alpha > 4, which are shown to be nearly the weakest possible conditions for spatial-sign covariance matrices from heavy-tailed data in the presence of dependence.
Conformal prediction (CP) offers a principled framework for quantifying predictive uncertainty with finite-sample coverage guarantees. However, when the calibration data are limited, the coverage of CP sets can deviate substantially from the nominal target. This paper introduces Enhanced Conformal Prediction (ECP), a new framework that incorporates abundant but potentially corrupted auxiliary data to recalibrate prediction sets by left-shifting the score thresholds derived from the clean calibration set, provably reducing set size while preserving finite-sample coverage guarantees. Further theoretical analysis demonstrates that ECP achieves a higher breakdown point than existing methods. Extensive experiments confirm the robustness and efficiency of ECP across a variety of settings.
In this paper, we study the spectral properties of a class of random matrices of the form Sn- = n(-1)(X 1X2 & lowast;- X 2X1 & lowast;), where X-k = Sigma(1/2)(k)Z( k), Z(k)'s are independent p x n complex-valued random matrices, and Sigma(k) are p x p positive semi-definite matrices that commute and are independent of the Z(k)'s for k = 1, 2. We assume that Z(k)'s have independent entries with zero mean and unit variance. The skew-symmetric/skew-Hermitian matrix S(n)(- )will be referred to as a random commutator matrix associated with the samples X-1 and X-2. We show that, when the dimension p and sample size n increase simultaneously, so that p/n -> c is an element of (0,infinity), there exists a limiting spectral distribution (LSD) for S-n(-), supported on the imaginary axis, under the assumptions that the joint spectral distribution of Sigma(1), Sigma(2) converges weakly. This nonrandom LSD can be described through its Stieltjes transform, which satisfies a system of Mar & ccaron;enko-Pastur-type functional equations. Moreover, we show that the companion matrix Sn+ = n(-1)(X X-1(2)& lowast; + X 2X1 & lowast;), under identical assumptions, has an LSD supported on the real line, which can be similarly characterized.
In this paper, we use dimensional reduction technique to study the central limit theory (CLT) of Hotelling's T-2 statistic. Specifically, we use a matrix denoted by U-pxq, to map q-dimensional sample vectors to a p-dimensional subspace, where q >= p or q >> p. Under the condition of p/n -> 0 as (p,n) ->infinity, we obtain the CLT of Hotelling's T-2 statistic. Moreover, the selection of U-pxq is also considered, which demonstrates good performance in the simulation studies.
This paper focuses on a large-dimensional approximate factor model with a general covariance matrix assumption of the idiosyncratic components when both the cross-section N and the time dimension T tend to infinity. First, a bias-corrected estimator for noise variance is proposed using random matrix theory, and its asymptotic normality is also established, which is free of the population distribution of the observations. Second, based on this bias-corrected noise variance estimator, the new information criteria are constructed to determine the number of factors. The consistency of the estimators for the number of factors is also proved as both N and T approach infinity. Finally, simulations and real data analysis are conducted to show the superiority and generality of our proposed estimators.
In this paper, we discuss the recurrence coefficients of the three-term recurrence relation for the orthogonal polynomials with respect to the weight functions [Formula: see text], where [Formula: see text]. In each case, [Formula: see text], [Formula: see text] and [Formula: see text] possess specific value ranges. Utilizing the ladder operator approach, we derive certain equations for the recurrence coefficients [Formula: see text] and [Formula: see text] from three compatibility conditions. Specifically, when [Formula: see text] and [Formula: see text], we formulate a difference equation for [Formula: see text] that connects it to the [Formula: see text]. Furthermore, we derive first-order (Toda evolution) and second-order differential equations for the [Formula: see text] and [Formula: see text], establishing the connection with the Painlevé IV equation when [Formula: see text]. Additionally, we further investigate the asymptotic behavior of the recurrence coefficients as [Formula: see text]. Following this, for each of the three cases, we establish bi-confluent Heun equations and demonstrate the transformation of the Heun equation into a Painlevé IV equation when [Formula: see text]. Subsequently, we explore the properties of the zeros of the orthogonal polynomials with the given weight function. Finally, we consider the asymptotic behavior of the smallest eigenvalue of large Hankel matrices generated by the weight function, specifically when [Formula: see text].
In the analysis of complex data from various fields like finance, imaging processing, and biomedical applications, the assumption regarding the structure of covariance holds a crucial role in ensuring accurate and efficient statistical inferences. We study the problem of evaluating whether a high-dimensional covariance matrix conforms to a linear structure defined by the linear combination of a predefined set of matrices. We introduce an innovative testing approach by integrating two Frobenius-type statistics, which capture both the difference and ratio between the unknown covariance matrix and the linearly structured matrix. Based on the joint distribution of the two test statistics, we derive the asymptotic null distribution and conduct a power analysis for this novel test under the high-dimensional setting. As evidenced by our extensive simulation studies, the proposed integrated test exhibits favorable control of the type I error rate under the null hypothesis, and demonstrates robust power across a spectrum of dense alternative hypotheses. Additionally, we further consider a power-enhanced test statistic that combines the proposed integrated test with a maximum-type test designed for sparse signals, enabling hypothesis testing against both dense and sparse alternatives.
Consider the complex Ginibre ensemble, whose eigenvalues are (lambda(i))1 <= i <= n and the spectral radius R-n =max(1 <= i <= n)|lambda i|. Set X-n = root 4 gamma n(R-n -n -(1 )/2 root gamma n) and F-n be its distribution function, where gamma(n) =log n - 2log(root 2 pi log n). It was proved in Rider [A limit theorem at the edge of a non-Hermitian random matrix ensemble, J. Phys. A 36 (2003) 3401-3409] that F(n )converges weakly to the Gumbel distribution Lambda, whose distribution function Lambda(x) = e-e-x. We prove in further in this paper that lim(n ->infinity) (log n) /loglog nW(1)(Fn, Lambda) = 2 and the Berry-Esseen bound lim(n ->infinity) (log n) /loglog nsupx is an element of & Ropf;|F-n(x) - Lambda(x)| = 2/ e.
Let M be an n & times; n random matrix with i.i.d. entries. This paper studies the deviation inequality of sn-k+1(M), the kth smallest singular value of M. In particular, when the entries of M are sub gaussian, we show that for any gamma is an element of (0, 1/2),epsilon > 0 and log n <= k <= cn P s(n-k+1)(M) <= epsilon/ root n } <= (C epsilon /k)(gamma k2) + e(-c1kn). This result improves an existing result due to Nguyen, which is a deviation inequality of s(n-k+1)(M) with (C epsilon/k)(gamma k2) + e(-cn )decay.
High-dimensional sample correlation matrices are a crucial class of random matrices in multivariate statistical analysis. The central limit theorem (CLT) provides a theoretical foundation for statistical inference. In this paper, assuming that the data dimension increases proportionally with the sample size, we derive the limiting spectral distribution of the matrix RnM and establish the CLT for the linear spectral statistics (LSS) of under the linear independent component structure. In contrast to the existing literature, our proposed spectral properties do not require M to be an identity matrix. Moreover, we also derive the joint limiting distribution of LSS of RnM1, . . . , RnMK and propose a numerical method for computing the asymptotic covariance of CLT of the LSS for rescaled sample correlation matrices. As an illustration, an application is given for the CLT.
This paper develops some theory of the Dyson equation for correlated linearizations and uses it to solve a problem on asymptotic deterministic equivalent for the test error in random features regression. The theory developed for the correlated Dyson equation includes existence-uniqueness, spectral support bounds, and stability properties. This theory is new for constructing deterministic equivalents for pseudo-resolvents of a class of linearizations with correlated entries. In the application, this theory is used to give a deterministic equivalent of the test error in random features ridge regression, in a proportional scaling regime, wherein we have conditioned on both training and test datasets.
We consider sparse inhomogeneous Erd\H{o}s-R\'enyi random graph ensembles where edges are connected independently with probability $p_{ij}$. We assume that $p_{ij}= \varepsilon_N f(w_i, w_j)$ where $(w_i)_{i\ge 1}$ is a sequence of deterministic weights, $f$ is a bounded function and $N\varepsilon_N\to \lambda\in (0,\infty)$. We characterise the limiting moments in terms of graph homomorphisms and also classify the contributing partitions. We present an analytic way to determine the Stieltjes transform of the limiting measure. The convergence of the empirical distribution function follows from the theory of local weak convergence in many examples but we do not rely on this theory and exploit combinatorial and analytic techniques to derive some interesting properties of the limit. We extend the methods of Khorunzhy et al. (2004) and show that a fixed point equation determines the limiting measure. The limiting measure crucially depends on $\lambda$ and it is known that in the homogeneous case, if $\lambda\to\infty$, the measure converges weakly to the semicircular law (Jung and Lee (2018)). We extend this result of interpolating between the sparse and dense regimes to the inhomogeneous setting and show that as $\lambda\to \infty$, the measure converges weakly to a measure which is known as the operator-valued semicircular law.
We continue our work [F. Bornemann, Asymptotic expansions of the limit laws of Gaussian and Laguerre (Wishart) ensembles at the soft edge, preprint (2024), arXiv:2403.07628] on asymptotic expansions at the soft edge for the classical n-dimensional Gaussian and Laguerre random matrix ensembles. By revisiting the construction of the associated skew-orthogonal polynomials in terms of wave functions, we obtain concise expressions for the level densities that are well suited for proving asymptotic expansions in powers of a certain parameter h asymptotic to n-2/3. In the unitary case, the expansion for the level density can be used to reconstruct the first correction term in an established asymptotic expansion of the associated generating function. In the orthogonal and symplectic cases, we can even reconstruct the conjectured first and second correction terms.
In this paper, we study a flexible functional linear regression model where the dependency of a scalar response on a functional predictor is function-valued process rather than conventional one-way processes. Additionally, we provide an intuitively appealing estimation approach to estimate the bivariate functional regression coefficient. We first represent the bivariate functional coefficient by using the data-driven bases function to achieve the dimension reduction, and then introduce an iterative least square procedure to estimate the coefficients after dimension reduction in the framework of low-rank structure. Theoretically, we investigate the convergence rate of bivariate functional coefficient estimator under mild conditions. Simulation studies indicate that the proposed methods perform well in finite samples and an empirical example is presented to illustrate its usefulness.
In this paper, we investigate a certain linear statistic of the unitary ensembles with two parameters, which is equivalent to the characterization of a sequence of polynomials orthogonal with respect to a semi-classical Laguerre weight [Formula: see text] with parameters [Formula: see text], [Formula: see text], [Formula: see text]. We explore certain transformations of the recurrence coefficients and the sub-leading coefficient for the semi-classical Laguerre polynomials, which can serve as solutions of the analogs of Painlevé IV, include the Jimbo–Miwa–Okamoto [Formula: see text]-form of Painlevé IV and the discrete [Formula: see text]-form of Painlevé IV. Using Dyson’s Coulomb fluid approach, we derive the asymptotic behaviors of the recurrence coefficients and the smallest eigenvalue of large Hankel matrices generated by the semi-classical Laguerre weight. Additionally, we reduce the second-order differential equation satisfied by the orthogonal polynomial generated by the semi-classical Laguerre weight to the biconfluent Heun equation.
Consider a random matrix of size [Formula: see text] as an additive deformation of the complex Ginibre ensemble under a deterministic matrix [Formula: see text]. Under certain assumptions on [Formula: see text], we prove that the local eigenvalue statistics in the bulk are governed by the Ginibre bulk statistics, which form a class of determinantal point process with the correlation kernel [Formula: see text] This is the continuation of the previous joint papers [D.-Z. Liu and L. Zhang, Critical edge statistics for deformed GinUEs, preprint (2023), arXiv:2311.13227v1] and [D.-Z. Liu and L. Zhang, Repeated erfc statistics for deformed GinUEs, preprint (2025), arXiv:2402.14362], which deal with the local eigenvalue statistics at the edge.
This paper discusses the multivariate Behrens-Fisher problem, which tests the equality of two population mean vectors under heteroscedasticity. The classical Wald-type test is not applicable in high dimensions owing to the singularity of sample covariance matrices. A ridgelized Wald-type (RIWT) test is then proposed. Using the exact four-moment theorem, its asymptotic null distribution is derived under moment conditions. Thus, our test can accommodate non-Gaussian distributions with general covariance matrices. Simulation results demonstrate the good performance of the proposed test.
Let Gamma(n) be an n x n Haar-invariant orthogonal matrix. Let Z(n)be the p x q upper-left submatrix of Gamma n and Gn be a p x q matrix whose pq entries are independent standard normals, where p and q are two positive integers. Let & Laplacetrf;((n)Z(n)) and & Laplacetrf;(G(n)) be their joint distributions, respectively. Consider the Fisher information I(& Laplacetrf;((n)Z(n))|& Laplacetrf;(G(n))) between the distributions of nZn and Gn. In this paper, we conclude that I(& Laplacetrf;((n)Z(n))|& Laplacetrf;(G(n)))-> 0 as n ->infinity if pq = o(n) and it does not tend to zero if c =limn ->infinity(pq)/(n )is an element of (0, +infinity). Precisely, we obtain that I(& Laplacetrf;((n)Z(n))|& Laplacetrf;(G(n))) = p(2)q(q + 1)/(4n)2 (1 + o(1)) when p = o(n).
We obtain bounds on the distribution of normalized gaps of eigenvalues of $N \times N$ GUE matrix in the bulk, that do not lose logarithmic factors of $N$ in the limit $N \to \infty$. As an application, we obtain fixed index universality results for the GUE minor process, which in turn are useful for establishing limiting results for random hives with GUE boundary data.