
Mathematical methods are developed to characterize the asymptotics ofrecurrent neural networks (RNN) as the number of hidden units, data samplesin the sequence, hidden state updates, and training steps simultaneously growto infinity. In the case of an RNN with a simplified weight matrix, we provethe convergence of the RNN to the solution of an infinite-dimensional ODEcoupled with the fixed point of a random algebraic equation. The analysisrequires addressing several challenges which are unique to RNNs. In typicalmean-field applications (e.g., feedforward neural networks), discrete updatesare of magnitude & Oscr;(1/N )and the number of updates is & Oscr;(N). Therefore,the system can be represented as an Euler approximation of an appropriateODE/PDE, which it will converge to asN ->infinity. However, the RNN hiddenlayer updates are & Oscr;(1). Therefore, RNNs cannot be represented as a dis-cretization of an ODE/PDE and standard mean-field techniques cannot beapplied. Instead, we develop a fixed point analysis for the evolution of theRNN memory states, with convergence estimates in terms of the number ofupdate steps and the number of hidden units. The RNN hidden layer is stud-ied as a function in a Sobolev space, whose evolution is governed by the datasequence (a Markov chain), the parameter updates, and its dependence on theRNN hidden layer at the previous time step. Due to the strong correlation be-tween updates, a Poisson equation must be used to bound the fluctuations ofthe RNN around its limit equation. These mathematical methods give rise tothe neural tangent kernel (NTK) limits for RNNs trained on data sequencesas the number of data samples and size of the neural network grow to infinity.
The constrained minimization-respectively maximization-of dissimilarity-modeling divergences (i.e., generally nonsymmetric, directed distances) and of related generalized entropies is a fundamental task in many areas of quantitative sciences. On R-K of arbitrary dimension K, we derive a method which tackles such kind of constrained optimization problems-and beyond-by limits of sequences of appropriately constructed, dimension-free, typically comfortably simulable random vectors; almost no assumptions (like convexity) on the set of constraints are needed. This very largely extends our recent results on f-divergences which can be connected to light-tailed probability distributions in a certain manner (cf. IEEE Trans. Inform. Theory (2023) 69 3062-3120). For instance, in the current paper we cover constrained optimizations of arbitrary f-divergences, Bregman distances, scaled Bregman distances and weighted & ell;(r)-distances.
We introduce a technique to establish limit profiles for Markov chains by leveraging tools from stochastic analysis-in particular, approximations via SDEs. We demonstrate our arguments by analysing the Bernoulli-Laplace urn model: initially, one urn contains k red balls and a second n - k blue balls; in each step, a pair of balls is chosen uniform and their locations are switched. Cutoff is known to occur at 1/2 n log min{k, root n} with window order n whenever 1 << k <= 1/2 n. We refine this by determining the limit profile: a function Phi such that d(TV) (1/2 n logmin{k, root n} + theta n) -> Phi(theta) as n -> infinity for all theta is an element of R. Its birth-death representation approximates an Ornstein-Uhlenbeck diffusion, once appropriately rescaled, started from the bulk (according to its equilibrium distribution), when k >> root n. Precise concentration results as the chain enters the bulk are needed. We also establish the limit profile when 1 << k less than or similar to root n, via a different method.
We study long time behavior of shear-thinning fluid flows ind >= 3di-mensions, driven by additive stochastic forcing of trace class, with power-lawindices ranging from 1 to2dd+2. We particularly focus on Leray-Hopf solu-tions, that is, on analytically weak solutions satisfying energy inequality.Introducing a new kind of energy related functional into the technique ofconvex integration enables the construction of infinitely many such solutionsthat are probabilistically strong for a certain initial value. Furthermore, weprovide global in time estimates which lead to the existence of infinitely manystationary and even ergodic Leray-Hopf solutions.These results represent the first construction of Leray-Hopf solutions inthe framework of stochastic shear-thinning fluids within this range of power-law indices
We investigate the asymptotic distribution of the largest off-diagonal en-try of the sample covariance matrix under the autoregressive structure drivenby a parameterr=rn is an element of(0,1). Our analysis reveals a phase transition at theorder of root logp/root nin the asymptotic distribution of the largest entries whenris bounded away from 1. Additionally, asrconverges to 1 with logarith-mic rate, we show that there is a subtle clustering effect among the entries,leading to an asymptotic distribution characterized by a Gumbel law with anextremal index strictly between 0 and 1. This unexpected behavior is coun-terintuitive and represents a significant departure from the standard Gumbellaw observed in (Ann. Statist.53(2025) 907-928) for sample correlation ma-trices: the seemingly simpler sample covariance matrices can exhibit morecomplex behavior than the sample correlation matrices. Our proof methodol-ogy relies on high-dimensional Gaussian approximation techniques and em-ploys a blocking argument to create independence.
This paper addresses a generalized filtering framework in which the signal process X and observation process Y are both driven by correlated Brownian motions, and the coefficients of their governing stochastic differential equations depend jointly on (X, Y), with the exception of the diffusion coefficient of the observation process, which does not depend upon the signal. Unlike many prior works, the observation equation may have a degenerate (noninvertible or even zero) diffusion coefficient. In this framework, we derive filtering equations and prove of equivalence between the uniqueness of the nonlinear Kushner-Stratonovich equation and the linear Zakai equation. Finally we give a novel proof of uniqueness for the Zakai equation using a backward stochastic partial differential equation (BSPDE), overcoming the limitations of classical duality arguments. This approach successfully handles the randomness and anticipation introduced by the observation-dependent coefficients, which are not tractable under traditional deterministic PDE methods.
In this paper we will introduce a class of forward-backward stochastic differential equations on tensor fields of Riemannian manifolds, which are related to semilinear parabolic partial differential equations on tensor fields. Moreover, we will use these forward-backward stochastic differential equations to give a stochastic representation of incompressible Navier-Stokes equations on Riemannian manifolds, where some extra conditions used in (Potential Anal. 48 (2018) 181-206) are not required.
The concept of downward conditional monotonicity for the Markov-modulated Poisson process (MMPP) is introduced and used to derive the opti-mal stochastic domination of a standard Poisson point process. The maximumarrival rate for the Poisson process which allows this domination to exist isshown to be related to an eigenvalue extracted from the generator matrix ofthe quasi-birth-death (QBD) formulation of the MMPP. This allows deriva-tion of survival and extinction regimes for a large family of contact processeswhose infection and recovery rates vary over time according to an underlyingrandom environment with a finite number of states. Direct comparison withstandard contact processes which dominate from above and below accom-plishes this
Motivated by the central role of social networks in the diffusion of information, the study of network valued data where nodes and/or edges have attributes, which modulate the dynamics of both network evolution, and information flow on the network itself, has witnessed significant research interest across multiple disciplines. A key ingredient of this general area comprises probabilistic network models that incorporate (a) heterogeneity in edge creation across different attribute groups; (b) temporal network evolution and (c) popularity bias. Such models are then used to understand a host of domain specific questions, including bias in network sampling, PageRank and degree centrality scores and their impact in network ranking and recommendation algorithms. Despite significant interest, for these network models, the main network functional amenable to analysis has so far been degree distribution asymptotics. In this paper, we analyze dynamic random network models where younger vertices connect to older ones with probabilities proportional to their degrees as well as a propensity kernel governed by their attribute types. Using stochastic approximation techniques we show that, in the large network limit, such networks converge in the local weak sense to limiting infinite random trees with an explicit description in terms of randomly stopped multi-type branching processes. This allows for the derivation of asymptotics for a wide class of network functionals implying, for example, that while degree distribution tail exponents depend on the attribute type (already derived by (Elec-tron. J. Probab. 18 (2013) 8)), PageRank centrality scores have the same tail exponent across attributes. The limit results also give explicit formulae for the performance of various network sampling mechanisms. One surprising consequence is the efficacy of PageRank and walk based network sampling schemes for directed networks in the setting of rare minorities.
We establish global two-sided heat kernel estimates (for full time and space) of the Schrodinger operator -1/2 Delta + V on & Ropf;(d), where the potential V (x) is locally bounded and behaves like c|x |(-alpha) near infinity with alpha is an element of (0, 2) and c > 0, or with alpha > 0 and c < 0. Our results improve all known results in the literature, and it seems that the current paper is the first one where two-sided matching heat kernel bounds for the long range potentials are established. The results of the paper mostly rely on probabilistic approaches.
A large class of statistics can be formulated as smooth functions of sample means of random vectors. In this paper, we propose a general partial Cram & eacute;r's condition (GPCC) and apply it to establish the validity of the Edgeworth expansion for the distribution function of these functions of sample means. Additionally, we apply the proposed theorems to several specific statistics. In particular, by verifying the GPCC, we demonstrate for the first time the validity of the formal Edgeworth expansion of Pearson's correlation coefficient between random variables with absolutely continuous and discrete components. Furthermore, we conduct a series of simulation studies that show the Edgeworth expansion has higher accuracy.
This work focuses on the mean field stochastic partial differential equations with nonlinear kernels. We first prove the existence and uniqueness of strong and weak solutions for mean field stochastic partial differential equations in the variational framework, then establish the convergence (in certain Wasserstein metric) of the empirical laws of interacting systems to the law of solutions of mean field equations, as the number of particles tends to infinity. The main challenge lies in addressing the inherent interplay between the high nonlinearity of operators and the non-local effect of coefficients that depend on the measure. In particular, we do not need to assume any exponential moment control condition of solutions, which extends the range of the applicability of our results. As applications, we first study a class of finite-dimensional interacting particle systems with polynomial kernels, which are commonly encountered in fields such as the data science and the machine learning. Subsequently, we present several illustrative examples of infinite-dimensional interacting systems with nonlinear kernels, such as the stochastic climate models, the stochastic Allen-Cahn equations, and the stochastic Burgers type equations.
We consider percolation of the vacant set of random interlacements at intensity u in dimensions three and higher, and derive lower bounds on the truncated two-point function for all values of u>0. These bounds are sharp up to principal exponential order for all u in dimension three and all u not equal u(& lowast;) in higher dimensions, where u & lowast; refers to the critical parameter of the model, and they match the upper bounds derived in the article (Goswami, Rodriguez, Shulzhenko (2025)). In dimension three, our results further imply that the truncated two-point function grows at large distances x at a rate that depends on x only through its Euclidean norm, which offers a glimpse of the expected (Euclidean) invariance of the scaling limit at criticality. The decay rate is atypical, it incurs a logarithmic correction and comes with an explicit pre-factor that converges to 0 as the parameter u approaches the critical point u(& lowast;) from either side. A particular challenge stems from the combined effects of lack of monotonicity due to the truncation in the super-critical phase, and the precise (rotationally invariant) controls we seek, that measure the effects of a certain "harmonic humpback" function. Among others, their derivation relies on rather fine estimates for hitting probabilities of the random walk in arbitrary direction e, which witness this invariance at the discrete level, and preclude straightforward applications of projection arguments.
For a > 0 and b > 0, let G(a,b) be the subgraph of Z(2) induced by the vertices between the first coordinate axis and the graph of the function f = f(a,b)(u) = a log(1 + u) + b log(1 + log(1 + u)), u >= 0. It is known that for a > 0, the critical value for Bernoulli percolation on G(f) = G(a,b) is strictly between 1/2 and 1, and that if b > 2a then the percolation phase transition is discontinuous. We study first-passage percolation (FPP) on G(a,b) with i.i.d. edge-weights (tau(e)) satisfying p = IP(tau(e )= 0) is an element of [1/2, 1) and the "gap condition" IP(tau(e) <= delta) = p for some delta > 0. We find the rate of growth of the expected passage time in G(f) from the origin to the line x = n, and show that, while when p = 1/2 it is of order n/(a log n), when p > 1/2 it can be of order (a) n(c1 )/(logn)(c2), (b) (log n)(c3), (c) log log n, or (d) constant, depending on the relationship between a, b, and p. For more general functions f, we prove a central limit theorem for the passage time and show that its variance grows at the same rate as the mean. As a consequence of our methods, we improve the percolation transition result by showing that the phase transition on G(a,b) is discontinuous if and only if b > a, and improve "sponge crossing dimensions" asymptotics from the 1980s on subcritical percolation crossing probabilities for tall thin rectangles.
We show the existence of a phase transition between a localisation and a nonlocalisation regime for a branching random walk with a catalyst at the origin. More precisely, we consider a continuous-time branching random walk that jumps at rate one, with simple random walk jumps on & Zopf;(d), and that branches (with binary branching) at rate lambda > 0 everywhere, except at the origin, where it branches at rate lambda(0 )> lambda. We show that, if lambda(0) is large enough, then the occupation measure of the branching random walk localises (i.e., when normalised by the total number of particles, it converges almost surely without spatial renormalisation), whereas, if lambda(0) is close enough to lambda, then the occupation measure delocalises, in the sense that the proportion of particles in any finite given set converges almost surely to zero. The case lambda = 0 (when branching only occurs at the origin) has been extensively studied in the literature and a transition between localisation and nonlocalisation was also exhibited in this case. Interestingly, the transition that we observe, conjecture, and partially prove in this paper occurs at the same threshold as in the case lambda = 0. One of the strengths of our result is that, in the localisation regime, we are able to prove convergence of the occupation measure, while existing results in the case lambda = 0 give convergence of moments instead.
We introduce a novel topology, called Kernel Mean Embedding Topology, for stochastic kernels, in a weak and strong form. This topology, defined on the spaces of Bochner integrable functions from a signal space to a space of probability measures endowed with a Hilbert space structure, allows for a versatile formulation. This construction allows one to obtain both a strong and weak formulation. (i) For its weak formulation, we highlight the utility on relaxed policy spaces, and investigate connections with the Young narrow topology and Borkar (or w^*)-topology, and establish equivalence properties. We report that, while both the w^*-topology and kernel mean embedding topology are relatively compact, they are not closed. Conversely, while the Young narrow topology is closed, it lacks relative compactness. (ii) We show that the strong form provides an appropriate formulation for placing topologies on spaces of models characterized by stochastic kernels with explicit robustness and learning theoretic implications on optimal stochastic control under discounted or average cost criteria. (iii) We thus show that this topology possesses several properties making it ideal to study optimality and approximations (under the weak formulation) and robustness (under the strong formulation) for many applications.
We study the uniform-in-time weak propagation of chaos for the consensus-based optimization (CBO) method on a bounded searching domain. We apply the methodology for studying long-time behaviors of interacting particle systems developed in the work of Delarue and Tse (ArXiv:2104.14973). Our work shows that the weak error has order O(N^-1) uniformly in time, where N denotes the number of particles. The main strategy behind the proofs are the decomposition of the weak errors using the linearized Fokker-Planck equations and the exponential decay of their Sobolev norms. Consequently, our result leads to the joint convergence of the empirical distribution of the CBO particle system to the Dirac-delta distribution at the global minimizer in population size and running time in Wasserstein-type metrics.
In this paper, it is shown that with large probability, the spectral radius of a large non-Hermitian random matrix with a general variance profile does not exceed the square root of the spectral radius of the variance profile matrix. A minimal moment assumption is considered and sparse variance profiles are covered. Following an approach developed recently by Bordenave, Chafaï and García-Zelada, the key theorem states the asymptotic equivalence between the reverse characteristic polynomial of the random matrix at hand and a random analytic function which depends on the variance profile matrix. The result is applied to the case of a non-Hermitian random matrix with a variance profile given by a piecewise constant or a continuous non-negative function, the inhomogeneous (centered) directed Erdős-Rényi model, and more.
The Robbins-Siegmund theorem is one of the most important results in stochastic optimization, where it is widely used to prove the convergence of stochastic algorithms. We provide a quantitative version of the theorem, establishing a bound on how far one needs to look in order to locate a region of metastability in the sense of Tao. Our proof involves a metastable analogue of Doob's theorem for L_1-supermartingales along with a series of technical lemmas that make precise how quantitative information propagates through sums and products of stochastic processes. In this way, our paper establishes a general methodology for finding metastable bounds for stochastic processes that can be reduced to supermartingales, and therefore for obtaining quantitative convergence information across a broad class of stochastic algorithms whose convergence proof relies on some variation of the Robbins-Siegmund theorem. We conclude by discussing how our general quantitative result might be used in practice.
In this paper, we establish the ergodic and mixing properties of stochastic 2D Navier-Stokes equations driven by a highly degenerate multiplicative Gaussian noise. The noise can appear in as few as four directions, and its intensity depends on the solution. The case of additive Gaussian noise was previously treated by Hairer and Mattingly (Ann. of Math. (2) 164 (2006) 993-1032). To derive the ergodic and mixing properties in the present setting, we employ Malliavin calculus to establish the asymptotically strong Feller property. The primary challenge lies in proving the "invertibility" of the Malliavin matrix, which differs fundamentally from the additive case.