Given data {( x_i,y_i): i≤ n}, with x_i standard d-dimensional Gaussian feature vectors, and y_i∈ℝ response variables, we study the general problem of learning a model parametrized by θ∈ℝ^d, by minimizing a loss function that depends on θ via the one-dimensional projections θ^ T x_i. While previous work mostly dealt with convex losses, our approach assumes general (non-convex) losses hence covering classical, yet poorly understood examples such as the perceptron and non-convex robust regression. We use the Kac-Rice formula to control the asymptotics of the expected number of local minima of the empirical risk, under the proportional asymptotics n,d→∞, n/d>1. Specifically, we prove a finite dimensional variational formula for the exponential growth rate of the expected number of local minima. Further we provide sufficient conditions under which the exponential growth rate vanishes and all empirical risk minimizers have the same asymptotic properties (in fact, we expect the minimizer to be unique in these circumstances). We refer to this phenomenon as `rate trivialization.' If the population risk has a unique minimizer, our sufficient condition for rate trivialization is typically verified when the samples/parameters ratio α is larger than a suitable constant α_⋆. Previous general results of this type required n≥ Cd log d. We illustrate our results in the case of non-convex robust regression. Based on heuristic arguments and numerical simulations, we present a conjecture for the exact location of the trivialization phase transition α_tr.
We analyze the asymptotics of a block-Wishart random matrix ensemble of the type W_k = ( X^* ⊗ I_k) T( X⊗ I_k) for X∈ℂ^n× p with i.i.d. rows satisfying a suitable concentration-of-measure property, and T := Diag( T_i)_i∈[n] a block diagonal matrix with self-adjoint blocks T_i∈ℂ^k× k, under the proportional asymptotics n/p with k fixed. These matrices play a prominent role in the analysis of k-index models in high-dimensional statistics. By studying the matrix Stieltjes transform of this random matrix model and its inverse (K-transform), we derive variational formulas for two functionals of the asymptotic spectral density of W_k: the left (equivalently right) edge of its support, and its logarithmic potential.
We consider a general model for high-dimensional empirical risk minimization whereby the data 𝐱_i are d-dimensional isotropic Gaussian vectors, the model is parametrized by ∈ℝ^d× k, and the loss depends on the data via the projection ^𝖳𝐱_i. This setting covers as special cases classical statistics methods (e.g. multinomial regression and other generalized linear models), but also two-layer fully connected neural networks with k hidden neurons. We use the Kac-Rice formula from Gaussian process theory to derive a bound on the expected number of local minima of this empirical risk, under the proportional asymptotics in which n,d→∞, with n≍ d. Via Markov's inequality, this bound allows to determine the positions of these minimizers (with exponential deviation bounds) and hence derive sharp asymptotics on the estimation and prediction error. In this paper, we apply our characterization to convex losses, where high-dimensional asymptotics were not (in general) rigorously established for k≥ 2. We show that our approach is tight and allows to prove previously conjectured results. In addition, we characterize the spectrum of the Hessian at the minimizer. A companion paper applies our general result to non-convex examples.
We consider learning an unknown target function $f_*$ using kernel ridge regression (KRR) given i.i.d. data $(u_i,y_i)$, $i\leq n$, where $u_i \in U$ is a covariate vector and $y_i = f_* (u_i) +\varepsilon_i \in \mathbb{R}$. A recent string of work has empirically shown that the test error of KRR can be well approximated by a closed-form estimate derived from an `equivalent' sequence model that only depends on the spectrum of the kernel operator. However, a theoretical justification for this equivalence has so far relied either on restrictive assumptions -- such as subgaussian independent eigenfunctions -- , or asymptotic derivations for specific kernels in high dimensions. In this paper, we prove that this equivalence holds for a general class of problems satisfying some spectral and concentration properties on the kernel eigendecomposition. Specifically, we establish in this setting a non-asymptotic deterministic approximation for the test error of KRR -- with explicit non-asymptotic bounds -- that only depends on the eigenvalues and the target function alignment to the eigenvectors of the kernel. Our proofs rely on a careful derivation of deterministic equivalents for random matrix functionals in the dimension free regime pioneered by Cheng and Montanari (2022). We apply this setting to several classical examples and show an excellent agreement between theoretical predictions and numerical simulations. These results rely on having access to the eigendecomposition of the kernel operator. Alternatively, we prove that, under this same setting, the generalized cross-validation (GCV) estimator concentrates on the test error uniformly over a range of ridge regularization parameter that includes zero (the interpolating solution). As a consequence, the GCV estimator can be used to estimate from data the test error and optimal regularization parameter for KRR.
Maximum margin binary classification is one of the most fundamental algorithms in machine learning, yet the role of featurization maps and the high-dimensional asymptotics of the misclassification error for non-Gaussian features are still poorly understood. We consider settings in which we observe binary labels $y_i$ and either $d$-dimensional covariates ${\boldsymbol z}_i$ that are mapped to a $p$-dimension space via a randomized featurization map ${\boldsymbol \phi}:\mathbb{R}^d \to\mathbb{R}^p$, or $p$-dimensional features of non-Gaussian independent entries. In this context, we study two fundamental questions: $(i)$ At what overparametrization ratio $p/n$ do the data become linearly separable? $(ii)$ What is the generalization error of the max-margin classifier? Working in the high-dimensional regime in which the number of features $p$, the number of samples $n$ and the input dimension $d$ (in the nonlinear featurization setting) diverge, with ratios of order one, we prove a universality result establishing that the asymptotic behavior is completely determined by the expected covariance of feature vectors and by the covariance between features and labels. In particular, the overparametrization threshold and generalization error can be computed within a simpler Gaussian model. The main technical challenge lies in the fact that max-margin is not the maximizer (or minimizer) of an empirical average, but the maximizer of a minimum over the samples. We address this by representing the classifier as an average over support vectors. Crucially, we find that in high dimensions, the support vector count is proportional to the number of samples, which ultimately yields universality.
We study a general class of optimization problems with decision variable ∈ℝ^p × k and cost function which is the sum of n terms, each dependent on through the k-dimensional projection ^⊤x_i, where x_i, i ≤ n are i.i.d. random vectors. This setting is general enough to include examples of current interest in statistical physics, high-dimensional statistics, and statistical learning theory. We consider the proportional asymptotics n, p →∞, with n/p = Θ(1), and prove that, whenever there exists a minimizer satisfying a suitable generalization of a "delocalization" condition, the minimum value is universal. Namely, (for subgaussian x_i) it depends on the distribution of x_i only through its asymptotic mean and covariance. This delocalization condition is essentially necessary. Earlier universality results for such problems were limited to strongly convex loss functions. We derive applications of our theory to statistical learning and prove general universality results both for train and (under additional conditions) test error. In particular, we establish universality for vectors x_i generated by random 1-layer neural networks (random features models) and first-order Taylor approximations of 2-layer networks (neural tangent models). Finally, we establish that the delocalization property holds for a class of statistical learning problems under a condition that is easy to verify.
Consider(1) supervised learning from i.i.d. samples {(y(i) is an element of x(i))}(i <= n) where x(i) is an element of R-p are feature vectors and y(i) is an element of R are labels. We study empirical risk minimization over a class of functions that are parameterized by k = O(1) vectors theta(1), ... theta(k) is an element of k 2 R-p, and prove universality results both for the training and test error. Namely, under the proportional asymptotics n; p -> infinity, with n/p = Theta(1), we prove that the the training error depends on the random features distribution only through its mean and covariance structure. We also prove that the minimum test error over near-empirical risk minimizers enjoys similar universality properties. Furthermore, we give conditions guaranteeing universality of the test error of the empirical risk minimizer that can be checked in a "Gaussian equivalent" model where the features are replaced with Gaussian features of the same (asymptotic) mean and covariance. In particular, the asymptotics of the train and test error can be computed -to leading order- under a simpler model in which the feature vectors xi are replaced by Gaussian vectors g(i) with the same mean and covariance. Earlier universality results were limited to strongly convex learning procedures, or to feature vectors xi with independent entries. Our main results hold for non-convex procedures and feature vectors with dependent entries as long as they have asymptotically Gaussian projections along vectors theta whose l(2) norm is "well-spread out" among its coordinates. We give examples showing that generally, for universality to hold, one needs to constrain the minimization of the empirical risk to this set of well-spread out vectors. Furthermore, we give examples of conditions under which the minimum over this restricted space converges to the unristricted minimum. Our distributional assumptions are general enough to include feature vectors xi that are produced by randomized featurization maps. In particular we explicitly check the assumptions for certain random features models (computing the output of a one-layer neural network with random weights) and neural tangent models (first-order Taylor approximation of two-layer networks).
We consider distributions arising from a mixture of causal models, where each model is represented by a directed acyclic graph (DAG). We provide a graphical representation of such mixture distributions and prove that this representation encodes the conditional independence relations of the mixture distribution. We then consider the problem of structure learning based on samples from such distributions. Since the mixing variable is latent, we consider causal structure discovery algorithms such as FCI that can deal with latent variables. We show that such algorithms recover a "union" of the component DAGs and can identify variables whose conditional distribution across the component DAGs vary. We demonstrate our results on synthetic and real data showing that the inferred graph identifies nodes that vary between the different mixture components. As an immediate application, we demonstrate how retrieval of this causal information can be used to cluster samples according to each mixture component.
We consider the problem of learning a causal graph in the presence of measurement error. This setting is for example common in genomics, where gene expression is corrupted through the measurement process. We develop a provably consistent procedure for estimating the causal structure in a linear Gaussian structural equation model from corrupted observations on its nodes, under a variety of measurement error models. We provide an estimator based on the method-of-moments, which can be used in conjunction with constraint-based causal structure discovery algorithms. We prove asymptotic consistency of the procedure and also discuss finite-sample considerations. We demonstrate our method's performance through simulations and on real data, where we recover the underlying gene regulatory network from zero-inflated single-cell RNA-seq data.
We consider the task of learning a causal graph in the presence of latent confounders given i.i.d.~samples from the model. While current algorithms for causal structure discovery in the presence of latent confounders are constraint-based, we here propose a score-based approach. We prove that under assumptions weaker than faithfulness, any sparsest independence map (IMAP) of the distribution belongs to the Markov equivalence class of the true model. This motivates the \emph{Sparsest Poset} formulation - that posets can be mapped to minimal IMAPs of the true model such that the sparsest of these IMAPs is Markov equivalent to the true model. Motivated by this result, we propose a greedy algorithm over the space of posets for causal structure discovery in the presence of latent confounders and compare its performance to the current state-of-the-art algorithms FCI and FCI+ on synthetic data.
The ability to estimate task difficulty is critical for many real-world decisions such as setting appropriate goals for ourselves or appreciating others' accomplishments. Here we give a computational account of how humans judge the difficulty of a range of physical construction tasks (e.g., moving 10 loose blocks from their initial configuration to their target configuration, such as a vertical tower) by quantifying two key factors that influence construction difficulty: physical effort and physical risk. Physical effort captures the minimal work needed to transport all objects to their final positions, and is computed using a hybrid task-and-motion planner. Physical risk corresponds to stability of the structure, and is computed using noisy physics simulations to capture the costs for precision (e.g., attention, coordination, fine motor movements) required for success. We show that the full effort-risk model captures human estimates of difficulty and construction time better than either component alone.
In this paper, we present a new task that investigates how people interact with and make judgments about towers of blocks. In Experiment 1, participants in the lab solved a series of problems in which they had to re-configure three blocks from an initial to a final configuration. We recorded whether they used one hand or two hands to do so. In Experiment 2, we asked participants online to judge whether they think the person in the lab used one or two hands. The results revealed a close correspondence between participants' actions in the lab, and the mental simulations of participants online. To explain participants' actions and mental simulations, we develop a model that plans over a symbolic representation of the situation, executes the plan using a geometric solver, and checks the plan's feasibility by taking into account the physical constraints of the scene. Our model explains participants' actions and judgments to a high degree of quantitative accuracy.
Network Coding is a relatively new forwarding paradigm where intermediate nodes perform a store, code, and forward operation on incoming packets. Traditional forwarding approaches, which employed a store and forward operation, suffered from the limitations of the max-flow min-cut theorem wherein sources transmitting information over bottleneck links had to compete for access to these links. With Network Coding, multiple sources are now able to transmit packets over bottleneck links simultaneously, increasing network capacity. While the majority of the contemporary literature has focused on the performance of Network Coding from a capacity perspective, the aim of this research has taken a new direction focusing on two Quality of Service metrics, Packet Delivery Ratio (PDR) and latency, in conjunction with Network Coding protocols in Mobile Ad-Hoc Networks (MANETs). Initial simulations will be performed on static environments to determine a Quality of Service baseline comparison between Network Coding protocols and traditional ad-hoc routing protocols. Additional simulations will then be performed for mobile scenarios to determine how the Network Coding protocols will compare to that of the standard ad-hoc routing protocols in the presence of mobility.
Over the past years, we have witnessed an explosive growth in the use of the multimedia applications such as audio streaming with mobile and static devices. Audio streaming applications demand new approaches to audio transmissions to meet the growing traffic volume, quantity and quality of audio traffic, and users' needs. This paper studies network coding which is a promising paradigm that has the potential to improve the performance of networks for audio streaming applications in terms of packet delivery ratio (PDR), latency and jitter. This paper examines several network coding protocols for ad hoc wireless mesh networks and compares their performances on audio streaming applications with optimized broadcast protocols, e. g., Simplified Multicast Forwarding and Partial Dominant Pruning (PDP). The results show that the performance increases significantly with Random Linear Network Coding (RLNC) scheme.
In general, node protection in a communication network guarantees the traffic flow from source to destination. Traditional protection schemes in wired networks introduce either resource-hungry solutions such as the (1+1) protection scheme or a delay and interrupt to the network operation as in the (1:N) protection scheme. Node protection using network coding could solve the above issues. But the existing research efforts are mostly concentrated on wired networks and not much research has been conducted on wireless mesh networks (WMNs). Implementing the traditional (1+1) protection scheme in wireless networks increases the capital cost and resources cannot be fully utilized. On the other hand, the design and implementation of the (1:N) protection scheme in wireless networks for greater value of N (more than 3) is very difficult. Network coding is a promising technique for relaying traffic and node protection in a wireless network. In multihop wireless networks, very few research efforts are concentrated on node protection schemes using network coding and no reports have been published on measuring the Quality of Service (QoS) performance of node protection using network coding. In this paper, first we introduce a protection scheme for a single relay node failure for WMNs using network coding. We measure the QoS performance, e.g., packet delivery ratio (PDR), latency and jitter, for a single relay node failure with and without our protection scheme. Next, we extend the same protection scheme against two relay node failures and compare QoS performance with two relay failures and no failure. The results show that the QoS performance of protection against single relay node failure and two relay node failures are very close to the no failure scheme, while the network reliability increases due to the protection mechanism.
Physical-Layer Network Coding (PLNC) has recently emerged as a promising new communications paradigm that has the ability to greatly increase the capacity of wireless networks. Current research has also demonstrated that PLNC can decrease the eavesdropping region of an external entity. In this paper we show how it is possible to mitigate passive eavesdropping in the presence of co-located nodes in a PLNC environment.
lnternet, preter, distribuer et vendre des theses partout dans le monde, a des fins commerciales ou autres, sur support microforme, papier, electronique et/ou autres formats.
Thomas Kunz合作论文数Department of Systems and Computer Engineering, Carleton University3