
Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly understood, and many theoretical works have focused on the easier-to-analyze linearized attention. In this work, we address this gap through a principled study of the single-location regression task, where the output depends on a linear transformation of a single input token at a random location. Building on ideas from statistical physics, we develop an analysis of attention-based predictors in the high-dimensional limit, where generalization performance is captured by a small set of order parameters. At the population level, we show that softmax achieves the Bayes risk, whereas linear attention fundamentally falls short. We then examine other activation functions to identify which properties are necessary for optimal performance. Finally, we analyze the finite-sample regime: we provide an asymptotic characterization of the test error and show that, while softmax is no longer Bayes-optimal, it consistently outperforms linear attention. We discuss the connection with optimization by gradient-based algorithms.
In this work, we revisit the Generalised Navier Boundary Condition (GNBC) introduced by Qian et al. in the sharp interface volume-of-fluid context. We replace the singular uncompensated Young stress by a smooth function with a characteristic width $\varepsilon \gt 0$ that is understood as a physical parameter of the model. Therefore, we call the model the 'contact region GNBC' (CR-GNBC). We show that the model is consistent with the fundamental kinematics of the contact angle transport described by Fricke, K & ouml;hne and Bothe. We implement the model in the geometrical volume-of-fluid solver Basilisk using a 'free angle' approach. This means that the dynamic contact angle is not prescribed, but reconstructed from the interface geometry and subsequently applied as an input parameter to compute the uncompensated Young stress. We couple this approach to the two-phase Navier-Stokes solver and study the withdrawing tape problem with a receding contact line. It is shown that the model allows for grid-independent solutions and leads to a full regularisation of the singularity at the moving contact line, which is in accordance with the thin film equation subject to this boundary condition. In particular, it is shown that the curvature at the moving contact line is finite and mesh converging. As predicted by the fundamental kinematics, the parallel shear stress component vanishes at the moving contact line for quasi-stationary states (i.e. for $\dot heta _d=0$ ), and the dynamic contact angle is determined by a balance between the uncompensated Young stress and an effective contact line friction. Furthermore, a nonlinear generalisation of the model is proposed, which aims at reproducing the molecular kinetic theory of Blake and Haynes for quasi-stationary states.
We prove that for almost all symmetric spaces X and for any sequence of compact locally symmetric spaces Y_n which is uniformly discrete, has a uniform spectral gap, and converges in the sense of Benjamini–Schramm to X, the joint eigenfunctions of all invariant differential operators on Y_n delocalize on average when their spectral parameters are taken to lie in a fixed spectral window.
This paper addresses the nonparametric estimation of the drift function over a compact domain for a time-homogeneous diffusion process, based on high-frequency discrete observations from N independent trajectories. We propose a neural network-based estimator and derive a non-asymptotic convergence rate, decomposed into a training error, an approximation error, and a diffusion-related term scaling as log N/N. For compositional drift functions, we establish an explicit rate. In the numerical experiments, we consider a drift function with local fluctuations generated by a double-layer compositional structure featuring local oscillations, and show that the empirical convergence rate becomes independent of the input dimension d. Compared to the B-spline method, the neural network estimator achieves better convergence rates and more effectively captures local features, particularly in higher-dimensional settings.
Our main result is to show that, if the p-th Bernstein polynomial of the (a, b)-module generated by a germ of a holomorphic volume form omega is an element of ohm n+1 0 in the (convergent) Brieskorn (a, b)-module associated to f, has a root-alpha - N, there exists a pole of order at least p for the meromorphic extension of an analytic functional associated to omega at some point in-alpha - N, under the hypothesis that f has an isolated singularity at the origin relative to the corresponding eigenvalue exp(2i pi alpha) of the monodromy. This implies the existence of at least p roots in-alpha - N (counting multiplicities) for the usual reduced Bernstein polynomial of the germ of f at 0. We also obtain in the case of an isolated singularity for f that the largest root-alpha-m inside {-alpha -N} of the reduced Bernstein polynomial of f produces a pole at the point lambda = -alpha - m for the meromorphic extension of the distribution f 2 lambda f & strns; -h for some h is an element of N.