Stochastic gradient descent (SGD) is a cornerstone algorithm for high-dimensional optimization, renowned for its empirical successes. Recent theoretical advances have provided a deep understanding of how SGD enables feature learning in high-dimensional nonlinear models, most notably the \emph{single-index model} with i.i.d. data. In this work, we study the sequential learning problem for single-index models, also known as generalized linear bandits or ridge bandits, where SGD is a simple and natural solution, yet its learning dynamics remain largely unexplored. We show that, similar to the optimal interactive learner, SGD undergoes a distinct "burn-in" phase before entering the "learning" phase in this setting. Moreover, with an appropriately chosen learning rate schedule, a single SGD procedure simultaneously achieves near-optimal (or best-known) sample complexity and regret guarantees across both phases, for a broad class of link functions. Our results demonstrate that SGD remains highly competitive for learning single-index models under adaptive data.
Towards understanding the fundamental limits of estimation from data of varied quality, we study the problem of estimating a mean parameter from heteroskedastic Gaussian observations where the variances are unknown and may vary arbitrarily across observations. While a simple linear estimator with known variances attains the smallest mean squared error, estimation without this knowledge is challenging due to the large number of nuisance parameters. We propose a simple and principled approach based on empirical Bayes: model the observations as if they were i.i.d. from a normal scale mixture and compute the profile maximum likelihood estimator (MLE) for the mean, treating the nonparametric mixing distribution as nuisance. Our result shows that this estimator achieves near-optimal error bounds across various heteroskedastic models in the literature. In particular, for the subset-of-signals problem where an unknown subset of observations has small variance, our estimator adaptively achieves the minimax rate for all signal sizes, including the sharp phase transition, without any tuning parameters. One of our key technical steps is a sharper metric entropy bound for normal scale mixtures, obtained via Chebyshev approximations on a transformed polynomial basis. This approach yields an improved polylogarithmic, rather than polynomial, dependence on the variance ratio, which could be of independent interest.
Existing auto-bidding algorithms in digital advertising often treat the value of an ad opportunity as the revenue obtained when an ad is shown and/or clicked, and bid accordingly. This can lead to wasteful spending because the true value is the marginal gain from paid exposure: even without winning a sponsored slot, an advertiser may still earn revenue via an organic search result (e.g., on Google or Amazon). Motivated by recent work, we model ad value as a treatment effect—the outcome difference between winning and losing the auction—and study online learning for bidding in second-price (Vickrey) auctions under this causal perspective. We develop algorithms that attain rate-optimal regret under several feedback models. A key ingredient exploits the information revealed by the second-price payment rule, which strictly improves regret relative to analogous learning problems in first-price auctions.
We study the contextual pricing problem, where in each round a seller observes a context, sets a price, and receives a binary purchase signal. We adopt a semi-parametric model in which the demand follows a linear parametric form composed with an unknown link function from a $\beta$-Hölder class. Prior work established regret rates of $\tilde{\mathcal{O}}(T^{2/3})$ for $\beta=1$ and $\tilde{\mathcal{O}}(T^{3/5})$ for $\beta=2$. Under a uni-modality condition, we propose a unified algorithm that combines the stationary subroutine of Wang & Chen (2025) with local polynomial regression, achieving the general rate $\tilde{\mathcal{O}}(T^{\frac{\beta+1}{2\beta+1}})$ for all $\beta \ge 1$. This recovers and strengthens existing results, while also addressing a gap in the prior analysis for $\beta=2$. Our analysis develops tighter semi-parametric confidence regions, removes derivative lower bound assumptions from earlier work, and offers a sharper exploration–exploitation trade-off. These insights not only extend theoretical guarantees to general $\beta$ but also improve practical performance by reducing the need for long forced-exploration phases.
We develop sharp bounds on the statistical distance between high-dimensional permutation mixtures and their i.i.d. counterparts. Our approach establishes a new geometric link between the spectrum of a complex channel overlap matrix and the information geometry of the channel, yielding tight dimension-independent bounds that close gaps left by previous work. Within this geometric framework, we also derive dimension-dependent bounds that uncover phase transitions in dimensionality for Gaussian and Poisson families. Applied to compound decision problems, this refined control of permutation mixtures enables sharper mean-field analyses of permutation-invariant decision rules, yielding strong non-asymptotic equivalence results between two notions of compound regret in Gaussian and Poisson models.
Let $(X_1,\ldots,X_n)$ be independent nonnegative random variables with $\mathbb{E} X_i\le1$, and write $S=\sum_iX_i$. For $δ>0$, we prove that \[ \mathbb{P}\left(S<\mathbb{E} S+δ\right)\ge b_{n,δ}, \] where $b_{n,δ}=δ(n/(n+δ))^n$ for $0<δ<1$ and $b_{n,δ}=(1-1/(n+δ))^n$ for $δ\ge1$. The bound is sharp for every $n$ and $δ\ge 1$. In particular, since $b_{n,δ} \ge e^{-1}$ for $δ\ge 1$, our result proves Feige's conjecture [Feige, 2004] in the affirmative for $δ\ge 1$. The proof is found by ChatGPT 5.6 Pro. It combines the exact Dirichlet calibration theorem of Vlassis and Thomas [Vlassis and Thomas, 2026], which resolves Gaffke's conjecture in statistics, with results in convex geometry including Grünbaum's centroid theorem [Grünbaum, 1960] and its generalization by Letwin and Yaskin [Letwin and Yaskin, 2024].
In this paper, we investigate personalized dynamic pricing and demand learning under high-dimensional models with growing effective dimensionality. Compared to strictly sparse models with a small and constant number of true predictors, allowing the number of effective features to grow can capture richer relationships through which an individual customer's demand can depend on the customer's features. Motivated by dynamic pricing applications where sellers typically have access to an expanding number of relevant customer features, we consider a seller that dynamically sets the price of a product for every individual customer. Customers arrive sequentially over time, and the seller observes a vector of features for each arriving customer. Demand is governed by a personalized model that depends on these customer features, and we measure the seller's performance by regret, defined as the revenue loss relative to a clairvoyant who knows the underlying personalized demand model. We establish theoretical regret bounds for this setting and complement them with numerical experiments that illustrate the role of growing effective dimensionality.
We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for 1-Lipschitz functions on the d-dimensional Euclidean ball. For time horizons n≥ d^10/3, we prove a lower bound of Ω(d^4/3√(n)), the first nontrivial bound that exceeds the d√(n) dependence of linear bandits, showing that stochastic bandit convex optimization is fundamentally harder than linear bandits. For d^2≤ n≤ d^10/3, we obtain a lower bound of Ω(√(d)n^3/4), matching the regret of the algorithm of Flaxman et al. (2005), establishing its optimality in this regime. The hard class of convex functions we construct takes the following form in dimension 2d: for an action a=(a^1,a^2)∈𝔹^2d, each function is the scaled soft maximum of a "tube", r^-1W^⋆ a^1-r/8εa^2 (hyperparameterized by ε,r), and a squared distance function, 1/2a^1-u^⋆^2-1/2u^⋆^2. Here u^⋆∈ℝ^d is the unknown target determining the minimizer, while W^⋆∈ℝ^d× d hides the region in which the quadratic curvature is observable. Indeed, observations reveal substantial information about u^⋆ only when the learner acts near the hidden tube a^2≈8ε/rW^⋆ a^1; away from it, the tube branch masks the quadratic branch. Thus the learner must pay to uncover the geometry encoded by W^⋆ before it can effectively exploit the curvature that identifies u^⋆. Formalizing this tradeoff yields a sample complexity lower bound of Ω(d^5/2/ε^2∧d^2/ε^4) for finding an ε-optimal action, and ultimately the Ω(d^4/3√(n)∧√(d)n^3/4) regret lower bound. The proof was developed by GPT-5.5 Pro and GPT-5.6 Sol Pro under the authors' guidance.
We theoretically justify the recent empirical finding of [Teh et al., 2025] that a transformer pretrained on synthetically generated data achieves strong performance on empirical Bayes (EB) problems. We take an indirect approach to this question: rather than analyzing the model architecture or training dynamics, we ask why a pretrained Bayes estimator, trained under a prespecified training distribution, can adapt to arbitrary test distributions. Focusing on Poisson EB problems, we identify the existence of universal priors such that training under these priors yields a near-optimal regret bound of O(1/n) uniformly over all test distributions. Our analysis leverages the classical phenomenon of posterior contraction in Bayesian statistics, showing that the pretrained transformer adapts to unknown test distributions precisely through posterior contraction. This perspective also explains the phenomenon of length generalization, in which the test sequence length exceeds the training length, as the model performs Bayesian inference using a generalized posterior.
We prove bounds on statistical distances between high-dimensional exchangeable mixture distributions (which we call permutation mixtures) and their i.i.d. counterparts. Our results are based on a novel method for controlling χ^2 divergences between exchangeable mixtures, which is tighter than the existing methods of moments or cumulants. At a technical level, a key innovation in our proofs is a new Maclaurin-type inequality for elementary symmetric polynomials of variables that sum to zero and an upper bound on permanents of doubly-stochastic positive semidefinite matrices. We obtain as a corollary a new de Finetti-style theorem (in the language of Diaconis and Freedman, 1987), as well as several new statistical results, including a differential privacy guarantee for the “shuffled privacy model” with Gaussian noise and improved generic consistency guarantees for empirical Bayes procedures in compound decision problems.
We prove new inequalities for elementary symmetric polynomials (ESPs) for vectors that sum to zero, and for square matrices with zero row and column sums. We apply these results to obtain a unified upper bound on the mean-field approximation guarantee for permutation mixtures, as well as a sharp $χ^2$ version of the de Finetti theorem for finite sequences over a small alphabet. The main proof ideas were developed by the GPT-5.5 Pro model.
When faced with a small sample from a large universe of possible outcomes, scientists often turn to the venerable Good–Turing estimator. Despite its pedigree, however, this estimator comes with considerable drawbacks, such as the need to hand-tune smoothing parameters and the lack of a precise optimality guarantee. We introduce a parameter-free estimator that bests Good–Turing in both theory and practice. Our method marries two classic ideas, namely Robbins's empirical Bayes and Kiefer–Wolfowitz non-parametric maximum likelihood estimation (NPMLE), to learn an implicit prior from data and then convert it into probability estimates. We prove that the resulting estimator attains the optimal instance-wise risk up to logarithmic factors in the competitive framework of Orlitsky and Suresh, and that the Good–Turing estimator is strictly suboptimal in the same framework. Our simulations on synthetic data and experiments with English corpora and U.S. Census data show that our estimator consistently outperforms both the Good–Turing estimator and explicit Bayes procedures.
We study the evolution of information in interactive decision making through the lens of a stochastic multi-armed bandit problem. Focusing on a fundamental example where a unique optimal arm outperforms the rest by a fixed margin, we characterize the optimal success probability and mutual information over time. Our findings reveal distinct growth phases in mutual information---initially linear, transitioning to quadratic, and finally returning to linear---highlighting curious behavioral differences between interactive and non-interactive environments. In particular, we show that optimal success probability and mutual information can be decoupled, where achieving optimal learning does not necessarily require maximizing information gain. These findings shed new light on the intricate interplay between information and learning in interactive decision making.
Feature alignment methods are used in many scientific disciplines for data pooling, annotation, and comparison. As an instance of a permutation learning problem, feature alignment presents significant statistical and computational challenges. In this work, we propose the covariance alignment model to study and compare various alignment methods and establish a minimax lower bound for covariance alignment that has a non-standard dimension scaling because of the presence of a nuisance parameter. This lower bound is in fact minimax optimal and is achieved by a natural quasi MLE. However, this estimator involves a search over all permutations which is computationally infeasible even when the problem has moderate size. To overcome this limitation, we show that the celebrated Gromov-Wasserstein algorithm from optimal transport which is more amenable to fast implementation even on large-scale problems is also minimax optimal. These results give the first statistical justification for the deployment of the Gromov-Wasserstein algorithm in practice.
We study regret minimization in repeated first-price auctions (FPAs), where a bidder observes only the realized outcome after each auction – win or loss. This setup reflects practical scenarios in online display advertising where the actual value of an impression depends on the difference between two potential outcomes, such as clicks or conversion rates, when the auction is won versus lost. We analyze three outcome models: (1) adversarial outcomes without features, (2) linear potential outcomes with features, and (3) linear treatment effects in features. For each setting, we propose algorithms that jointly estimate private values and optimize bidding strategies, achieving near-optimal regret bounds. Notably, our framework enjoys a unique feature that the treatments are also actively chosen, and hence eliminates the need for the overlap condition commonly required in causal inference.
Combinatorial bandits extend the classical bandit framework to settings where the learner selects multiple arms in each round, motivated by applications such as online recommendation and assortment optimization. While extensions of upper confidence bound (UCB) algorithms arise naturally in this context, adapting arm elimination methods has proved more challenging. We introduce a novel elimination scheme that partitions arms into three categories (confirmed, active, and eliminated), and incorporates explicit exploration to update these sets. We demonstrate the efficacy of our algorithm in two settings: the combinatorial multi-armed bandit with general graph feedback, and the combinatorial linear contextual bandit. In both cases, our approach achieves near-optimal regret, whereas UCB-based methods can provably fail due to insufficient explicit exploration. Matching lower bounds are also provided.
The "sample amplification" problem formalizes the following question: Given n i.i.d. samples drawn from an unknown distribution P, when is it possible to produce a larger set of n + m samples which cannot be distinguished from n + m i.i.d. samples drawn from P? In this work, we provide a firm statistical foundation for this problem by deriving generally applicable amplification procedures, lower bound techniques and connections to existing statistical notions. Our techniques apply to a large class of distributions including the exponential family, and establish a rigorous connection between sample amplification and distribution learning.
In “Optimal No-Regret Learning in Repeated First-Price Auctions,” Y. Han, W. Tsachy, and Z. Zhou study online learning in repeated first-price auctions where a bidder, only observing the winning bid at the end of each auction, learns to adaptively bid to maximize her cumulative payoff. To achieve this goal, the bidder faces censored feedback: If she wins the bid, then she is not able to observe the highest bid of the other bidders, which we assume is i.i.d. drawn from an unknown distribution. In this paper, they develop the first learning algorithm that achieves a near-optimal regret bound, by exploiting two structural properties of first-price auctions, that is, the specific feedback structure and payoff function.
We develop a unifying framework for information-theoretic lower bound in statistical estimation and interactive decision making. Classical lower bound techniques---such as Fano's inequality, Le Cam's method, and Assouad's lemma---are central to the study of minimax risk in statistical estimation, yet are insufficient to provide tight lower bounds for \emph{interactive decision making} algorithms that collect data interactively (e.g., algorithms for bandits and reinforcement learning). Recent work of Foster et al. provides minimax lower bounds for interactive decision making using seemingly different analysis techniques from the classical methods. These results---which are proven using a complexity measure known as the \emph{Decision-Estimation Coefficient} (DEC)---capture difficulties unique to interactive learning, yet do not recover the tightest known lower bounds for passive estimation. We propose a unified view of these distinct methodologies through a new lower bound approach called \emph{interactive Fano method}. As an application, we introduce a novel complexity measure, the \emph{Decision Dimension}, which facilitates the new lower bounds for interactive decision making that extend the DEC methodology by incorporating the complexity of estimation. Using the Decision Dimension, we (i) provide a unified characterization of learnability for \emph{any} structured bandit problem, (ii) close the remaining gap between the upper and lower bounds in Foster et al. (up to polynomial factors) for any interactive decision making problem in which the underlying model class is convex.