Generative Pre-trained Transformer (GPT) architectures are the most popular design for language modeling. Energy-based modeling is a different paradigm that views inference as a dynamical process operating on an energy landscape. We propose a minimal modification of the GPT setting to unify it with the EBM framework. The inference step of our model, which we call eNeRgy-GPT (NRGPT), is conceptualized as an exploration of the tokens on the energy landscape. We prove, and verify empirically, that under certain circumstances this exploration becomes gradient descent, although they don't necessarily lead to the best performing models. We demonstrate that our model performs well for simple language (Shakespeare dataset), algebraic ListOPS tasks, and richer settings such as OpenWebText language modeling. We also observe that our models may be more resistant to overfitting, doing so only during very long training.
Both the backpropagation algorithm in machine learning and the maximum principle in optimal control theory are posed as a two-point boundary problem, resulting in a "forward-backward" lock. We derive a reformulation of the maximum principle in optimal control theory as a hyperbolic initial value problem by introducing an additional "optimization time" dimension. We introduce counter-propagating wave variables with finite propagation speed and recast the optimization problem in terms of scattering relationships between them. This relaxation of the original problem can be interpreted as a physical system that equilibrates and changes its physical properties in order to minimize reflections. We discretize this continuum theory to derive a family of fully unlocked algorithms suitable for training neural networks. Different parameter dynamics, including gradient descent, can be derived by demanding dissipation and minimization of reflections at parameter ports. These results also imply that any physical substrate that supports the scattering and dissipation of waves can be interpreted as solving an optimization problem.
The recent posting arXiv:2605.02621 [14], commenting on the article rspa.2025.0413 [7], argues that the proof of Lemma 3.1 in [7] is missing the spatial derivative of the density, which would lead to a Bohm-like quantum potential. This technical note shows why the propagated density is independent of space in the Feynman propagator construction of Lemma 3.1. This is done by extending the proof of Lemma 3.1 explicitly with Bohm-like quantum potential terms along the stationary action paths, and then showing that these terms are exactly zero. In [7], this property can also be verified directly on most examples (double slit, Aharonov-Bohm, potential well, harmonic oscillator, tunneling, EPR, QED), as well as in the derivations of the Pauli, Dirac, and Maxwell equations. For more general nonlinear actions, a time rescaling may be required to guarantee this space independence along stationary paths. In the hydrogen atom example, this time rescaling can be computed in closed form. In contrast to the general wave of the Madelung solution [9] Lemma 3.1 of [7] is defined first for a propagator, and a general wave is then constructed in a second step. Recall that a propagator is a specific quantum wave, which is initialized at t=0 with a Dirac impulse at a given initial position or momentum. In turn, a general wave is constructed in a second step by superposing a distribution of initial conditions using the propagator. This key difference is why the Bohm-like quantum potential terms disappear in the construction [7] (specifically, in the first step) while the Bohm potential in the Madelung analysis does not. This fundamental difference is also consistent with the fact that the wave construction in [7] extends naturally to relativistic contexts, while Bohmian non-locality notoriously prevents such extensions. Keywords - Response to arXiv:2605.02621, in relation to rspa.2025.0413
Many energy-based control strategies for mechanical systems require the choice of a Coriolis factorization satisfying a skew-symmetry property. This paper (a) explores if and when a control designer has flexibility in this choice, (b) develops a canonical choice related to the Christoffel symbols, and (c) describes how to efficiently perform control computations with it for constrained mechanical systems. We link the choice of a Coriolis factorization to the notion of an affine connection on the configuration manifold and show how properties of the connection relate with the associated factorization. In particular, the factorization based on the Christoffel symbols is linked with a torsion-free property that can limit the twisting of system trajectories during passivity-based control. We then develop a way to induce Coriolis factorizations for constrained mechanisms from unconstrained ones, which provides a pathway to use the theory for efficient control computations with high-dimensional systems such as humanoids and quadruped robots with open- and closed-chain mechanisms. A collection of algorithms is provided (and made available open source) to support the recursive computation of passivity-based control laws, adaptation laws, and regressor matrices in future applications.
We show that the Schr & ouml;dinger equation can be solved exactly based only on classical least action. Fundamental postulates of quantum mechanics can in turn be derived directly from this construction. The results extend to the relativistic Klein-Gordon, Pauli, and Dirac equations, and suggest a smooth transition between physics across scales. Most quantum mechanics problems have classical versions which involve multiple least action solutions. The associated classical multipaths stem either from the initial position or momentum distribution, or from branch points, generated, e.g. by a multiply connected manifold (double slit experiment), by spatial inequality constraints (particle in a box), or by a singularity (Coulomb potential). We show that the exact Schr & ouml;dinger wave function psi can be constructed by combining this classical multi-valued action phi with the classical density rho, computed analytically from phi along each extremal action path. The construction is general and does not involve any semi-classical approximation. Quantum wave collapse at measurement can be derived from the classical density change. Entanglement corresponds to a sum of classical particle actions mapping to a tensor product of spinors. The results also provide a simpler computational alternative to Feynman path integrals, as they use only a minimal subset of classical paths.
Understanding how systems built out of modular components can be jointly optimized is an important problem in biology, engineering, and machine learning. The backpropagation algorithm is one such solution and has been instrumental in the success of neural networks. Despite its empirical success, a strong theoretical understanding of it is lacking. Here, we combine tools from Riemannian geometry, optimal control theory, and theoretical physics to advance this understanding. We make three key contributions: First, we revisit the derivation of backpropagation as a constrained optimization problem and combine it with the insight that Riemannian gradient descent trajectories can be understood as the minimum of an action. Second, we introduce a recursively defined layerwise Riemannian metric that exploits the modular structure of neural networks and can be efficiently computed using the Woodbury matrix identity, avoiding the O(n^3) cost of full metric inversion. Third, we develop a framework of composable “Riemannian modules” whose convergence properties can be quantified using nonlinear contraction theory, providing algorithmic stability guarantees of order O(κ^2 L/(ξμ√(n))) where κ and L are Lipschitz constants, μ is the mass matrix scale, and ξ bounds the condition number. Our layerwise metric approach provides a practical alternative to natural gradient descent. While we focus here on studying neural networks, our approach more generally applies to the study of systems made of modules that are optimized over time, as it occurs in biology during both evolution and development.
We derive a sufficient condition for $k$-contraction in a generalized Lurie system (GLS), that is, the feedback connection of a nonlinear dynamical system and a memoryless nonlinear function. For $k=1$, this reduces to a sufficient condition for standard contraction. For $k=2$, this condition implies that every bounded solution of the GLS converges to an equilibrium, which is not necessarily unique. We demonstrate the theoretical results by analyzing $k$-contraction in a biochemical control circuit with nonlinear dissipation terms.
Entropy-regularized optimal transport, which has strong links to the Schrödinger bridge problem in statistical mechanics, enjoys a variety of applications from trajectory inference to generative modeling. A major driver of renewed interest in this problem is the recent development of fast matrix-scaling algorithms—known as iterative proportional fitting or the Sinkhorn algorithm—for entropic optimal transport, which have favorable complexity over traditional approaches to the unregularized problem. Here, we take a perspective on this algorithm rooted in the thermodynamic origins of Schrödinger's problem and inspired by the modern geometric theory of diffusion: is the Sinkhorn flow (viewed in continuous-time as a mirror descent by recent results) the gradient flow of entropy in a formal Riemannian geometry? We answer this question affirmatively, finding a nonlocal Wasserstein gradient structure in the dynamics of its free marginal. This offers a physical interpretation of the Sinkhorn flow as the stochastic dynamics of a particle with law evolving by the nonlocal diffusion of a chemical potential. Simultaneously, it brings a standard suite of functional inequalities characterizing Markov diffusion processes to bear upon its geometry and convergence. We prove an entropy-energy (de Bruijn) identity, a Poincaré inequality, and a Bakry-Émery-type condition under which a logarithmic Sobolev inequality (LSI) holds and implies exponential convergence of the Sinkhorn flow in entropy. We lastly discuss computational applications such as stopping heuristics and latent-space design criteria leveraging the LSI and, returning to the physical interpretation, the possibility of natural systems whose relaxation to equilibrium inherently solves entropic optimal transport or Schrödinger bridge problems.
Astrocytes, the most abundant type of glial cell, play a fundamental role in memory. Despite most hippocampal synapses being contacted by an astrocyte, there are no current theories that explain how neurons, synapses, and astrocytes might collectively contribute to memory function. We demonstrate that fundamental aspects of astrocyte morphology and physiology naturally lead to a dynamic, high-capacity associative memory system. The neuron-astrocyte networks generated by our framework are closely related to popular machine learning architectures known as Dense Associative Memories. Adjusting the connectivity pattern, the model developed here leads to a family of associative memory networks that includes a Dense Associative Memory and a Transformer as two limiting cases. In the known biological implementations of Dense Associative Memories, the ratio of stored memories to the number of neurons remains constant, despite the growth of the network size. Our work demonstrates that neuron-astrocyte networks follow a superior memory scaling law, outperforming known biological implementations of Dense Associative Memory. Our model suggests an exciting and previously unnoticed possibility that memories could be stored, at least in part, within the network of astrocyte processes rather than solely in the synaptic weights between neurons.
Understanding the asymptotic behavior of reaction-diffusion (RD) systems is crucial for modeling processes ranging from species coexistence in ecology to biochemical interactions within cells. In this work, we analyze RD systems in which diffusion is modeled using the θ-diffusion framework, while the reaction dynamics are spatially varying. We demonstrate that spatial heterogeneity affects the asymptotic behavior of such systems. Using contraction theory, we derive conditions that guarantee the exponential convergence of system trajectories, regardless of initial conditions. These conditions explicitly account for the influence of spatial heterogeneity in both the diffusion and reaction terms. As an application, we study a biochemical system and derive the quasi-steady-state (QSS) approximation, illustrating how spatial heterogeneity modulates the effective binding rates of biomolecular species.
We show that the Schrödinger equation can be solved exactly based only on classical least action. Fundamental postulates of quantum mechanics can in turn be derived directly from this construction. The results extend to the relativistic Klein-Gordon, Pauli, Dirac, and Maxwell equations, and suggest a smooth transition between physics across scales. Most quantum mechanics problems have classical versions which involve multiple least action solutions. The associated classical multipaths stem either from the initial position or momentum distribution, or from branch points, generated, e.g., by a multiply connected manifold (double slit experiment), by spatial inequality constraints (particle in a box), or by a singularity (Coulomb potential). We show that the exact Schrödinger wave function ψ of the original quantum problem can be constructed by combining this classical multi-valued action ϕ with the density ρ of the classical position dynamics, where a key point is that ρ can be easily computed from ϕ along each extremal action path. The construction is general and does not involve any quasi-classical approximation. Examples illustrate how the quantum wave functions for the double-slit experiment or e.g., the hydrogen atom can be computed exactly from their classical least action counterparts. In a quantum measurement process, randomness originates from the determined forward mapping of an initial classical density distribution. In the Einstein-Podolsky-Rosen experiment, while Bell's inequalities are violated, from this perspective there is indeed a hidden variable in the form of a complex spinor. These results also provide a simpler computational alternative to Feynman path integrals, as they use only a minimal subset of classical paths and avoid zig-zag paths and time-slicing altogether.
Evaluating and updating the obstacle avoidance velocity for an autonomous robot in real-time ensures robustness against noise and disturbances. A passive damping controller can obtain the desired motion with a torque-controlled robot, which remains compliant and ensures a safe response to external perturbations. Here, we propose a novel approach for designing the passive control policy. Our algorithm complies with obstacle-free zones while transitioning to increased damping near obstacles to ensure collision avoidance. This approach ensures stability across diverse scenarios, effectively mitigating disturbances. Validation on a 7DoF robot arm demonstrates superior collision rejection capabilities compared to the baseline, underlining its practicality for real-world applications. Our obstacle-aware damping controller represents a substantial advancement in secure robot control within complex and uncertain environments.
This paper presents a modular framework for motion planning using movement primitives. Central to the approach is Contraction Theory, a modular stability tool for nonlinear dynamical systems. The approach extends prior methods by achieving parallel and sequential combinations of both discrete and rhythmic movements, while enabling independent modulation of each movement. This modular framework enables a divide-and-conquer strategy to simplify the programming of complex robot motion planning. Simulation examples illustrate the flexibility and versatility of the framework, highlighting its potential to address diverse challenges in robot motion planning.
A key property of successful learning algorithms is generalization. In classical supervised learning, generalization can be achieved by ensuring that the empirical error converges to the expected error as the number of training samples goes to infinity. Within this classical setting, we analyze the generalization properties of iterative optimizers such as stochastic gradient descent and natural gradient flow through the lens of dynamical systems and control theory. Specifically, we use contraction analysis to show that generalization and dynamical robustness are intimately related through the notion of algorithmic stability. In particular, we prove that Riemannian contraction in a supervised learning setting implies generalization. We show that if a learning algorithm is contracting in some Riemannian metric with rate $\lambda > 0$, it is uniformly algorithmically stable with rate $\mathcal{O}(1/\lambda n)$, where $n$ is the number of examples in the training set. The results hold for stochastic and deterministic optimization, in both continuous and discrete-time, for convex and non-convex loss surfaces. The associated generalization bounds reduce to well-known results in the particular case of gradient descent over convex or strongly convex loss surfaces. They can be shown to be optimal in certain linear settings, such as kernel ridge regression under gradient flow. Finally, we demonstrate that the well-known Polyak-Lojasiewicz condition is intimately related to the contraction of a model's outputs as they evolve under gradient descent. This correspondence allows us to derive uniform algorithmic stability bounds for nonlinear function classes such as wide neural networks.
Designs incorporating kinematic loops are becoming increasingly prevalent in the robotics community. Despite the existence of dynamics algorithms to deal with the effects of such loops, many modern simulators rely on dynamics libraries that require robots to be represented as kinematic trees. This requirement is reflected in the de facto standard format for describing robots, the Universal Robot Description Format (URDF), which does not support kinematic loops resulting in closed chains. This paper introduces an enhanced URDF, termed URDF+, which addresses this key shortcoming of URDF while retaining the intuitive design philosophy and low barrier to entry that the robotics community values. The URDF+ keeps the elements used by URDF to describe open chains and incorporates new elements to encode loop joints. We also offer an accompanying parser that processes the system models coming from URDF+ so that they can be used with recursive rigid-body dynamics algorithms for closed-chain systems that group bodies into local, decoupled loops. This parsing process is fully automated, ensuring optimal grouping of constrained bodies without requiring manual specification from the user. We aim to advance the robotics community towards this elegant solution by developing efficient and easy-to-use software tools.
The wave-particle duality and its probabilistic interpretation are at the heart of quantum mechanics. Here we show that, in some standard contexts like the double slit experiment, a deterministic interpretation can be provided. This interpretation arises from the fact that if one mininizes a deterministic action subject to spatial inequality contraints, in general the solution is not unique. Thus, the probabilistic setting is replaced in part by the non-uniqueness of solutions of certain constrained optimization problems, itself often a result of the non-Lipschitzness of the constrained dynamics. While the approach leaves the results of associated Feynman integrals unchanged, it may considerably simplify their computation as the problems grow more complex, since only specific deterministic paths computed from the constrained minimization need to be included in the integral.
We present a new direct adaptive control approach for nonlinear systems with unmatched and matched uncertainties. The method relies on adjusting the adaptation gains of individual unmatched parameters whose adaptation transients would otherwise destabilize the closed-loop system. The approach also guarantees the restoration of the adaptation gains to their nominal values and can readily incorporate direct adaptation laws for matched uncertainties. The proposed framework is general as it only requires stabilizability for all possible models.
This paper presents an algorithm to geometrically characterize inertial parameter identifiability for an articulated robot. The geometric approach tests identifiability across the infinite space of configurations using only a finite set of conditions and without approximation. It can be applied to general open-chain kinematic trees ranging from industrial manipulators to legged robots, and it is the first solution for this broad set of systems that is provably correct. The high-level operation of the algorithm is based on a key observation: Undetectable changes in inertial parameters can be represented as sequences of inertial transfers across the joints. Drawing on the exponential parameterization of rigid-body kinematics, undetectable inertial transfers are analyzed in terms of observability from linear systems theory. This analysis can be applied recursively, and lends an overall complexity of $O(N)$ to characterize parameter identifiability for a system of $N$ bodies. Matlab source code for the new algorithm is provided.
Controlling complex tasks in robotic systems, such as circular motion for cleaning or following curvy lines, can be dealt with using nonlinear vector fields. This article introduces a novel approach called the rotational obstacle avoidance method (ROAM) for adapting the initial dynamics when obstacles partially occlude the workspace. ROAM presents a closed-form solution that effectively avoids star-shaped obstacles in spaces of arbitrary dimensions by rotating the initial dynamics toward the tangent space. The algorithm enables navigation within obstacle hulls and can be customized to actively move away from surfaces while guaranteeing the presence of only a single saddle point on the boundary of each obstacle. We introduce a sequence of mappings to extend the approach for general nonlinear dynamics. Moreover, ROAM extends its capabilities to handle multiobstacle environments and provides the ability to constrain dynamics within a safe tube. By utilizing weighted vector-tree summation, we successfully navigate around general concave obstacles represented as a tree-of-stars. Through experimental evaluation, ROAM demonstrates superior performance in minimizing occurrences of local minima and maintaining similarity to the initial dynamics, outperforming existing approaches in multiobstacle simulations. Due to its simplicity, the proposed method is highly reactive and can be applied effectively in dynamic environments. This was demonstrated during the collision-free navigation of a 7-degree-of-freedom robot arm around dynamic obstacles.
We present two quaternion-based sliding variables for controlling the orientation of a manipulator's end-effector. Both sliding variables are free of singularities and represent global exponentially convergent error dynamics that do not exhibit unwinding when used in feedback. The choice of sliding variable is dictated by whether the end-effector's angular velocity vector is expressed in a local or global frame, and is a matter of convenience. Using quaternions allows the end-effector to move in its full operational envelope, which is not possible with other representations, e.g., Euler angles, that introduce representation-specific singularities. Further, the presented stability results are global rather than almost global, where the latter is often the best one can achieve when using rotation matrices to represent orientation.
Nicolas Tabareau合作论文数Departement Informatique
Ecole des Mines de Nantes14
Alain Berthoz合作论文数Laboratoire de Physiologie de la Perception et de l'Action8