Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fine-tuning or scientific discovery demand diversity. Existing remedies such as entropy regularization or diversity bonuses often require fragile trade-offs that sacrifice performance for stochasticity or rely on heuristic metrics that can misalign policy rankings. We argue that diversity is more naturally understood as the rational response to uncertainty in the reward. When the reward function is not perfectly known–as is the case with ambiguous preferences or imperfect reward models–committing to a single action can be sub-optimal. Building on this, we propose a fundamental reformulation of the RL objective by replacing the scalar reward with a distribution over reward functions, and applying a non-linear objective over sets of actions. The result is a framework in which calibrated behavioural diversity emerges naturally, remains controllable through the reward function distribution, and is obtained without sacrificing expected reward. Focusing on the contextual bandit setting, we derive a principled gradient estimator for this objective and prove that our formulation naturally generalizes both vanilla policy gradient and more recently developed action-set approaches. Our empirical results demonstrate that this framework offers a robust and theoretically grounded alternative for complex RL tasks where the traditional formulation of the problem fails to induce the desired breadth of agent behaviour.
Reinforcement Learning from Human Feedback (RLHF) effectively aligns Large Language Models (LLMs) with aggregate human preferences but often fails to address the diverse and conflicting needs of individual users. To overcome this issue, we introduce Spectral Souping, a unified framework for efficient, online preference alignment. Our contribution is the discovery of a universal spectral representation within LLMs, which is proven to be highly amenable to model merging. This theoretical insight enables a two-phase methodology: we first learn a basis of specialized policies offline, each focused on a distinct, fine-grained preference dimension. An online adaptation algorithm then efficiently “soups” these policies at inference time, either by merging their outputs or parameters, enabling rapid model adaptation without the need for costly online retraining w.r.t. tailored preference rewards. Experiments on online preference alignment benchmarks demonstrate that our method achieves significant performance improvements over existing state-of-the-art approaches, presenting a scalable and computationally efficient solution for dynamically adapting LLMs to individual user preferences.
We introduce distributional dynamic programming (DP) methods for optimizing statistical functionals of the return distribution, with standard reinforcement learning as a special case. Previous distributional DP methods could optimize the same class of expected utilities as classic DP. To go beyond, we combine distributional DP with stock augmentation, a technique previously introduced for classic DP in the context of risk-sensitive RL, where the MDP state is augmented with a statistic of the rewards obtained since the first time step. We find that a number of recently studied problems can be formulated as stock-augmented return distribution optimization, and we show that we can use distributional DP to solve them. We analyze distributional value and policy iteration, with bounds and a study of what objectives these distributional DP methods can or cannot optimize. We describe a number of applications outlining how to use distributional DP to solve different stock-augmented return distribution optimization problems, for example maximizing conditional value-at-risk, and homeostatic regulation. To highlight the practical potential of stock-augmented return distribution optimization and distributional DP, we introduce an agent that combines DQN and the core ideas of distributional DP, and empirically evaluate it for solving instances of the applications discussed.
Reinforcement learning from human feedback usually models preferences using a reward function that does not distinguish between people. We argue that this is unlikely to be a good design choice in contexts with high potential for disagreement, like in the training of large language models. We formalise and analyse the problem of learning a reward model that can be specialised to a user. Using the principle of empirical risk minimisation, we derive a probably approximately correct (PAC) bound showing the dependency of the approximation error on the number of training examples, as usual, and also on the number of human raters who provided feedback on them. Based on our theoretical findings, we discuss how to best collect pairwise preference data and argue that adaptive reward models should be beneficial when there is considerable disagreement among users. We also propose a concrete architecture for an adaptive reward model. Our approach leverages the observation that individual preferences can be captured as a linear combination of a set of general reward features. We show how to learn such features and subsequently use them to quickly adapt the reward model to a specific individual, even if their preferences are not reflected in the training data. We present experiments with large language models illustrating our theoretical results and comparing the proposed architecture with a non-adaptive baseline. Consistent with our analysis, the benefits provided by our model increase with the number of raters and the heterogeneity of their preferences. We also show that our model compares favourably to adaptive counterparts, including those performing in-context personalisation.
Agents are minimally entities that are influenced by their past observations and act to influence future observations. This latter capacity is captured by empowerment, which has served as a vital framing concept across artificial intelligence and cognitive science. This former capacity, however, is equally foundational: In what ways, and to what extent, can an agent be influenced by what it observes? In this paper, we ground this concept in a universal agent-centric measure that we refer to as plasticity, and reveal a fundamental connection to empowerment. Following a set of desiderata on a suitable definition, we define plasticity using a new information-theoretic quantity we call the generalized directed information. We show that this new quantity strictly generalizes the directed information introduced by Massey (1990) while preserving all of its desirable properties. Under this definition, we find that plasticity is well thought of as the mirror of empowerment: The two concepts are defined using the same measure, with only the direction of influence reversed. Our main result establishes a tension between the plasticity and empowerment of an agent, suggesting that agent design needs to be mindful of both characteristics. We explore the implications of these findings, and suggest that plasticity, empowerment, and their relationship are essential to understanding agency
Agency is a system's capacity to steer outcomes toward a goal, and is a central topic of study across biology, philosophy, cognitive science, and artificial intelligence. Determining if a system exhibits agency is a notoriously difficult question: Dennett (1989), for instance, highlights the puzzle of determining which principles can decide whether a rock, a thermostat, or a robot each possess agency. We here address this puzzle from the viewpoint of reinforcement learning by arguing that agency is fundamentally frame-dependent: Any measurement of a system's agency must be made relative to a reference frame. We support this claim by presenting a philosophical argument that each of the essential properties of agency proposed by Barandiaran et al. (2009) and Moreno (2018) are themselves frame-dependent. We conclude that any basic science of agency requires frame-dependence, and discuss the implications of this claim for reinforcement learning.
Multi-task reinforcement learning aims to quickly identify solutions for new tasks with minimal or no additional interaction with the environment. Generalized Policy Improvement (GPI) addresses this by combining a set of base policies to produce a new one that is at least as good—though not necessarily optimal—as any individual base policy. Optimality can be ensured, particularly in the linear-reward case, via techniques that compute a Convex Coverage Set (CCS). However, these are computationally expensive and do not scale to complex domains. The Option Keyboard (OK) improves upon GPI by producing policies that are at least as good—and often better. It achieves this through a learned meta-policy that dynamically combines base policies. However, its performance critically depends on the choice of base policies. This raises a key question: is there an optimal set of base policies—an optimal *behavior basis*—that enables zero-shot identification of optimal solutions for *any* linear tasks? We solve this open problem by introducing a novel method that efficiently constructs such an optimal behavior basis. We show that it significantly reduces the number of base policies needed to ensure optimality in new tasks. We also prove that it is strictly more expressive than a CCS, enabling particular classes of *non-linear* tasks to be solved optimally. We empirically evaluate our technique in challenging domains and show that it outperforms state-of-the-art approaches, increasingly so as task complexity increases.
The whiteleg shrimp, Penaeus vannamei, is a highly valued and globally produced crustacean species. However, the rising cost of shrimp feed, exacerbated by increasing cereal prices, prompts the exploration of cost-effective and sustainable formulations. This study investigates the potential of Salicornia ramosissima biomass by-product (the non-edible part) as a substitute for wheat meal in juvenile shrimp diets, aiming to create sustainable formulations. Particularly to assess the impact of incorporating S. ramosissima into shrimp aquafeeds on various aspects of shrimp development, including growth performance, survival, immune status, and oxidative status. A commercial-like diet was formulated and served as control, whereas four other diets contained S. ramosissima stems or a combination of leaves and seeds, both at inclusion levels of 5% and 10%. Shrimps were fed the experimental diets for 31 and 55 days, followed by a bacterial bath challenge test to gauge their immune response to pathogens. At the end of the feeding period, growth performance and survival rates remained consistent across all diets. However, shrimp fed diets with S. ramosissima consumed more feed to achieve similar weights of those fed the control diet, particularly in diets containing leaves and seeds at a 10% inclusion level, likely due to lower digestibility of dry matter, lipids, and energy. While S. ramosissima biomass inclusion did not affect shrimp weight, relative growth rate, or survival, it did lead to higher feed conversion ratios and feed intake. Additionally, S. ramosissima inclusion affected shrimps’ overall body composition, particularly moisture and ash content. S. ramosissima inclusion modulated antioxidant enzyme activity in the shrimp’s hepatopancreas, indicating potential health improvements. The observed gene expression changes related to antioxidant enzymes, points to an overall down-regulation with the inclusion of S. ramosissima. Despite challenges in feeding efficiency, the inclusion of S. ramosissima, especially stems, shows promise in reducing feed costs by utilizing a food agro-industrial by-products (non edible parts). Furthermore, S. ramosissima inclusion led to subtle changes in certain plasma humoral parameters. In conclusion, this study highlights the potential of this halophyte as a functional feed ingredient capable of enhancing shrimp’s antioxidant response, aligning with global resource optimization and sustainability initiatives.
The green tips of Salicornia ramosissima are used for human consumption, while, in a production scenario, the rest of the plant is considered a residue. This study evaluated the potential of incorporating salicornia by-products in diets for juvenile European seabass, partially replacing wheat meal, aspiring to contribute to their valorization. A standard diet and three experimental diets including salicornia in 2.5%, 5% and 10% inclusion levels were tested in triplicate. After 62 days of feeding, no significant differences between treatments were observed in fish growth performances, feeding efficiency and economic conversation ratio. Nutrient digestibility of the experimental diets was unaffected by the inclusion of salicornia when compared to a standard diet. Additionally, salicornia had significant modulatory effects on the fish muscle biochemical profiles, namely by significantly decreasing lactic acid and increasing succinic acid levels, which can potentially signal health-promoting effects for the fish. Increases in DHA levels in fish fed a diet containing 10% salicornia were also shown. Therefore, the results suggest that salicornia by-products are a viable alternative to partially replace wheat meal in diets for juvenile European seabass, contributing to the valorization of a residue and the implementation of a circular economy paradigm in halophyte farming and aquaculture.
Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction. However, they have not been equally leveraged: language models have had significant real-world impact, whereas video generation has remained largely limited to media entertainment. Yet video data captures important information about the physical world that is difficult to express in language. To address this gap, we discuss an under-appreciated opportunity to extend video generation to solve tasks in the real world. We observe how, akin to language, video can serve as a unified interface that can absorb internet knowledge and represent diverse tasks. Moreover, we demonstrate how, like language models, video generation can serve as planners, agents, compute engines, and environment simulators through techniques such as in-context learning, planning and reinforcement learning. We identify major impact opportunities in domains such as robotics, self-driving, and science, supported by recent work that demonstrates how such advanced capabilities in video generation are plausibly within reach. Lastly, we identify key challenges in video generation that mitigate progress. Addressing these challenges will enable video generation models to demonstrate unique value alongside language models in a wider array of AI applications.
This paper contributes a new approach for distributional reinforcement learning which elucidates a clean separation of transition structure and reward in the learning process. Analogous to how the successor representation (SR) describes the expected consequences of behaving according to a given policy, our distributional successor measure (SM) describes the distributional consequences of this behaviour. We formulate the distributional SM as a distribution over distributions and provide theory connecting it with distributional and model-based reinforcement learning. Moreover, we propose an algorithm that learns the distributional SM from data by minimizing a two-level maximum mean discrepancy. Key to our method are a number of algorithmic techniques that are independently valuable for learning generative models of state. As an illustration of the usefulness of the distributional SM, we show that it enables zero-shot risk-sensitive policy evaluation in a way that was not previously possible.
Whiteleg shrimp Penaeus vannamei farming in clear water recirculating aquaculture systems (RAS) is relatively recent, and consequently, knowledge on the shrimp dietary demands is still insufficient, particularly in the initial developmental stages. This study aimed at assessing the dietary protein requirement of whiteleg shrimp post-larvae (PL) in a clear-water recirculating aquaculture system (RAS). Six microdiets were formulated to contain 34%, 44%, 49%, 54%, 58%, and 63% crude protein (P34, P44, P49, P54, P58 and P63, respectively) and were evaluated in triplicates. Whiteleg shrimp PL (3.2 mg wet weight) were reared for 21 days in a clear-water RAS at Riasearch Lda. At the end of the feeding period, the optimal protein requirement was estimated at 47.1%, 46.4%, 47.2%, and 44.0% for weight gain, relative growth rate (RGR), feed conversion ratio (FCR), and survival, respectively. PL fed the P54, P58, and P63 diets achieved significantly higher final body weights than those fed P34. PL fed P34 showed significantly lower RGR and survival and significantly higher FCR values than those fed the remaining diets, suggesting that low protein diets may not be adequate to be used in this stage of shrimp development and/or for the clear-water RAS husbandry conditions. Moreover, diet P34 seemingly reduced the overall antioxidant status of the PL when compared to P44, P49, and P54. However, the P34 diet seems to have stimulated the PL immune mechanisms when compared to P44, P49, and P54, possibly due to increased levels of fish and algae oil. Similarly, despite the good growth performances, a diet containing 63% of protein also seemed to have compromised the overall shrimp PL antioxidant status and stimulate their immune system. Shrimp fed diet P54 showed an apparent overall superior antioxidant status when compared to the remaining diets, evidencing that using protein inclusion levels up to 54% in aquafeeds not only potentiates growth performances and survival but also can potentially be beneficial to the health status of P. vannamei PL grown in a clear-water RAS. Hence, results from this study suggest that a minimum of approximately 47% of protein should be considered when tailoring microdiets for whiteleg shrimp PL grown in a clear-water RAS, but inclusion levels up to 54% can be used with benefits to the PL antioxidant status.
A growing body of evidence suggests that neural networks employed in deep reinforcement learning (RL) gradually lose their plasticity, the ability to learn from new data; however, the analysis and mitigation of this phenomenon is hampered by the complex relationship between plasticity, exploration, and performance in RL. This paper introduces plasticity injection, a minimalistic intervention that increases the network plasticity without changing the number of trainable parameters or biasing the predictions. The applications of this intervention are two-fold: first, as a diagnostic tool $\unicode{x2014}$ if injection increases the performance, we may conclude that an agent's network was losing its plasticity. This tool allows us to identify a subset of Atari environments where the lack of plasticity causes performance plateaus, motivating future studies on understanding and combating plasticity loss. Second, plasticity injection can be used to improve the computational efficiency of RL training if the agent has to re-learn from scratch due to exhausted plasticity or by growing the agent's network dynamically without compromising performance. The results on Atari show that plasticity injection attains stronger performance compared to alternative methods while being computationally efficient.
Planning is one of the oldest and most important problems in artificial intelligence. Simulation-based search algorithms, such as AlphaZero, have achieved superhuman performance in chess and Go and are used widely in real-world applications of planning. In this paper we provide a unified framework for simulation-based search. Algorithms in this framework interleave operators for policy evaluation (better estimating the value function of the current policy) and policy improvement (using the value function to form a better policy). These operators are applied to states and actions that are sampled in sequential trajectories, and that may branch recursively into other sampled trajectories. The value function and policy may also be represented by a function approximator. Our framework includes a broad family of search algorithms that includes Monte-Carlo tree search, sparse sampling, nested Monte-Carlo search, classification-based policy iteration, and AlphaZero.
In a standard view of the reinforcement learning problem, an agent's goal is to efficiently identify a policy that maximizes long-term reward. However, this perspective is based on a restricted view of learning as finding a solution, rather than treating learning as endless adaptation. In contrast, continual reinforcement learning refers to the setting in which the best agents never stop learning. Despite the importance of continual reinforcement learning, the community lacks a simple definition of the problem that highlights its commitments and makes its primary concepts precise and clear. To this end, this paper is dedicated to carefully defining the continual reinforcement learning problem. We formalize the notion of agents that "never stop learning" through a new mathematical language for analyzing and cataloging agents. Using this new language, we define a continual learning agent as one that can be understood as carrying out an implicit search process indefinitely, and continual reinforcement learning as the setting in which the best agents are all continual learning agents. We provide two motivating examples, illustrating that traditional views of multi-task reinforcement learning and continual supervised learning are special cases of our definition. Collectively, these definitions and perspectives formalize many intuitive concepts at the heart of learning, and open new research pathways surrounding continual learning agents.
The identification of novel feed materials as a source of functional ingredients is a topical priority in the finfish aquaculture sector. Due to the agrotechnical practices associated and phytochemical profiling, halophytes emerge as a new source of feedstuff for aquafeeds, with the potential to boost productivity and environmental sustainability. Therefore, the present study aimed to assess the potential of Salicornia ramosissima incorporation (2.5, 5, and 10%), for 2 months, in the diet of juvenile European seabass, seeking antioxidant (in the liver, gills, and blood) and genoprotective (DNA and chromosomal integrity in blood) benefits. Halophyte inclusion showed no impairments on growth performance. Moreover, a tissue-specific antioxidant improvement was apparent, namely through the GSH-related defense subsystem, but revealing multiple and complex mechanisms. A genotoxic trigger (regarded as a pro-genoprotective mechanism) was identified in the first month of supplementation. A clear protection of DNA integrity was detected in the second month, for all the supplementation levels (and the most prominent melioration at 10%). Overall, these results pointed out a functionality of S. ramosissima-supplemented diets and a promising way to improve aquaculture practices, also unraveling a complementary novel, low-value raw material, and a path to its valorization.
Dietary additives have the potential to stimulate the whiteleg shrimp immune system, but information is scarce on their use in diets for larval/post-larval stages. The potential beneficial effects of vitamins C and E, β-glucans, taurine, and methionine were evaluated. Four experimental microdiets were tested: a positive control diet (PC); the PC with decreased levels of vitamin C and E as negative control (NC); the PC with increased taurine and methionine levels (T + M); and the PC supplemented with β-glucans (BG). No changes in growth performance and survival were observed. However, post-larvae shrimp fed the NC had lower relative expressions of pen-3 than those fed the PC, suggesting that lower levels of vitamins C and E may impact the shrimp immune status. Lipid peroxidation levels dropped significantly in the BG compared to the PC, indicating that β-glucans improved the post-larvae antioxidant mechanisms. Furthermore, when compared with the NC diet, PL fed with BG showed significant increases in tGSH levels and in the relative expression of crus and pen-3, suggesting a synergistic effect between vitamins C and E and β-glucans. Amongst the additives tested, β-glucans seems to be the most promising even when compared to a high-quality control diet.
When has an agent converged? Standard models of the reinforcement learning problem give rise to a straightforward definition of convergence: An agent converges when its behavior or performance in each environment state stops changing. However, as we shift the focus of our learning problem from the environment's state to the agent's state, the concept of an agent's convergence becomes significantly less clear. In this paper, we propose two complementary accounts of agent convergence in a framing of the reinforcement learning problem that centers around bounded agents. The first view says that a bounded agent has converged when the minimal number of states needed to describe the agent's future behavior cannot decrease. The second view says that a bounded agent has converged just when the agent's performance only changes if the agent's internal state changes. We establish basic properties of these two definitions, show that they accommodate typical views of convergence in standard settings, and prove several facts about their nature and relationship. We take these perspectives, definitions, and analysis to bring clarity to a central idea of the field.
Reasoning at multiple levels of temporal abstraction is one of the key attributes of intelligence. In reinforcement learning, this is often modeled through temporally extended courses of actions called options. Options allow agents to make predictions and to operate at different levels of abstraction within an environment. Nevertheless, approaches based on the options framework often start with the assumption that a reasonable set of options is known beforehand. When this is not the case, there are no definitive answers for which options one should consider. In this paper, we argue that the successor representation (SR), which encodes states based on the pattern of state visitation that follows them, can be seen as a natural substrate for the discovery and use of temporal abstractions. To support our claim, we take a big picture view of recent results, showing how the SR can be used to discover options that facilitate either temporally-extended exploration or planning. We cast these results as instantiations of a general framework for option discovery in which the agent's representation is used to identify useful options, which are then used to further improve its representation. This results in a virtuous, never-ending, cycle in which both the representation and the options are constantly refined based on each other. Beyond option discovery itself, we also discuss how the SR allows us to augment a set of options into a combinatorially large counterpart without additional learning. This is achieved through the combination of previously learned options. Our empirical evaluation focuses on options discovered for exploration and on the use of the SR to combine them. The results of our experiments shed light on important design decisions involved in the definition of options and demonstrate the synergy of different methods based on the SR, such as eigenoptions and the option keyboard.
Using a model of the environment and a value function, an agent can construct many estimates of a state's value, by unrolling the model for different lengths and bootstrapping with its value function. Our key insight is that one can treat this set of value estimates as a type of ensemble, which we call an implicit value ensemble (IVE). Consequently, the discrepancy between these estimates can be used as a proxy for the agent's epistemic uncertainty; we term this signal model-value inconsistency or self-inconsistency for short. Unlike prior work which estimates uncertainty by training an ensemble of many models and/or value functions, this approach requires only the single model and value function which are already being learned in most model-based reinforcement learning algorithms. We provide empirical evidence in both tabular and function approximation settings from pixels that self-inconsistency is useful (i) as a signal for exploration, (ii) for acting safely under distribution shifts, and (iii) for robustifying value-based planning with a learned model.