Fleets of robo-taxis offering on-demand transportation services, commonly known as Autonomous Mobility-on-Demand (AMoD) systems, hold significant promise for societal benefits, such as reducing pollution, energy consumption, and urban congestion. However, orchestrating these systems at scale remains a critical challenge, with existing coordination algorithms often failing to exploit the systems' full potential. This work introduces a novel decision-making framework that unites mathematical modeling with data-driven techniques. In particular, we present the AMoD coordination problem through the lens of reinforcement learning and propose a graph network-based framework that exploits the main strengths of graph representation learning, reinforcement learning, and classical operations research tools. Extensive evaluations across diverse simulation fidelities and scenarios demonstrate the flexibility of our approach, achieving superior system performance, computational efficiency, and generalizability compared to prior methods. Finally, motivated by the need to democratize research efforts in this area, we release publicly available benchmarks, datasets, and simulators for network-level coordination alongside an open-source codebase designed to provide accessible simulation platforms and establish a standardized validation process for comparing methodologies. Code available at: https://github.com/StanfordASL/RL4AMOD
Autonomous mobility-on-demand (AMoD) systems, powered by advances in robotics, control, and machine learning (ML), offer a promising paradigm for future urban transportation. AMoD offers fast and personalized travel services by leveraging centralized control of autonomous vehicle fleets to optimize operations and enhance service performance. However, the rapid growth of this field has outpaced the development of standardized practices for evaluating and reporting results, leading to significant challenges in reproducibility. As AMoD control algorithms become increasingly complex and data-driven, a lack of transparency in modeling assumptions, experimental setups, and algorithmic implementation hinders scientific progress and undermines confidence in the results. This article presents a systematic study of reproducibility in AMoD research. We identify key components across the research pipeline, spanning system modeling, control problems, simulation design, algorithm specification, and evaluation, and analyze common sources of irreproducibility. We survey prevalent practices in the literature, highlight gaps, and propose a structured framework to assess and improve reproducibility. While focused on AMoD, the principles and practices we advocate generalize to a broader class of cyber-physical systems that rely on networked autonomy and data-driven control. This work aims to lay the foundation for a more transparent and reproducible research culture in the design and deployment of intelligent mobility systems.
Learned optimization aims to improve upon hand-designed optimizers (e.g., Adam and Muon) by meta-learning small neural network optimizers over a distribution of tasks. While recent work has greatly advanced the architectural design and inductive biases of learned optimizers (LOs), current meta-training approaches still suffer from two main difficulties: (1) they cannot efficiently scale meta-training to long-horizon inner problems and (2) they often fail to surpass comparable hand-designed optimizers. To address these limitations, we propose Efficient Long-hOrizon (ELO) learning, an efficient meta-training algorithm that (1) reallocates redundant meta-training compute to longer failure regimes, achieving efficient long-horizon learning, and (2) enforces decoupled progressive expert supervision, providing stable meta-learning signals that additionally improve the generalization of LOs. Our empirical study evaluates ELO for meta-training both element-wise and matrix-based LOs. Across downstream language modeling (GPT-2-124M/350M on FineWeb) and image classification (ViT-B/16, ResNet-50 on ImageNet-1K) tasks, ELO substantially improves the long-unroll performance and out-of-distribution generalization of the base LOs. In particular, ELO-Celo2 consistently outperforms well-tuned AdamW across all evaluated tasks, while remaining competitive with Muon on language modeling. Notably, all ELO baselines require less than 7 H100 GPU-hours for meta-training.
Gaussian Processes (GPs) are widely seen as the state-of-the-art surrogate models for Bayesian optimization (BO) due to their ability to model uncertainty and their performance on tasks where correlations are easily captured (such as those defined by Euclidean metrics) and their ability to be efficiently updated online. However, the performance of GPs depends on the choice of kernel, and kernel selection for complex correlation structures is often difficult or must be made bespoke. While Bayesian neural networks (BNNs) are a promising direction for higher capacity surrogate models, they have so far seen limited use due to poor performance on some problem types. In this paper, we propose an approach which shows competitive performance on many problem types, including some that BNNs typically struggle with. We build on variational Bayesian last layers (VBLLs), and connect training of these models to exact conditioning in GPs. We exploit this connection to develop an efficient online training algorithm that interleaves conditioning and optimization. Our findings suggest that VBLL networks significantly outperform GPs and other BNN architectures on tasks with complex input correlations, and match the performance of well-tuned GPs on established benchmark tasks.
Hierarchical policies enable strong performance in many sequential decision-making problems, such as those with high-dimensional action spaces, those requiring long-horizon planning, and settings with sparse rewards. However, learning hierarchical policies from static offline datasets presents a significant challenge.Crucially, actions taken by higher-level policies may not be directly observable within hierarchical controllers, and the offline dataset might have been generated using a different policy structure, hindering the use of standard offline learning algorithms.In this work, we propose $\textit{OHIO}$: a framework for offline reinforcement learning (RL) of hierarchical policies. Our framework leverages knowledge of the policy structure to solve the $\textit{inverse problem}$, recovering the unobservable high-level actions that likely generated the observed data under our hierarchical policy.This approach constructs a dataset suitable for off-the-shelf offline training.We demonstrate our framework on robotic and network optimization problems and show that it substantially outperforms end-to-end RL methods and improves robustness. We investigate a variety of instantiations of our framework, both in direct deployment of policies trained offline and when online fine-tuning is performed. Code and data are available at https://ohio-offline-hierarchical-rl.github.io.
Background: The Johns Hopkins Activity and Mobility Program is a systematic approach to measure and improve patient mobility. Purpose: The purpose of this study was to evaluate the relationship between mobility loss and quality outcomes. Methods: A retrospective cohort study design was used. Patients were categorized into 3 groups (gain, loss, no change in mobility) using the Johns Hopkins Highest Level of Mobility (JH-HLM) scores. The association between mobility loss and falls risk, in-hospital mortality, delirium, discharge to a facility, length of stay, and 30 day readmissions were assessed. Results: Those who lost mobility were more at risk of being a high fall risk, in-hospital mortality, delirium, discharging to a facility, and had 48% longer lengths of stay. There was no association between mobility loss and 30-day readmissions. Conclusions: Loss of mobility assessed using JH-HLM scores is associated with worse patient outcomes.
Background A better understanding of how chronic physical health conditions affect long-term outcomes following injury is essential for quantifying the burden of serious orthopaedic injuries. We aimed to describe the association between the presence of post-injury chronic physical health conditions and (i) the change in health status from before injury to six different follow-up time points after injury; and (ii) survival time. Methods A cohort study was conducted using linked data from the REcovery after Serious Trauma: Outcomes, Resource Use, and Patient Experiences study, the Victorian Registry of Births, Deaths and Marriages (BDM) (2009–2017), the victorian admitted episodes dataset (2009–2017) and the victorian emergency minimum dataset (2009–2017). Adults (≥ 18 years old) with serious orthopaedic injuries who survived to discharge from their trauma admission were included. Multivariable linear regression analysis was conducted to evaluate the association between post-injury chronic physical health conditions and the mean change in health status (EuroQol-Visual Analogue Scale) from before injury to six follow-up time points post-injury. Survival analysis was conducted to estimate the probability of survival for people with and without chronic physical health conditions following injury. Results Out of 894 participants, 177 (19.8%) had at least one chronic physical health condition recorded up to five years post-injury. People with post-injury conditions reported a greater mean decline in health status than people without post-injury conditions (difference, (95% CI): −6.9 (−9.7, −4.2), p = 0.01). Over the study period, almost six times as many people with chronic physical health conditions post-injury died as people without these conditions (AHR (95% CI): 5.7 (2.9, 11.3), p < 0.01). Conclusions Chronic physical conditions after serious orthopaedic injuries were associated with a lower survival probability and a deteriorating health status. Orthopaedic injury survivors may benefit from early detection and treatment of chronic conditions.
Fine-tuning language models~(LMs) on human-generated data remains a prevalent practice. However, the performance of such models is often limited by the quantity and diversity of high-quality human data. In this paper, we explore whether we can go beyond human data on tasks where we have access to scalar feedback, for example, on math problems where one can verify correctness. To do so, we investigate a simple self-training method based on expectation-maximization, which we call ReST$^{EM}$, where we (1) generate samples from the model and filter them using binary feedback, (2) fine-tune the model on these samples, and (3) repeat this process a few times. Testing on advanced MATH reasoning and APPS coding benchmarks using PaLM-2 models, we find that ReST$^{EM}$ scales favorably with model size and significantly surpasses fine-tuning only on human data. Overall, our findings suggest self-training with feedback can substantially reduce dependence on human-generated data.
INTRODUCTION:Road safety has been a long-enduring policy concern in Australia, with significant financial burden of road trauma and evident socioeconomic disparities. Transport injuries disproportionately impact individuals in remote areas, those in lower socioeconomic situations, and Aboriginal and Torres Strait Islander populations. There is a lack of insight into transport injuries in Aboriginal and Torres Strait Islander communities, absence of Indigenous perspective in published research and limited utilisation of linked data assets to address the inequity. Aim 1 is to determine the breadth, cost and causal factors of serious injury from road traffic crashes in South Australia (SA) and New South Wales (NSW) with a focus on injury prevention. Aim 2 is to identify enablers and barriers to compensation schemes for Aboriginal and Torres Strait Islander patients in SA and NSW. METHODS AND ANALYSIS:This study will be guided by an Aboriginal and Torres Strait Islander Governance Group, applying Knowledge Interface Methodology and Indigenous research principles to ensure Indigenous Data Sovereignty and incorporation of informed perspectives. A mixed-method approach will be undertaken to explore study aims including using big data assets and mapping patient journey. CONCLUSION:The results of this study will provide valuable insights for the development of focused injury prevention strategies and policies tailored to Aboriginal and Torres Strait Islander communities. By addressing the specific needs and challenges faced by these communities, the study aims to enhance road safety outcomes and promote equitable access to healthcare and compensation for affected individuals and their families.
We study the robustness of deep reinforcement learning algorithms against distribution shifts within contextual multi-stage stochastic combinatorial optimization problems from the operations research domain. In this context, risk-sensitive algorithms promise to learn robust policies. While this field is of general interest to the reinforcement learning community, most studies up-to-date focus on theoretical results rather than real-world performance. With this work, we aim to bridge this gap by formally deriving a novel risk-sensitive deep reinforcement learning algorithm while providing numerical evidence for its efficacy. Specifically, we introduce discrete Soft Actor-Critic for the entropic risk measure by deriving a version of the Bellman equation for the respective Q-values. We establish a corresponding policy improvement result and infer a practical algorithm. We introduce an environment that represents typical contextual multi-stage stochastic combinatorial optimization problems and perform numerical experiments to empirically validate our algorithm's robustness against realistic distribution shifts, without compromising performance on the training distribution. We show that our algorithm is superior to risk-neutral Soft Actor-Critic as well as to two benchmark approaches for robust deep reinforcement learning. Thereby, we provide the first structured analysis on the robustness of reinforcement learning under distribution shifts in the realm of contextual multi-stage stochastic combinatorial optimization problems.
A challenging problem in many modern machine learning tasks is to process weight-space features, i.e., to transform or extract information from the weights and gradients of a neural network. Recent works have developed promising weight-space models that are equivariant to the permutation symmetries of simple feedforward networks. However, they are not applicable to general architectures, since the permutation symmetries of a weight space can be complicated by recurrence or residual connections. This work proposes an algorithm that automatically constructs permutation equivariant models, which we refer to as universal neural functionals (UNFs), for any weight space. Among other applications, we demonstrate how UNFs can be substituted into existing learned optimizer designs, and find promising improvements over prior methods when optimizing small image classifiers and language models. Our results suggest that learned optimizers can benefit from considering the (symmetry) structure of the weight space they optimize.
Many patients are unable to identify members of their hospital care team and experience confusion regarding some medical terminology used during hospitalization, including descriptions of the structure of their inpatient care team. This cross-sectional study sought to (1) examine inpatients' understanding of the role of a hospitalist and (2) assess inpatients' familiarity with other medical terminology commonly used in the hospital. We surveyed 172 patients admitted to the hospital medicine service at two academic medical centers. We found that almost half (47%) of respondents were unfamiliar with the term and/or role of a hospitalist, while the remaining patients had varied understanding of the role. Several other medical terms were frequently misunderstood (such as "NPO," "PA," and "Attending"). Ongoing efforts are needed to improve communication to ensure that hospitalized patients understand the hospitalist's role and the medical terms shared with them.
We introduce a deterministic variational formulation for training Bayesian last layer neural networks. This yields a sampling-free, single-pass model and loss that effectively improves uncertainty estimation. Our variational Bayesian last layer (VBLL) can be trained and evaluated with only quadratic complexity in last layer width, and is thus (nearly) computationally free to add to standard architectures. We experimentally investigate VBLLs, and show that they improve predictive accuracy, calibration, and out of distribution detection over baselines across both regression and classification. Finally, we investigate combining VBLL layers with variational Bayesian feature learning, yielding a lower variance collapsed variational inference method for Bayesian neural networks.
As autonomous decision-making agents move from narrow operating environments to unstructured worlds, learning systems must move from a closed-world formulation to an open-world and few-shot setting in which agents continuously learn new classes from small amounts of information. This stands in stark contrast to modern machine learning systems that are typically designed with a known set of classes and a large number of examples for each class. In this work we extend embedding-based few-shot learning algorithms to the open-world recognition setting. We combine Bayesian non-parametric class priors with an embedding-based pre-training scheme to yield a highly flexible framework which we refer to as few-shot learning for open world recognition (FLOWR). We benchmark our framework on open-world extensions of the common MiniImageNet and TieredImageNet few-shot learning datasets. Our results show, compared to prior methods, strong classification accuracy performance and up to a 12% improvement in H-measure (a measure of novel class detection) from our non-parametric open-world few-shot learning scheme.
The authors declare no conflict of interest.
Fractional gradient descent has been studied extensively, with a focus on its ability to extend traditional gradient descent methods by incorporating fractional-order derivatives. This approach allows for more flexibility in navigating complex optimization landscapes and offers advantages in certain types of problems, particularly those involving non-linearities and chaotic dynamics. Yet, the challenge of fine-tuning the fractional order parameters remains unsolved. In this work, we demonstrate that it is possible to train a neural network to predict the order of the gradient effectively.
Modern machine learning requires system designers to specify aspects of the learning pipeline, such as losses, architectures, and optimizers. Meta-learning, or learning-to-learn, instead aims to learn those aspects, and promises to unlock greater capabilities with less manual effort. One particularly ambitious goal of meta-learning is to train general-purpose in-context learning algorithms from scratch, using only black-box models with minimal inductive bias. Such a model takes in training data, and produces test-set predictions across a wide range of problems, without any explicit definition of an inference model, training loss, or optimization algorithm. In this paper we show that Transformers and other black-box models can be meta-trained to act as general-purpose in-context learners. We characterize transitions between algorithms that generalize, algorithms that memorize, and algorithms that fail to meta-train at all, induced by changes in model size, number of tasks, and meta-optimization. We further show that the capabilities of meta-trained algorithms are bottlenecked by the accessible state size (memory) determining the next prediction, unlike standard models which are thought to be bottlenecked by parameter count. Finally, we propose practical interventions such as biasing the training distribution that improve the meta-training and meta-generalization of general-purpose in-context learning algorithms.
Background Australian road safety remains a major policy concern from premature mortality and disability. Inequities exit in transport injuries with greater burden in Aboriginal and Torres Strait Islander communities. Objective Focus on Indigenous Knowledges to characterise the journey for Aboriginal and Torres Strait Islander patient journeys from road traffic injuries. Aims 1. examine prevention strategies through causal factors from road traffic crashes in South Australia (SA) and New South Wales (NSW), Aim 2. examine access to road injury compensation schemes. Methods Indigenous Governance of Project Data occurred through Aboriginal research leadership and an Aboriginal and Torres Strait Islander Governance Group. This acted to protect Indigenous Knowledges and enact Data Governance processes for Data Sovereignty. Knowledge interface methodology informed a mixed-methods approach. Participants were Aboriginal and Torres Strait Islander individuals involved in a road traffic crash above 18 years of age. Administration data over 10 years (June 2012-June 2022) from Transport for NSW and the SA Trauma Registry, were analysed using a strength-based multinomial logistic regression model. The Indigenous data collection method of yarning occurred with participants, thematic analysis identified enablers and barriers to compensation schemes. Results Multinomial logistic regression identified factors which decreased the odds of no/minor injuries: head on crashes, 10 km/h increase in speed, non-intersection crashes, yearly increases in age or driving unauthorised. Metropolitan crashes increased the odds of no/minor injuries after a crash to country areas (OR 1.34; 95% CI 1.07–1.67). No difference in injury severity was found between crashing on a curved compared to a straight stretch of road. Aboriginal and Torres Strait Islander participants reported strong impacts by their traffic injuries across physical, psychosocial financial, logistical and time domains. Limited participants accessed road traffic injury compensation, barriers included claim time-frames, and access to culturally appropriate compensation support which was further impacted by in hospital care received. Conclusions Co-designed strength-based road safety campaigns with Aboriginal and Torres Strait Islander communities needs to focus on a variety of contexts, for example country driving, or diving across the life span. Similarly, culturally safe and appropriate compensation support is urgently needed for Aboriginal and Torres Strait Islander communities.
Background While injuries can impact on children's educational achievements (with threats to their development and employment prospects), these risks are poorly quantified. This population-based longitudinal study investigated the impact of an injury-related hospital admission on Welsh children's academic performance.Methods The Secure Anonymised Information Linkage databank, 55 587 children residing in Wales from 2006 to 2016 who had an injury hospital admission (58.2% males; 16.8% born in most deprived Wales area; 80.1% one injury hospital admission) were linked to data from the Wales Electronic Cohort for Children. The primary outcome was the Core Subject Indicator reflecting educational achievement at key stages 2 (school years 3-6), 3 (school years 7-9) and 4 (school years 10-11). Covariates in models included demographic, birth, injury and school characteristics.Results Educational achievement of children was negatively associated with: pedestrian injuries (adjusted risk ratio, (95% CIs)) (0.87, (0.83 to 0.92)), cyclist (0.96, (0.94 to 0.99)), high fall (0.96, (0.94 to 0.97)), fire/flames/smoke (0.85, (0.73 to 0.99)), cutting/piercing object (0.96, (0.93 to 0.99)), intentional self-harm (0.86, (0.82 to 0.91)), minor traumatic brain injury (0.92, (0.86 to 0.99)), contusion/open wound (0.93, (0.91 to 0.95)), fracture of vertebral column (0.78, (0.64 to 0.95)), fracture of femur (0.88, (0.84 to 0.93)), internal abdomen/pelvic haemorrhage (0.82, (0.69 to 0.97)), superficial injury (0.94, (0.92 to 0.97)), young maternal age (<18 years: 0.91, (0.88 to 0.94); 19-24 years: 0.94, (0.93 to 0.96)); area based socioeconomic status (0.98, (0.97 to 0.98)); moving to a more deprived area (0.95, (0.93 to 0.97)); requiring special educational needs (0.46, (0.44 to 0.47)). Positive associations were: being female (1.04, (1.03 to 1.06)); larger pupil school sizes and maternal age 30+ years.Conclusion This study highlights the importance on a child's education of preventing injuries and implementing intervention programmes that support injured children. Greater attention is needed on equity-focused educational support and social policies addressing needs of children at risk of underachievement, including those from families experiencing poverty.
The ability of Language Models (LMs) to understand natural language makes them a powerful tool for parsing human instructions into task plans for autonomous robots. Unlike traditional planning methods that rely on domain-specific knowledge and handcrafted rules, LMs generalize from diverse data and adapt to various tasks with minimal tuning, acting as a compressed knowledge base. However, LMs in their standard form face challenges with long-horizon tasks, particularly in partially observable multi-agent settings. We propose an LM-based Long-Horizon Planner for Multi-Agent Robotics (LLaMAR), a cognitive architecture for planning that achieves state-of-the-art results in long-horizon tasks within partially observable environments. LLaMAR employs a plan-act-correct-verify framework, allowing self-correction from action execution feedback without relying on oracles or simulators. Additionally, we present MAP-THOR, a comprehensive test suite encompassing household tasks of varying complexity within the AI2-THOR environment. Experiments show that LLaMAR achieves a 30\% higher success rate than other state-of-the-art LM-based multi-agent planners in MAP-THOR and Search \& Rescue tasks. Code can be found at [https://github.com/nsidn98/LLaMAR](https://github.com/nsidn98/LLaMAR)