Generative models have emerged as a powerful paradigm for AI planning, yet their performance remains constrained by training data distribution. One approach is to improve generated solutions during inference by scaling test-time compute. A more efficient alternative is to optimize the inferential process itself. In this paper, we show that a modified version of a classical Open-Closed List (OCL) search provides just such an efficient inferential procedure. Our algorithm synergizes two learned components: a generative model that performs fast rollouts from specific reasoning paths and a value model that manages which of many possible reasoning lines to follow. We present novel contributions in exploration control and how learned models are integrated within the OCL framework. Experimental evaluation across multiple combinatorial planning domains shows that our approach consistently outperforms baseline search algorithms in both computational efficiency and solution quality.
AbstractA body of work proposes that social-norm change can be explained in terms of game theory. These game theoretic models, however, don't fully account for how and why utterances are used to change social norms. This paper describes the problem and some of the solution elements. There are three existing, relevant, game-based models. The first is a game theoretic model of social norm change (Bicchieri, 2005, 2016). This accounts for how individuals make decisions to adhere to or violate norms, based on empirical expectations of how others will behave. The second is the idea of a conversational game (Lewis, 1979) and its extensions. This posits that speech acts are accommodated in a conversation to make what is said correct play. This feature can explain how some speech acts, such as slurring utterances, change the dynamics of a conversation. The third is a theory of pragmatic inference, known as Rational Speech Act theory (Goodman and Frank, 2016). This is a computational theory of pragmatics, of how listeners interpret utterances and how speakers construct utterances that can be understood. This paper proposes, without setting out the full formal model, that elements of these three theories need to be incorporated together into a game theoretic model of how utterances change long-term social norms.
Unfairness emerges in bargaining games under a variety of conditions. Two effects increase the probability that a simulated society converges to unfairness: (i) the Red King effect produces advantage in favor of members of the larger group; (ii) the Bargaining Power effect produces advantage in favor of the group with more powerful individuals. This paper shows how these effects are modulated by relationships between unobservable and observable traits (signals). We investigate three models. First, when an agent can choose signals, we find a relationship between unfairness and distinctive signaling. Second, we introduce stochastic signaling, in which signals are somewhat (de-)correlated with an underlying trait. Third, we introduce adaptive tolerance, in which agents create a signal using a partition of the underlying trait space. It is shown computationally that, in each case, there is a fall in the proportion of societies converging on unequal resource distribution. Analysis supports the idea that the underlying phenomenon linking all these is the mutual information between the signal and the unobservable trait.
Dexterous grasping of a novel object given a single view is an open problem. This paper makes several contributions to its solution. First, we present a simulator for generating and testing dexterous grasps. Second, we present a dataset, generated by this simulator, of 2.4 million simulated dexterous grasps of variations of 294 base objects drawn from 20 categories. Third, we combine an existing approach to learn a grasp generation model with three different learned evaluative models employing ResNet-50 or VGG16 as their visual backbone. Fourth, we train, and evaluate 17 variants of the resulting generative-evaluative architectures on the simulated dataset, showing improvement from 69.53% grasp success rate to 90.49%. Fifth, we present a real robot implementation and evaluate the four most promising variants, executing 196 real robot grasps in total. We show that our best architectural variant achieves a grasp success rate of 87.8% on real novel objects seen from a single view, improving on a baseline of 57.1%. Finally, we explore the inner workings of our best evaluative model and perform an extensive analysis of its results on the simulated dataset.
Consider a mobile robot exploring an office building with the aim of observing as much human activity as possible over several days. It must learn where and when people are to be found, count the observed activities, and revisit popular places at the right time. In this paper we present a series of Bayesian estimators for the levels of human activity that improve on simple counting. We then show how these estimators can be used to drive efficient exploration for human activities. The estimators arise from modelling the human activity counts as a partially observable Poisson process (POPP). This paper presents novel extensions to POPP for the following cases: (i) the robot’s sensors are correlated, (ii) the robot’s sensor model, itself built from data, is also unreliable, (iii) both are combined. It also combines the resulting Bayesian estimators with a simple, but effective solution to the exploration-exploitation trade-off faced by the robot in a real deployment. A series of 15 day robot deployments show how our approach boosts the number of human activities observed by 70% relative to a baseline and produces more accurate estimates of the level of human activity in each place and time.
Manipulation of deformable objects has given rise to an important set of open problems in the field of robotics. Application areas include robotic surgery, household robotics, manufacturing, logistics, and agriculture, to name a few. Related research problems span modeling and estimation of an object's shape, estimation of an object's material properties, such as elasticity and plasticity, object tracking and state estimation during manipulation, and manipulation planning and control. In this survey article, we start by providing a tutorial on foundational aspects of models of shape and shape dynamics. We then use this as the basis for a review of existing work on learning and estimation of these models and on motion planning and control to achieve desired deformations. We also discuss potential future lines of work.
Investigation of scenarios to drive ontology-formation, perceptual development, development of action, and growth of an information processing architecture in an ’altricial’ robot in a 3-D environment. An artificial 3-D domain with a number of interesting and challenging features is proposed as a domain for robotics research in perception, learning, action, planning, ontology formation, and perhaps later on natural language communication (after the robot has something to communicate about).
The daily working hours of mobile robots are limited primarily by battery life. Most systems use a combination of thresholds and fixed periods to decide when to charge. This produces charging behaviour that ignores high-value tasks that must be performed within time-windows or by deadlines. Instead the robot should schedule charging adaptively, taking into account the times of day when it is expected to be given more valuable tasks to perform. This paper proposes an approach that exploits the fact that, during long-term deployments, the robot can learn when it is most probable that valuable tasks are added to the system, enabling it to schedule charging at times that are expected to be less busy. We pose the problem of scheduling battery charging as a multi-objective sequential decision making problem over a time-dependent Markov decision process model of expected task rewards and battery dynamics. We evaluate the scalability and solution quality of our multi-objective scheduler, and compare it with a typical rule-based approach. Empirical results show that our approach enables more flexible and efficient robot behaviour, which takes into account both the value of current available tasks and the predicted value of future tasks to decide whether to charge at a given time.
In the last years, museums around the world began using museum guide robots as an alternative to audio guides. Thus, it is important to identify how these new museum guides can optimally interact with visitors and whether they are more efficient that conventional audio guides. In this paper, we present the results from an experimental study which compared conventional audio guides with robots with different 'personality'. These results demonstrate that people remember significantly more information when they are guided by a cheerful robot than when their guide is a serious one or a conventional audio system. Importantly, we also introduce the idea of two collaborative tour guide robots. This idea was inspired by cognitive studies showing that people remember more when they receive information from two different human speakers. Interestingly, our participants also liked more our collaborative robots, than any version of a single robot and/or audio system.
The daily working hours of long-life mobile robots are limited primarily by battery life. Most systems use a combination of hard thresholds and fixed periods to decide when to charge. This produces charging behaviour that ignores high-value tasks that must be performed within time-windows or by deadlines. Instead the robot should schedule charging adaptively, taking into account the times of day when it is expected to be given more valuable tasks to perform. This paper proposes an approach that exploits the fact that, during long-term deployments, the robot can learn when it is most probable that valuable tasks are added to the system, thus it can plan to charge on times that are expected to be less busy. We pose the problem of scheduling battery charging as a multi-objective sequential decision making problem over a time-dependent Markov decision process model of expected task rewards and battery behaviour. We compare a typical rule-based approach to our multi-objective scheduler and show that our approach enables for more flexible and efficient robot behaviour, which takes into account both the value of current available tasks and the predicted value of future tasks to decide whether to charge at a given time.
During the initial trials of a manipulation task, humans tend to keep their arms stiff in order to reduce the effects of any unforeseen disturbances. After a few repetitions, humans perform the task accurately with much lower stiffness. Research in human motor control indicates that this behavior is supported by learning and continuously revising internal models of the manipulation task. These internal models predict future states of the task, anticipate necessary control actions, and adapt impedance quickly to match task requirements. Drawing inspiration from these findings, we propose a framework for online learning of a time-independent forward model of a manipulation task from a small number of examples. The measured inaccuracies in the predictions of this model dynamically update the forward model and modify the impedance parameters of a feedback controller during task execution. Furthermore, our framework includes a hybrid force-motion controller that provides compliance in particular directions while adapting the impedance in other directions. These capabilities are evaluated on continuous contact tasks such as pulling non-linear springs, polishing a board, and stirring porridge.
In the coming years tour guide robots will be widely used in museums and exhibitions. Therefore, it is important to identify how these new museum guides can optimally interact with visitors. In this paper, we introduce the idea of two collaborative tour guide robots. We have been inspired by evidence from cognitive studies stating that people remember more when they receive information from two different human speakers. Our collaborative tour guides were benchmarked against single robot guides. Our study initially proved, through real-world experiments, previous proposals stating that the personality of the robot affects the human learning process; our results demonstrate that people remember significantly more information when they are guided by a cheerful robot than when their guide is a serious one. Moreover, another important outcome of our study is that our visitors tend to like more our collaborative robots, than any referenced single robot, as demonstrated by the higher scores in the aesthetic-related questions. Hence our results suggest that a cheerful robot is more suitable for learning purposes while two robots are more suitable for entertainment purposes.
This article describes REBA, a knowledge representation and reasoning architecture for robots that is based on tightly-coupled transition diagrams of the domain at two different levels of granularity. An action language is extended to support non-boolean fluents and non-deterministic causal laws, and used to describe the domain's transition diagrams, with the fine-resolution transition diagram being defined as a refinement of the coarse-resolution transition diagram. The coarse-resolution system description, and a history that includes prioritized defaults, are translated into an Answer Set Prolog (ASP) program. For any given goal, inference in the ASP program provides a plan of abstract actions. To implement each such abstract action, the robot automatically zooms to the part of the fine-resolution transition diagram relevant to this action. The zoomed fine-resolution system description, and a probabilistic representation of the uncertainty in sensing and actuation, are used to construct a partially observable Markov decision process (POMDP). The policy obtained by solving the POMDP is invoked repeatedly to implement the abstract action as a sequence of concrete actions. The fine-resolution outcomes of executing these concrete actions are used to infer coarse-resolution outcomes that are added to the coarse-resolution history and used for subsequent coarse-resolution reasoning. The architecture thus combines the complementary strengths of declarative programming and probabilistic graphical models to represent and reason with non-monotonic logic-based and probabilistic descriptions of uncertainty and incomplete domain knowledge. In addition, we describe a general methodology for the design of software components of a robot based on these knowledge representation and reasoning tools, and provide a path for proving the correctness of these components. The architecture is evaluated in simulation and on a mobile robot finding and moving target objects to desired locations in indoor domains, to show that the architecture supports reliable and efficient reasoning with violation of defaults, noisy observations and unreliable actions, in complex domains.
Belief space planning is a viable alternative to formalise partially observable control problems and, in the recent years, its application to robot manipulation problems has grown. However, this planning approach was tried successfully only on simplified control problems. In this paper, we apply belief space planning to the problem of planning dexterous reach-to-grasp trajectories under object pose uncertainty. In our framework, the robot perceives the object to be grasped on-the-fly as a point cloud and compute a full 6D, non-Gaussian distribution over the object's pose (our belief space). The system has no limitations on the geometry of the object, i.e., non-convex objects can be represented, nor assumes that the point cloud is a complete representation of the object. A plan in the belief space is then created to reach and grasp the object, such that the information value of expected contacts along the trajectory is maximised to compensate for the pose uncertainty. If an unexpected contact occurs when performing the action, such information is used to refine the pose distribution and triggers a re-planning. Experimental results show that our planner (IR3ne) improves grasp reliability and compensates for the pose uncertainty such that it doubles the proportion of grasps that succeed on a first attempt.
This paper concerns the problem of how to learn to grasp dexterously, so as to be able to then grasp novel objects seen only from a single viewpoint. Recently, progress has been made in data-efficient learning of generative grasp models that transfer well to novel objects. These generative grasp models are learned from demonstration (LfD). One weakness is that, as this paper shall show, grasp transfer under challenging single-view conditions is unreliable. Second, the number of generative model elements increases linearly in the number of training examples. This, in turn, limits the potential of these generative models for generalization and continual improvement. In this paper, it is shown how to address these problems. Several technical contributions are made: (i) a view-based model of a grasp; (ii) a method for combining and compressing multiple grasp models; (iii) a new way of evaluating contacts that is used both to generate and to score grasps. Together, these improve grasp performance and reduce the number of models learned. These advances, in turn, allow the introduction of autonomous training, in which the robot learns from self-generated grasps. Evaluation on a challenging test set shows that, with innovations (i)–(iii) deployed, grasp transfer success increases from 55.1% to 81.6%. By adding autonomous training this rises to 87.8%. These differences are statistically significant. In total, across all experiments, 539 test grasps were executed on real objects.
The hype about sensorimotor learning is currently reaching high fever, thanks to the latest advancement in deep learning. In this paper, we present an open-source framework for collecting large-scale, time-synchronised synthetic data from highly disparate sensory modalities, such as audio, video, and proprioception, for learning robot manipulation tasks. We demonstrate the learning of non-linear sensorimotor mappings for a humanoid drumming robot that generates novel motion sequences from desired audio data using cross-modal correspondences. We evaluate our system through the quality of its cross-modal retrieval, for generating suitable motion sequences to match desired unseen audio or video sequences.
We present a parametric formulation for learning generative models for grasp synthesis from a demonstration. We cast new light on this family of approaches, proposing a parametric formulation for grasp synthesis that is computationally faster compared to related work and indicates better grasp success rate performance in simulated experiments, showing a gain of at least 10% success rate (p < 0.05) in all the tested conditions. The proposed implementation is also able to incorporate arbitrary constraints for grasp ranking that may include task-specific constraints. Results are reported followed by a brief discussion on the merits of the proposed methods noted so far.
The original article [1] contained a minor error in the following sentence in the Discussion.
A challenge in robot manipulation is how to learn tasks efficiently. We combine learning from demonstration with data efficient exploration guided by Bayesian optimisation. We use dynamic movement primitives to encode manipulation actions. These permit temporal and spatial scaling of the demonstrated trajectory. We demonstrate the effectiveness of direct policy search for the scaling parameters with Bayesian optimisation (BO). We evaluate BO against random search on two real robot tasks: a ‘throw object to target’ task, and a ‘flip object to target’ task. We are able to obtain good policy parameters despite large amounts of noise and a weak relationship between the parameters and the policy score.
Geert-Jan M. Kruijff合作论文数German Research Center for Artificial Intelligence (DFKI GmbH);Language Technology group 8
Henrik Jacobsson合作论文数Language Technology Lab4