We present a Bayesian approach to learning flexible safety constraints and subsequently verifying whether plans satisfy these constraints. Our approach, called the Safety Constraint Learner/Checker (SCLC), infers safety constraints from a single expert demonstration trace and minimal background knowledge, and applies these constraints to the solutions proposed by multiple planning agents in an integrated and heterogeneous ensemble. The SCLC calculates how much to blame plan fragments (partial solutions) generated by the individual planning agents. This information is used when composing these fragments into a final overall plan. In particular, fragments whose safety violations exceed a threshold are rejected. This facilitates the generation of safe plans. We have integrated the SCLC within the Generalized Integrated Learning Architecture, which was designed for Defense Advanced Research Projects Agency (DARPA)’s Integrated Learning (IL) program. The main goal of the IL program is to promote the development and success of sophisticated systems that learn to solve challenging real‐world problems based on a simple demonstration by a human expert and exiguous domain knowledge. We present experimental results showing the advantages of the SCLC on two multiagent problem‐solving tasks that were benchmark applications in DARPA’s IL program.
We present a novel ensemble architecture for learning problem-solving techniques from a very small number of expert solutions and demonstrate its effectiveness in a complex real-world domain. The key feature of our “Generalized Integrated Learning Architecture” (GILA) is a set of heterogeneous independent learning and reasoning (ILR) components, coordinated by a central meta-reasoning executive (MRE). The ILRs are weakly coupled in the sense that all coordination during learning and performance happens through the MRE. Each ILR learns independently from a small number of expert demonstrations of a complex task. During performance, each ILR proposes partial solutions to subproblems posed by the MRE, which are then selected from and pieced together by the MRE to produce a complete solution. The heterogeneity of the learner-reasoners allows both learning and problem solving to be more effective because their abilities and biases are complementary and synergistic. We describe the application of this novel learning and problem solving architecture to the domain of airspace management, where multiple requests for the use of airspaces need to be deconflicted, reconciled, and managed automatically. Formal evaluations show that our system performs as well as or better than humans after learning from the same training data. Furthermore, GILA outperforms any individual ILR run in isolation, thus demonstrating the power of the ensemble architecture for learning and problem solving.
We propose an intelligent tutoring system that constructs a curriculum of hints and problems in order to teach a student skills with a rich dependency structure. We provide a template for building a multi-layered Dynamic Bayes Net to model this problem and describe how to learn the parameters of the model from data. Planning with the DBN then produces a teaching policy for the given domain. We test this end-to-end curriculum design system in two human-subject studies in the areas of finite field arithmetic and artificial language and show this method performs on par with hand-tuned expert policies.
Students interacted with an intelligent tutoring system to learn grammatical rules for an artificial language. Six tutoring policies were explored. One, based on a Dynamic Bayes' Network model of skills, was learned from the performance of previous students. Overall, this policy and other intelligent policies outperformed random policies. Some policies allowed students to choose one of three problems to work on, while others presented a single problem at each iteration. The benefit of choice was not apparent in group statistics; however, there was a strong interaction with gender. Overall, women learned less than men, but they learned different amounts in the choice and no choice conditions, whereas men seemed unaffected by choice. We explore reasons for these interactions between gender, choice and learning.
The most common versions of particle swarm optimization PSO algorithms are rotationally variant. It has also been pointed out that PSO algorithms can concentrate particles along paths parallel to the coordinate axes. In this paper, the authors explicitly connect these two observations by showing that the rotational variance is related to the concentration along lines parallel to the coordinate axes. Based on this explicit connection, the authors create fitness functions that are easy or hard for PSO to solve, depending on the rotation of the function.
In this paper we describe the application of a novel learning and problem solving architecture to the domain of airspace management, where multiple requests for the use of airspace need to be reconciled and managed automatically. The key feature of our Generalized Integrated Learning Architecture (GILA) is a set of integrated learning and reasoning (ILR) systems coordinated by a central meta-reasoning executive (MRE). Each ILR learns independently from the same training example and contributes to problem-solving in concert with other ILRs as directed by the MRE. Formal evaluations show that our system performs as well as or better than humans after learning from the same training data. Further, GILA outperforms any individual ILR run in isolation, thus demonstrating the power of the ensemble architecture for learning and problem solving.
We present a Bayesian approach to learning exible safety constraints and subsequently verifying whether plans satisfy these constraints. Our approach, called the Safety Constraint Learner/Checker (SCLC), is embedded within the Generalized Integrated Learning Architecture (GILA), which is an integrated, heterogeneous, multi-agent ensemble architecture designed for learning complex problem solving techniques from demonstration by human experts. The SCLC infers safety constraints from a single expert demonstration trace, and applies these constraints to the solutions proposed by the agents in the ensemble. Blame for constraint violations is then transmitted to the individual learning/planning/reasoning agents, thereby facilitating new problem-solving episodes. We discuss the advantages of the SCLC and demonstrate empirical results on an Airspace Planning and Deconiction Task, which was a benchmark application in the DARPA Integrated Learning Program.
In this paper we describe the application of a novel learning and problem solving architecture to the domain of airspace management, where multiple requests for the use of airspace need to be reconciled and managed automatically. The key feature of our “Generalized Integrated Learning Architecture” (GILA) is a set of integrated learning and reasoning (ILR) systems coordinated by a central meta-reasoning executive (MRE). Each ILR learns independently from the same training example and contributes to problem-solving in concert with other ILRs as directed by the MRE. Formal evaluations show that our system performs as well as or better than humans after learning from the same training data. Further, GILA outperforms any individual ILR run in isolation, thus demonstrating the power of the ensemble architecture for learning and problem solving.
This paper focusses on an automated learning/reasoning system for inferring mission safety in airspace operations. We describe a simplified version of a realistic airspace operations scenario that inspired our work. Our domain knowledge about airspace operations and mission safety is expressed qualitatively. We define and describe a way to construct explanations of missions that are translated into a Neural Network representation which is fit to the data and scored. We describe a pruning algorithm to select a greedy-best explanation structure of mission safety. Our experimental evaluation demonstrates the effectiveness of using domain knowledge in learning compared to the standard hidden-layered Artificial Neural Networks. 1 Problem and Significance One of the most important factors in military decisionmaking is safety, i.e., whether a high-valued asset participating in a mission is safe or not, given the enemy threats and our capabilities in that mission. For example, consider the following fictitious but very realistic scenario, simplified substantially here: Intelligence confirms that the militant power in the recently captured Area 6 have escaped to neighboring countries such as Areas 1 and 7 with large numbers of weapons, including SCUD missiles and hand held surface-to-air missiles (SAMs). Despite this intelligence, the exact location of these missiles is still unknown. The Joint Force Air Component Commander (JFACC) directs a large proportion of airpower from Areas 3, 5, and 6 to be devoted to finding and eliminating the threat of SCUD University of Illinois at Urban, Computer Science Department, Siebel Center for Computer Science, 201 N Goodwin Ave, Urbana, IL 61801-2302, USA. Email: {levine,mrebl}@cs.uiuc.edu University of Maryland, Institute for Advanced Computer Studies, College Park, MD 20742, USA. Email: ukuter@cs.umd.edu Analytic Services Inc., 2900 South Quincy St., Arlington, Va 22206, USA. Email: Kevin.VanSloten@anser.org University of Wyoming, Department of Computer Science, 1000 E. University Avenue, Laramie, WY 82071, USA. {anton,derekg,dspears}@cs.uwyo.edu Figure 1: An illustration of the military planning scenario described in the text. The map is based on a screenshot of the map of the well-known board game Risk [Wikipedia, 2009]. missiles. Assets involved in this mission include F-15E Strike Eagles, F/A-18 Hornets, JSTARS, E-2C Hawkeye, E-3A AWACS, and KC-135 tankers. The Rules of Engagement (ROE) require that (1) the coalition forces are not authorized to engage or otherwise cross the border into any neighboring country, and (2) a visual confirmation of any SCUD launched is established before any action – thus, a SCUD hunting mission will only occur during daylight hours or under three-quarter or more moon clear nights. The air defenses of the potential host countries that would interfere with any kind of SCUD hunting mission include multiple shoulder launched surface-to-air missiles, along with the TOR and Hawk (export variant) missile systems. Figure 1 shows some Hawk and TOR missile ranges, and some suspected launcher sights depicted by yellow stars. Hawk missile sights are shown in red, and
This paper presents a Bayesian approach to learning flexible safety constraints in a coordinated, multi-planner ensemble, along with stochastic and active experimentation approaches for assigning degrees of blame when these constraints are violated. The blame is subsequently translated and conveyed to planners, for the purpose of improved overall system performance. To illustrate the advantages of our framework, we provide and discuss examples on a real test application for Airspace Control Order (ACO) planning and deconfliction, which is a benchmark application in the DARPA Integrated
A key challenge of automated planning, including “safe planning,” is the requirement of a domain expert to provide the background knowledge, including some set of safety constraints. To alleviate the infeasibility of acquiring complete and correct knowledge from human experts in many complex, real-world domains, this paper investigates a technique for automated extraction of safety constraints by observing a user demonstration trace. In particular, we describe a new framework based on maximum likelihood learning for generating constraints on the concepts and properties in a domain ontology for a planning domain. Then, we describe a generalization of this framework that involves Bayesian learning of such constraints. To illustrate the advantages of our framework, we provide and discuss examples on a real test application for Airspace Control Order (ACO) planning, a benchmark application in the DARPA Integrated Learning Program.
This paper introduces a novel framework for designing multi-agent systems, called “Distributed Agent Evolution with Dynamic Adaptation to Local Unexpected Scenarios” (DAEDALUS). Traditional approaches to designing multi-agent systems are offline (in simulation), and assume the presence of a global observer. In the online (real world), there may be no global observer, performance feedback may be delayed or perturbed by noise, agents may only interact with their local neighbors, and only a subset of agents may experience any form of performance feedback. Under these circumstances, it is much more difficult to design multi-agent systems. DAEDALUS is designed to address these issues, by mimicking more closely the actual dynamics of populations of agents moving and interacting in a task environment. We use two case studies to illustrate the feasibility of this approach.