Recent advancements in large foundation models have remarkably enhanced our understanding of sensory information in open-world environments. In leveraging the power of foundation models, it is crucial for AI research to pivot away from excessive reductionism and toward an emphasis on systems that function as cohesive wholes. Specifically, we emphasize developing Agent AI -- an embodied system that integrates large foundation models into agent actions. The emerging field of Agent AI spans a wide range of existing embodied and agent-based multimodal interactions, including robotics, gaming, and healthcare systems, etc. In this paper, we propose a novel large action model to achieve embodied intelligent behavior, the Agent Foundation Model. On top of this idea, we discuss how agent AI exhibits remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Furthermore, we discuss the potential of Agent AI from an interdisciplinary perspective, underscoring AI cognition and consciousness within scientific discourse. We believe that those discussions serve as a basis for future research directions and encourage broader societal engagement.
The development of artificial intelligence systems is transitioning from creating static, task-specific models to dynamic, agent-based systems capable of performing well in a wide range of applications. We propose an Interactive Agent Foundation Model that uses a novel multi-task agent training paradigm for training AI agents across a wide range of domains, datasets, and tasks. Our training paradigm unifies diverse pre-training strategies, including visual masked auto-encoders, language modeling, and next-action prediction, enabling a versatile and adaptable AI framework. We demonstrate the performance of our framework across three separate domains -- Robotics, Gaming AI, and Healthcare. Our model demonstrates its ability to generate meaningful and contextually relevant outputs in each area. The strength of our approach lies in its generality, leveraging a variety of data sources such as robotics sequences, gameplay data, large-scale video datasets, and textual information for effective multimodal and multi-task learning. Our approach provides a promising avenue for developing generalist, action-taking, multimodal systems.
Standard deep reinforcement learning (DRL) aims to maximize expected reward, considering collected experiences equally in formulating a policy. This differs from human decision-making, where gains and losses are valued differently and outlying outcomes are given increased consideration. It also fails to capitalize on opportunities to improve safety and/or performance through the incorporation of distributional context. Several approaches to distributional DRL have been investigated, with one popular strategy being to evaluate the projected distribution of returns for possible actions. We propose a more direct approach whereby risk-sensitive objectives, specified in terms of the cumulative distribution function (CDF) of the distribution of full-episode rewards, are optimized. This approach allows for outcomes to be weighed based on relative quality, can be used for both continuous and discrete action spaces, and may naturally be applied in both constrained and unconstrained settings. We show how to compute an asymptotically consistent estimate of the policy gradient for a broad class of risk-sensitive objectives via sampling, subsequently incorporating variance reduction and regularization measures to facilitate effective on-policy learning. We then demonstrate that the use of moderately "pessimistic" risk profiles, which emphasize scenarios where the agent performs poorly, leads to enhanced exploration and a continual focus on addressing deficiencies. We test the approach using different risk profiles in six OpenAI Safety Gym environments, comparing to state of the art on-policy methods. Without cost constraints, we find that pessimistic risk profiles can be used to reduce cost while improving total reward accumulation. With cost constraints, they are seen to provide higher positive rewards than risk-neutral approaches at the prescribed allowable cost.
Multi-objective reinforcement learning (MORL) has been proposed to learn control policies over multiple competing objectives with each possible preference over returns. However, current MORL algorithms fail to account for distributional preferences over the multi-variate returns, which are particularly important in real-world scenarios such as autonomous driving. To address this issue, we extend the concept of Pareto-optimality in MORL into distributional Pareto-optimality, which captures the optimality of return distributions, rather than the expectations. Our proposed method, called Distributional Pareto-Optimal Multi-Objective Reinforcement Learning~(DPMORL), is capable of learning distributional Pareto-optimal policies that balance multiple objectives while considering the return uncertainty. We evaluated our method on several benchmark problems and demonstrated its effectiveness in discovering distributional Pareto-optimal policies and satisfying diverse distributional preferences compared to existing MORL methods.
Advances in reinforcement learning (RL) have resulted in recent breakthroughs in the application of artificial intelligence (AI) across many different domains. An emerging landscape of development environments is making powerful RL techniques more accessible for a growing community of researchers. However, most existing frameworks do not directly address the problem of learning in complex operating environments, such as dense urban settings or defense-related scenarios, that incorporate distributed, heterogeneous teams of agents. To help enable AI research for this important class of applications, we introduce the AI Arena: a scalable framework with flexible abstractions for associating agents with policies and policies with learning algorithms. Our results highlight the strengths of our approach, illustrate the importance of curriculum design, and measure the impact of multi-agent learning paradigms on the emergence of cooperation.
Advances in reinforcement learning (RL) have resulted in recent breakthroughs in the application of artificial intelligence (AI) across many different domains. An emerging landscape of development environments is making powerful RL techniques more accessible for a growing community of researchers. However, most existing frameworks do not directly address the problem of learning in complex operating environments, such as dense urban settings or defense-related scenarios, that incorporate distributed, heterogeneous teams of agents. To help enable AI research for this important class of applications, we introduce the AI Arena: a scalable framework with flexible abstractions for distributed multi-agent reinforcement learning. The AI Arena extends the OpenAI Gym interface to allow greater flexibility in learning control policies across multiple agents with heterogeneous learning strategies and localized views of the environment. To illustrate the utility of our framework, we present experimental results that demonstrate performance gains due to a distributed multi-agent learning approach over commonly-used RL techniques in several different learning environments.
Recent breakthroughs and rapid progress in AI will impact, if not transform, every mission. JHU/APL developed an AI Technology Roadmap to guide the Laboratory’s contributions to the critical challenges the nation will face developing and implementing intelligent systems for these missions over the coming decades. We began this exercise by describing a series of envisioned futures for intelligent systems across sea, land, air, space and information, and examined them to identify the common AI technology vectors needed to achieve each vision: (1) Autonomous Perception, describing the path to intelligent systems that perceive in the context of the extreme uncertainty and complexity of the real world; (2) Superhuman Decision-Making and Autonomous Action, to realize the potential for intelligent systems to reason over more information than any team of analysts or operators and act in ways systems under manned control cannot; (3) Human-Machine Teaming at the Speed of Thought, to ensure humans can stay involved at speed and scale; and of particular importance for national security applications, (4) Safe and Assured Operation, so these systems can be trusted to stay true to commander’s intent, in adversarial and sensitive contexts. Each technology vector is aligned with a targeted goal, and with each goal we provide a roadmap in the form of near-, mid-, and long-term AI advances critical to reaching the goal. This paper describes the JHU/APL AI Technology Roadmap and presents key examples of recent progress and forward-looking research and exploratory development along each vector.
Intelligent systems are already having a remarkable impact on society. Future advancements could have an even greater impact by empowering people through human-machine teaming, addressing challenges with vast geographic scales, and accelerating interstellar discovery. Creating intelligent systems that can be trusted to operate autonomously is a grand challenge for humanity. In this article, we explore potential futures for trustworthy autonomous systems, identify some of the significant challenges, and illustrate potential pathways by describing developments underway at the Johns Hopkins University Applied Physics Laboratory (APL).
AbstractThere is currently a global arms race for the development of artificial intelligence (AI) and unmanned robotic systems that are empowered by AI (AI-robots). This paper examines the current use of AI-robots on the battlefield and offers a framework for understanding AI and AI-robots. It examines the limitations and risks of AI-robots on the battlefield and posits the future direction of battlefield AI-robots. It then presents research performed at the Johns Hopkins University Applied Physics Laboratory (JHU/APL) related to the development, testing, and control of AI-robots, as well as JHU/APL work on human trust of autonomy and developing self-regulating and ethical robotic systems. Finally, it examines multiple possible future paths for the relationship between humans and AI-robots.
The Active Sensing Testbed (AST) is a novel framework for research in machine perception and world-view reasoning. The AST supports exploratory development of perception systems that can build internal models of the world by combining multi-view and multi-modal analytics, utilize these models to form hypotheses about a scene, and potentially take action to fill in gaps in knowledge or make predictions about future world states. As a modular software framework, the AST is intended to lower the barrier to entry for researchers and developers in applying state-of-the-art computer vision techniques to real-world problems.
Highly automated systems are becoming omnipresent. They range in function from self-driving vehicles to advanced medical diagnostics and afford many benefits. However, there are assurance challenges that have become increasingly visible in high-profile crashes and incidents. Governance of such systems is critical to garner widespread public trust. Governance principles have been previously proposed offering aspirational guidance to automated system developers; however, their implementation is often impractical given the excessive costs and processes required to enact and then enforce the principles. This Perspective, authored by an international and multidisciplinary team across government organizations, industry and academia, proposes a mechanism to drive widespread assurance of highly automated systems: independent audit. As proposed, independent audit of AI systems would embody three ‘AAA’ governance principles of prospective risk Assessments, operation Audit trails and system Adherence to jurisdictional requirements. Independent audit of AI systems serves as a pragmatic approach to an otherwise burdensome and unenforceable assurance challenge. As highly automated systems become pervasive in society, enforceable governance principles are needed to ensure safe deployment. This Perspective proposes a pragmatic approach where independent audit of AI systems is central. The framework would embody three AAA governance principles: prospective risk Assessments, operation Audit trails and system Adherence to jurisdictional requirements.
We propose an ensemble approach for multi-target binary classification, where the target class breaks down into a disparate set of pre-defined target-types. The system goal is to maximize the probability of alerting on targets from any type while excluding background clutter. The agent-classifiers that make up the ensemble are binary classifiers trained to classify between one of the target-types vs. clutter. The agent ensemble approach offers several benefits for multi-target classification including straightforward in-situ tuning of the ensemble to drift in the target population and the ability to give an indication to a human operator of which target-type causes an alert. We propose a combination strategy that sums weighted likelihood ratios of the individual agent-classifiers, where the likelihood ratio is between the target-type for the agent vs. clutter. We show that this combination strategy is optimal under a conditionally non-discriminative assumption. We compare this combiner to the common strategy of selecting the maximum of the normalized agent-scores as the combiner score. We show experimentally that the proposed combiner gives excellent performance on the multi-target binary classification problems of pin-less verification of human faces and vehicle classification using acoustic signatures.
The challenge of establishing assurance in autonomy is rapidly attracting increasing interest in the industry, government, and academia. Autonomy is a broad and expansive capability that enables systems to behave without direct control by a human operator. To that end, it is expected to be present in a wide variety of systems and applications. A vast range of industrial sectors, including (but by no means limited to) defense, mobility, health care, manufacturing, and civilian infrastructure, are embracing the opportunities in autonomy yet face the similar barriers toward establishing the necessary level of assurance sooner or later. Numerous government agencies are poised to tackle the challenges in assured autonomy. Given the already immense interest and investment in autonomy, a series of workshops on Assured Autonomy was convened to facilitate dialogs and increase awareness among the stakeholders in the academia, industry, and government. This series of three workshops aimed to help create a unified understanding of the goals for assured autonomy, the research trends and needs, and a strategy that will facilitate sustained progress in autonomy. The first workshop, held in October 2019, focused on current and anticipated challenges and problems in assuring autonomous systems within and across applications and sectors. The second workshop held in February 2020, focused on existing capabilities, current research, and research trends that could address the challenges and problems identified in workshop. The third event was dedicated to a discussion of a draft of the major findings from the previous two workshops and the recommendations.
The ability to create artificial intelligence (AI) capable of performing complex tasks is rapidly outpacing our ability to ensure the safe and assured operation of AI-enabled systems. Fortunately, a landscape of AI safety research is emerging in response to this asymmetry and yet there is a long way to go. In particular, recent simulation environments created to illustrate AI safety risks are relatively simple or narrowly-focused on a particular issue. Hence, we see a critical need for AI safety research environments that abstract essential aspects of complex real-world applications. In this work, we introduce the AI safety TanksWorld as an environment for AI safety research with three essential aspects: competing performance objectives, human-machine teaming, and multi-agent competition. The AI safety TanksWorld aims to accelerate the advancement of safe multi-agent decision-making algorithms by providing a software framework to support competitions with both system performance and safety objectives. As a work in progress, this paper introduces our research objectives and learning environment with reference code and baseline performance metrics to follow in a future work.
Reconnaissance blind chess (RBC) is a chess variant in which a player cannot see her opponent's pieces but can learn about them through private, explicit sensing actions. The game presents numerous research challenges, and was the focus of a competition held in conjunction with of the 2019 Conference on Neural Information Processing Systems (NeurIPS). The 22 bots that played in the tournament leveraged a diverse set of algorithms, including variations of multi-state tracking, piece-wise probability estimation, Gibbs sampling, bandit algorithms, tree search, counterfactual regret minimization (CFR), deep learning, and others. None of the algorithms of which we are aware converges to an optimal strategy. Top algorithms generally incorporated sensing strategies that successfully minimized uncertainty (as measured in the number of possible opponent states). The top two approaches reduced this raw uncertainty metric less than some others. Successful strategies sometimes defied conventional wisdom in chess, as evidenced by deviations between win rate and aggregate move strength as assessed by the leading available chess engine.
This paper provides a complexity analysis for the game of reconnaissance blind chess (RBC), a recently-introduced variant of chess where each player does not know the positions of the opponent's pieces a priori but may reveal a subset of them through chosen, private sensing actions. In contrast to many commonly studied imperfect information games like poker, an RBC player does not know what the opponent knows or has chosen to learn, exponentially expanding the size of the game's information sets (i.e., the number of possible game states that are consistent with what a player has observed). Effective RBC sensing and moving strategies must account for the uncertainty of both players, an essential element of many real-world decision-making problems. Here we evaluate RBC from a game theoretic perspective, tracking the proliferation of information sets from the perspective of selected canonical bot players in tournament play. We show that, even for effective sensing strategies, the game sizes of RBC compare to those of Go while the average size of a player's information set throughout an RBC game is much greater than that of a player in Heads-up Limit Hold 'Em. We compare these measures of complexity among different playing algorithms and provide cursory assessments of the various sensing and moving strategies.
This paper considers online classification problems where each object to be classified consists of a sequence of measurements, termed here a track. We present an approach that combines ideas from sequential hypothesis testing with those from conformal prediction to address track level outliers - entire measurement sequences that are novel relative to the statistical model. We show with analysis and empirical results that this approach preserves the optimal performance of the underlying sequential hypothesis testing when outliers are absent and provides an error rate guarantee in the presence of contamination by novel tracks.
The any-combiner is a classifier combination approach for target classification problems in which the target class can be naturally decomposed into multiple subclasses. This kind of classification problem can often occur in sensor-based system applications, such as biometric user verification, biosurveillance or underwater mine detection, in which the system goal is to identify a test exemplar as belonging to a category of objects of interest to the exclusion of all other exemplars (clutter). We propose an approach to the target classification problem in which an ensemble of classifier agents are trained to distinguish individual target subclasses from clutter. The any-combiner is then trained by optimizing the multi-agent ensemble for maximum recognition performance across all target subclasses over a range of acceptable operating points. Once deployed, the any-combiner classifies a test example as a target if any of the agents indicates a true positive classification for its target subclass. Experiments show that the any-combiner yields excellent performance on the tasks of biometric verification using face images and underwater object classification using acoustic features.
Conformal prediction is a relatively recent approach to classification that offers a theoretical framework for generating predictions with precise levels of confidence. For each new object encountered, a conformal predictor outputs a set of class labels that contains the true label with probability at least 1 - ∈, where ∈ is a user-specified error rate. The ability to predict with confidence can be extremely useful, but in many real-world applications unambiguous predictions consisting of a single class label are preferred. Hence it is desirable to design conformal predictors to maximize the rate of singleton predictions, termed the efficiency of the predictor. In this paper we derive a novel criterion for maximizing efficiency for a certain class of conformal predictors, show how concepts from local distance metric learning can provide a useful bound for maximizing this criterion, and demonstrate efficiency gains on real-world datasets.
In classification problems, the goal is often to maximize the probability of a true positive while satisfying a system design constraint on the number of false alarms. Current classifiers, which are based on minimizing classification error, achieve this goal indirectly by either shifting the bias of the resulting solution or via tuning of an asymmetry parameter. While these approaches may be used to satisfy a false alarm constraint, they may produce classifiers with suboptimal target classification performance at the desired operating point. We propose a large margin classifier that directly maximizes true positive classifications at a desired false alarm rate via the inclusion of an optimization constraint on the estimated probability of false alarms. Unlike existing false alarm constrained classifiers (also called Neyman-Pearson or minimax classifiers) our approach allows the use of an arbitrary loss function and is applicable to nonlinear datasets that exhibit a high degree of class imbalance. Our technique achieves a robust solution in small-sample training scenarios via the use of tail-estimation techniques to predict the probability of false alarms when estimates based on training error may be unreliable.