As we consider entrusting large language models (LLMs) with key societal and decision-making roles, measuring their alignment with human cognition becomes critical. This requires methods that can assess how these systems represent information and facilitate comparisons with human understanding across diverse tasks. To meet this need, we adapted representational similarity analysis (RSA), using pairwise ratings to help quantify alignment between AIs and humans. Among the models we studied, GPT-5-mini and Claude Sonnet 4.5 showed the strongest alignment with human text ratings. Llama-4 was the best aligned open-source model. However, gaps between LLM and human behavior remain. No model we studied adequately captured the inter-individual variability observed among human participants, and models only moderately aligned with individual human responses. We demonstrate the utility of this approach across multiple modalities (words, sentences, and images), helping further our understanding of how LLMs encode knowledge, and enabling an examination of alignment with human cognition.
The IARPA Space-Based Machine Automated Recognition Technique (SMART) program recently posed a new challenge for computer vision in satellite imagery - searching for and characterizing anthropogenic activity, i.e. gradual changes to the Earth's surface caused by humans. While numerous solutions exist for detection and classification of objects in single images, characterizing activity requires spatio-temporal reasoning over long time scales. Given the novelty of this problem, standard performance metrics and benchmarks had not yet been established. This work proposes a suite of application-generic performance metrics for detecting and classifying sites where activities of interest have taken place, and presents initial results of several different algorithms aimed at spatially and temporally localizing and classifying heavy construction activity. An implementation of these performance metrics along with sample system outputs is available with the annotation dataset on Github(1) and IEEE DataPort.(2)
As we consider entrusting Large Language Models (LLMs) with key societal and decision-making roles, measuring their alignment with human cognition becomes critical. This requires methods that can assess how these systems represent information and facilitate comparisons to human understanding across diverse tasks. To meet this need, we developed Turing Representational Similarity Analysis (RSA), a method that uses pairwise similarity ratings to quantify alignment between AIs and humans. We tested this approach on semantic alignment across text and image modalities, measuring how different Large Language and Vision Language Model (LLM and VLM) similarity judgments aligned with human responses at both group and individual levels. GPT-4o showed the strongest alignment with human performance among the models we tested, particularly when leveraging its text processing capabilities rather than image processing, regardless of the input modality. However, no model we studied adequately captured the inter-individual variability observed among human participants. This method helped uncover certain hyperparameters and prompts that could steer model behavior to have more or less human-like qualities at an inter-individual or group level. Turing RSA enables the efficient and flexible quantification of human-AI alignment and complements existing accuracy-based benchmark tasks. We demonstrate its utility across multiple modalities (words, sentences, images) for understanding how LLMs encode knowledge and for examining representational alignment with human cognition.
Satellite-based remote sensing imagery is an effective means for detecting objects and structures in support of many applications. However, detecting the spatial and temporal bounds of a specific activity in satellite imagery is inherently more complex and research in this area is nascent. One reason for this is that describing an activity implies defining both spatial and temporal bounds and while activity is inherently continuous in nature, the geospatial (imagery) time series for any particular swath of ground provided by satellite imagery is relatively sparse and discrete in comparison. The IARPA Space-Based Machine Automated Recognition Technique (SMART)1 program is the first large-scale research program to target advancing the state of the art for automatically detecting, characterizing, and monitoring large-scale anthropogenic activity in global, multispectral satellite imagery. The program has two primary research objectives: 1) the "harmonization" of multiple imagery sources and 2) automated reasoning at scale to detect, characterize, and monitor activities of interest. This paper provides details on the goals, dataset, metrics, and lessons learned of the IARPA SMART program. By releasing the annotated dataset, the program aims to foster additional research in this area by the community at large.
Although satellite-based remote sensing imagery has been widely demonstrated to be an effective means of detecting objects and structures, detecting the spatial and temporal bounds of human activity is inherently more complex, and research in this area is nascent. In this paper, we describe the performance metrics established for the IARPA Space-based Machine-Automated Recognition Technique (SMART) to track progress toward meeting this challenge. We also share the latest algorithm performance results from the SMART program.
Naval intelligence plays a critical role in multi-domain operations by identifying and tracking vessels of interest, especially suspected "dark ships" operating in an emissions-controlled (EMCON) state. While applying machine learning (ML) to maritime satellite imagery could enable an automated open-ocean search capability for dark ships, ensuring the robustness of ML models to environmental variations in the maritime domain remains a challenge because training sets do not encapsulate all possible environmental conditions. To address the challenge of unsupervised domain adaptation (UDA) in ship classification, i.e. transferring a ML model from a labeled source domain to an unlabeled target domain, we propose employing combinations of semi-supervised learning (SSL) techniques with standalone UDA approaches. Specifically, we incorporate combinations of FixMatch, minimum class confusion, gradient reversal, and mixup augmentation into the standard cross-entropy supervised loss function. These interventions were compared in two domain shift settings, one in which the source and target domains are both comprised of simulated data, and another in which the source domain consists of only simulated data, and the target domain consists of only real data. Experimental results comparing the combinations of interventions to a regularized fine-tuning baseline demonstrate that the greatest improvements in model robustness were achieved when combinations of our SSL strategy (FixMatch) and UDA algorithms were incorporated into training.
Discovery of novel materials is slow but necessary for societal progress. Here, we demonstrate a closed-loop machine learning (ML) approach to rapidly explore a large materials search space, accelerating the intentional discovery of superconducting compounds. By experimentally validating the results of the ML-generated superconductivity predictions and feeding those data back into the ML model to refine, we demonstrate that success rates for superconductor discovery can be more than doubled. Through four closed-loop cycles, we report discovery of a superconductor in the Zr-In-Ni system, re-discovery of five superconductors unknown in the training datasets, and identification of two additional phase diagrams of interest for new superconducting materials. Our work demonstrates the critical role experimental feedback provides in ML-driven discovery, and provides a blueprint for how to accelerate materials progress.
The goal of meta-learning is to generalize to new tasks and goals as quickly as possible. Ideally, we would like approaches that generalize to new goals and tasks on the first attempt. Requiring a policy to perform on a new task on the first attempt without even a single example trajectory is a zero-shot problem formulation. When tasks are identified by goal images, the tasks can be considered visually goal-directed. In this work, we explore the problem of visual goal-directed zero-shot meta-imitation learning. Inspired by several popular approaches to Meta-RL, we composed several core ideas related to task-embedding and planning by gradient descent to attempt to explore this problem. To evaluate these approaches, we adapted the Meta-world benchmark tasks to create 24 distinct visual goal-directed manipulation tasks. We found that 7 out of 24 tasks could be successfully completed on the first attempt by at least one of the approaches we tested. We demonstrated that goal-directed zero-shot approaches can translate to a physical robot with a demonstration based on Jenga block manipulation tasks using a Kinova Jaco robotic arm.
The discovery of novel materials drives industrial innovation, although the pace of discovery tends to be slow due to the infrequency of "Eureka!" moments. These moments are typically tangential to the original target of the experimental work: "accidental discoveries". Here we demonstrate the acceleration of intentional materials discovery - targeting material properties of interest while generalizing the search to a large materials space with machine learning (ML) methods. We demonstrate a closed-loop ML discovery process targeting novel superconducting materials, which have industrial applications ranging from quantum computing to sensors to power delivery. By closing the loop, i.e. by experimentally testing the results of the ML-generated superconductivity predictions and feeding data back into the ML model to refine, we demonstrate that success rates for superconductor discovery can be more than doubled. In four closed-loop cycles, we discovered a new superconductor in the Zr-In-Ni system, re-discovered five superconductors unknown in the training datasets, and identified two additional phase diagrams of interest for new superconducting materials. Our work demonstrates the critical role experimental feedback provides in ML-driven discovery, and provides definite evidence that such technologies can accelerate discovery even in the absence of knowledge of the underlying physics.
Recent breakthroughs and rapid progress in AI will impact, if not transform, every mission. JHU/APL developed an AI Technology Roadmap to guide the Laboratory’s contributions to the critical challenges the nation will face developing and implementing intelligent systems for these missions over the coming decades. We began this exercise by describing a series of envisioned futures for intelligent systems across sea, land, air, space and information, and examined them to identify the common AI technology vectors needed to achieve each vision: (1) Autonomous Perception, describing the path to intelligent systems that perceive in the context of the extreme uncertainty and complexity of the real world; (2) Superhuman Decision-Making and Autonomous Action, to realize the potential for intelligent systems to reason over more information than any team of analysts or operators and act in ways systems under manned control cannot; (3) Human-Machine Teaming at the Speed of Thought, to ensure humans can stay involved at speed and scale; and of particular importance for national security applications, (4) Safe and Assured Operation, so these systems can be trusted to stay true to commander’s intent, in adversarial and sensitive contexts. Each technology vector is aligned with a targeted goal, and with each goal we provide a roadmap in the form of near-, mid-, and long-term AI advances critical to reaching the goal. This paper describes the JHU/APL AI Technology Roadmap and presents key examples of recent progress and forward-looking research and exploratory development along each vector.
Most approaches to deep reinforcement learning (DRL) attempt to solve a single task at a time. As a result, most existing research benchmarks consist of individual games or suites of games that have common interfaces but little overlap in their perceptual features, objectives, or reward structures. To facilitate research into knowledge transfer among trained agents (e.g. via multi-task and metalearning), more environment suites that provide configurable tasks with enough commonality to be studied collectively are needed. In this paper we present Meta Arcade, a tool to easily define and configure custom 2D arcade games that share common visuals, state spaces, action spaces, game components, and scoring mechanisms. Meta Arcade differs from prior environments in that both task commonality and configurability are prioritized: entire sets of games can be constructed from common elements, and these elements are adjustable through exposed parameters. We include a suite of 24 predefined games that collectively illustrate the possibilities of this framework and discuss how these games can be configured for research applications. We provide several experiments that illustrate how Meta Arcade could be used, including single-task benchmarks of predefined games, sample curriculum-based approaches that change game parameters over a set schedule, and an exploration of transfer learning between games.
Physics-based models for ocean dynamics and optical raytracing are used extensively for rendering maritime scenes in computer graphics [Darles et al. 2011]. Raytracing models can provide high-fidelity representations of an ocean image with full control of the underlying environmental conditions, sensor specifications, and viewing geometry. However, the computational expense of rendering ocean scenes can be high. This work demonstrates an alternative approach to ocean raytracing via machine learning, specifically Generative Adversarial Networks (GANs) [Goodfellow et al. 2014]. In this paper, we demonstrate that a GAN trained on several thousand small scenes produced by a raytracing model can be used to generate megapixel scenes roughly an order of magnitude faster with a consistent wave spectrum and minimal processing artifacts.
Feature selection is a common problem in pattern recognition. Though often motivated by the curse of dimensionality, feature selection also has the added benefit of reducing the cost of extracting features from test data. In this work, sparse probit models are modified to incorporate feature costs. A single-classifier approach, Cost-Constrained Feature optimization (CCFO), is compared to a new ensemble method referred to as the Cost-Constrained Classifier Cascade (C4). The C4 method utilizes a boosting framework that accommodates per-sample feature selection. Experimental results compare C4, CCFO, and baseline sparse kernel classification on two data sets with asymmetric feature costs, illustrating that C4 can yield similar or better accuracy and more economical use of expensive features.
Physics-based models for ocean dynamics and optical raytracing are used extensively for rendering maritime scenes in computer graphics [Darles et al. 2011]. Raytracing models can provide high-fidelity representations of an ocean image with full control of the underlying environmental conditions, sensor specifications, and viewing geometry. However, the computational expense of rendering ocean scenes can be high. This work demonstrates an alternative approach to ocean raytracing via machine learning, specifically Generative Adversarial Networks (GANs) [Goodfellow et al. 2014]. In this paper, we demonstrate that a GAN trained on several thousand small scenes produced by a raytracing model can be used to generate megapixel scenes roughly an order of magnitude faster with a consistent wave spectrum and minimal processing artifacts.
This paper considers attacks against machine learning algorithms used in remote sensing applications. The remote sensing domain presents a suite of challenges that are not fully addressed by current research focused on natural image data. In this paper we present a new study of adversarial examples in the context of satellite image classification problems. Using a recently curated data set and associated classifier, we provide a preliminary analysis of adversarial examples in settings where the targeted classifier is permitted multiple observations of the same location over time. While our experiments to date are purely digital, our problem setup incorporates a number of practical considerations that an attacker would need to take into account when mounting physical attacks.
Most Brain-Computer Interface (BCI) work has focused on detecting specific sensory or motor information, but BCIs are beginning to be applied to more abstract domains like covert speech and communication of semantic thought. One potential approach to decoding more abstract information is linear zero-shot classification via semantic attributes, which is computationally efficient and may facilitate real-time processing. In this work, several variations of this model are applied to electrocorticography (ECoG) data recorded during a picture-naming task with nine patients. Performances of encoding and decoding models are compared, and results are discussed in the context of BCI applications.
Dimensionality poses a serious challenge when making predictions from human neuroimaging data. Across imaging modalities, large pools of potential neural features (e.g., responses from particular voxels, electrodes, and temporal windows) have to be related to typically limited sets of stimuli and samples. In recent years, zero-shot prediction models have been introduced for mapping between neural signals and semantic attributes, which allows for classification of stimulus classes not explicitly included in the training set. While choices about feature selection can have a substantial impact when closed-set accuracy, open-set robustness, and runtime are competing design objectives, no systematic study of feature selection for these models has been reported. Instead, a relatively straightforward feature stability approach has been adopted and successfully applied across models and imaging modalities. To characterize the tradeoffs in feature selection for zero-shot learning, we compared correlation-based stability to several other feature selection techniques on comparable data sets from two distinct imaging modalities: functional Magnetic Resonance Imaging and Electrocorticography. While most of the feature selection methods resulted in similar zero-shot prediction accuracies and spatial/spectral patterns of selected features, there was one exception; A novel feature/attribute correlation approach was able to achieve those accuracies with far fewer features, suggesting the potential for simpler prediction models that yield high zero-shot classification accuracy.
Non-invasive neuroimaging studies have shown that semantic category and attribute information are encoded in neural population activity. Electrocorticography (ECoG) offers several advantages over non-invasive approaches, but the degree to which semantic attribute information is encoded in ECoG responses is not known. We recorded ECoG while patients named objects from 12 semantic categories and then trained high-dimensional encoding models to map semantic attributes to spectral-temporal features of the task-related neural responses. Using these semantic attribute encoding models, untrained objects were decoded with accuracies comparable to whole-brain functional Magnetic Resonance Imaging (fMRI), and we observed that high-gamma activity (70-110Hz) at basal occipitotemporal electrodes was associated with specific semantic dimensions (manmade-animate, canonically large-small, and places-tools). Individual patient results were in close agreement with reports from other imaging modalities on the time course and functional organization of semantic processing along the ventral visual pathway during object recognition. The semantic attribute encoding model approach is critical for decoding objects absent from a training set, as well as for studying complex semantic encodings without artificially restricting stimuli to a small number of semantic categories.
Feature selection is often necessary when implementing classifiers in practice. Most approaches to feature selection are motivated by the curse of dimensionality, but few seek to mitigate the overall computational cost of feature extraction. In this work, we propose a model-based approach for addressing both objectives. The model is based around a sparse kernel machine with feature scaling parameters controlled by a beta-Bernoulli prior. The hyperparameters are controlled by each feature's computational cost. Experiments were carried out using publicly-available data sets, and the proposed Cost-Constrained Feature Optimization (CCFO) was compared to related methods in terms of accuracy and computational reduction.