We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in materials science, each annotated with a ground-truth feasibility score on a five-point scale along with an explanation of that assessment. The collection differs from previous collections in several important ways: 1) it defines a complex task that requires reasoning over claims of varying scientific feasibility; 2) its claims are not extracted from existing scientific publications but are created de novo, greatly reducing the chances that LLMs have trained on them; 3) claims and ground truth are established by subject matter experts, not by artificial intelligence; and 4) unlike many benchmarks that ask about question/answer pairs, provide multiple choice answers, or ask questions requiring short, fixed answers, SFBench explanations are completely open-ended. We describe the benchmark design, data creation process, and evaluation metrics, and we report baseline results using recent GPT models.
A crucial part of facilitating the cooperation of multi-robot and human-robot teams is a Common World Model – a shared knowledge base with both physical information (e.g., ground map) and semantic information (e.g, locations of threats and goals) – that can be used to provide high-level guidance to heterogeneous robot teams. Past work performed by Johns Hopkins University Applied Physics Laboratory (JHU/APL) has shown that the Advanced Explosive Ordnance Disposal Robotic System (AEODRS) architecture – a Modular Open Systems Approach (MOSA) architecture leveraging the JAUS (Joint Architecture for Unmanned Systems) standard for definition of its logical interfaces – can be effectively used to develop and integrate the subsystems of a teleoperated ground vehicle for use in a complex environment. This demonstration tackles the next challenge, which is to extend the AEODRS architecture to facilitate multi-robot and human-robot teams.
Generative machine learning models can use data generated by scientific modeling to create large quantities of novel material structures. Here, we assess how one state-of-the-art generative model, the physics-guided crystal generation model (PGCGM), can be used as part of the inverse design process. We show that the default PGCGM's input space is not smooth with respect to parameter variation, making material optimization difficult and limited. We also demonstrate that most generated structures are predicted to be thermodynamically unstable by a separate property-prediction model, partially due to out-of-domain data challenges. Our findings suggest how generative models might be improved to enable better inverse design.
Ships inside the Arctic basin require high-resolution (1-5 km), near-term (days to semimonthly) forecasts for guidance on scales of interest to their operations where forecast model predictions are insufficient due to their coarse spatial and temporal resolutions. Deep learning techniques offer the capability of rapid assimilation and analysis of multiple sources of information for improved forecasting. Data from the National Oceanographic and Atmospheric Administration's Global Forecast System, Multi-scale Ultra-high Resolution Sea Surface Temperature (MEaSUREs), and the National Snow and Ice Data Center's Multisensor Analyzed Sea ice Extent (MASIE) were used to develop the sea ice extent deep learning forecast model, over the freeze-up periods of 2016, 2018, 2019, and 2020 in the Beaufort Sea. Sea ice extent forecasts were produced for 1-7 days in the future. The approach was novel for sea ice extent forecasting in using forecast data as model input to aid in the prediction of sea ice extent. Model accuracy was assessed against a persistence model. While the average accuracy of the persistence model dropped from 97% to 90% for forecast days 1-7, the deep learning model accuracy dropped only to 93%. A k-fold (fourfold) cross-validation study found that on all except the first day, the deep learning model, which includes a U-Net architecture with an 18-layer residual neural network (Resnet-18) backbone, does better than the persistence model. Skill scores improve the farther out in time to 0.27. The model demonstrated success in predicting changes in ice extent of significance for navigation in the Amundsen Gulf. Extensions to other Arctic seas, seasons, and sea ice parameters are under development.
Discovery of novel materials is slow but necessary for societal progress. Here, we demonstrate a closed-loop machine learning (ML) approach to rapidly explore a large materials search space, accelerating the intentional discovery of superconducting compounds. By experimentally validating the results of the ML-generated superconductivity predictions and feeding those data back into the ML model to refine, we demonstrate that success rates for superconductor discovery can be more than doubled. Through four closed-loop cycles, we report discovery of a superconductor in the Zr-In-Ni system, re-discovery of five superconductors unknown in the training datasets, and identification of two additional phase diagrams of interest for new superconducting materials. Our work demonstrates the critical role experimental feedback provides in ML-driven discovery, and provides a blueprint for how to accelerate materials progress.
Despite the advancement of machine learning techniques in recent years, state-of-the-art systems lack robustness to "real world" events, where the input distributions and tasks encountered by the deployed systems will not be limited to the original training context, and systems will instead need to adapt to novel distributions and tasks while deployed. This critical gap may be addressed through the development of "Lifelong Learning" systems that are capable of 1) Continuous Learning, 2) Transfer and Adaptation, and 3) Scalability. Unfortunately, efforts to improve these capabilities are typically treated as distinct areas of research that are assessed independently, without regard to the impact of each separate capability on other aspects of the system. We instead propose a holistic approach, using a suite of metrics and an evaluation framework to assess Lifelong Learning in a principled way that is agnostic to specific domains or system techniques. Through five case studies, we show that this suite of metrics can inform the development of varied and complex Lifelong Learning systems. We highlight how the proposed suite of metrics quantifies performance trade-offs present during Lifelong Learning system development - both the widely discussed Stability-Plasticity dilemma and the newly proposed relationship between Sample Efficient and Robust Learning. Further, we make recommendations for the formulation and use of metrics to guide the continuing development of Lifelong Learning systems and assess their progress in the future.
Properties of interest for crystals and molecules, such as band gap, elasticity, and solubility, are generally related to each other: they are governed by the same underlying laws of physics. However, when state-of-the-art graph neural networks attempt to predict multiple properties simultaneously (the multi-task learning (MTL) setting), they frequently underperform a suite of single property predictors. This suggests graph networks may not be fully leveraging these underlying similarities. Here we investigate a potential explanation for this phenomenon: the curvature of each property's loss surface significantly varies, leading to inefficient learning. This difference in curvature can be assessed by looking at spectral properties of the Hessians of each property's loss function, which is done in a matrix-free manner via randomized numerical linear algebra. We evaluate our hypothesis on two benchmark datasets (Materials Project (MP) and QM8) and consider how these findings can inform the training of novel multi-task learning models.
The discovery of novel materials drives industrial innovation, although the pace of discovery tends to be slow due to the infrequency of "Eureka!" moments. These moments are typically tangential to the original target of the experimental work: "accidental discoveries". Here we demonstrate the acceleration of intentional materials discovery - targeting material properties of interest while generalizing the search to a large materials space with machine learning (ML) methods. We demonstrate a closed-loop ML discovery process targeting novel superconducting materials, which have industrial applications ranging from quantum computing to sensors to power delivery. By closing the loop, i.e. by experimentally testing the results of the ML-generated superconductivity predictions and feeding data back into the ML model to refine, we demonstrate that success rates for superconductor discovery can be more than doubled. In four closed-loop cycles, we discovered a new superconductor in the Zr-In-Ni system, re-discovered five superconductors unknown in the training datasets, and identified two additional phase diagrams of interest for new superconducting materials. Our work demonstrates the critical role experimental feedback provides in ML-driven discovery, and provides definite evidence that such technologies can accelerate discovery even in the absence of knowledge of the underlying physics.
Earth and Space Science Open Archive PosterOpen AccessYou are viewing the latest version by default [v1]Short-Term Sea Ice Extent Forecasting with Deep LearningAuthorsMaryKelleriDChristinePiatkoMaryClemens-SewallRebeccaEagerKevinFosterChristopherGiffordiDDerekRollendiDJenniferSleemanSee all authors Mary KelleriDCorresponding Author• Submitting AuthorJohns Hopkins University Applied Physics LaboratoryiDhttps://orcid.org/0000-0003-0669-1298view email addressThe email was not providedcopy email addressChristine PiatkoThe Johns Hopkins University/Applied Physics Laboratoryview email addressThe email was not providedcopy email addressMary Clemens-SewallThe Johns Hopkins University/Applied Physics Laboratoryview email addressThe email was not providedcopy email addressRebecca EagerThe Johns Hopkins University/Applied Physics Laboratoryview email addressThe email was not providedcopy email addressKevin FosterThe Johns Hopkins University/Applied Physics Laboratoryview email addressThe email was not providedcopy email addressChristopher GiffordiDThe John Hopkins University/Applied Physics LaboratoryiDhttps://orcid.org/0000-0002-3848-6267view email addressThe email was not providedcopy email addressDerek RollendiDJohns Hopkins University Applied Physics LaboratoryiDhttps://orcid.org/0000-0003-0826-859Xview email addressThe email was not providedcopy email addressJennifer SleemanThe Johns Hopkins University/Applied Physics Laboratoryview email addressThe email was not providedcopy email address
Cyber defense requires decision making under uncertainty, yet this critical area has not been a focus of research in judgment and decision-making. Future defense systems, which will rely on software-defined networks and may employ “moving target” defenses, will increasingly automate lower level detection and analysis, but will still require humans in the loop for higher level judgment. We studied the decision making process and outcomes of 17 experienced network defense professionals who worked through a set of realistic network defense scenarios. We manipulated gain versus loss framing in a cyber defense scenario, and found significant effects in one of two focal problems. Defenders that began with a network already in quarantine (gain framing) used a quarantine system more, as measured by cost, than those that did not (loss framing). We also found some difference in perceived workload and efficacy. Alternate explanations of these findings and implications for network defense are discussed.
The linear least trimmed squares (LTS) estimator is a statistical technique for fitting a linear model to a set of points. It was proposed by Rousseeuw as a robust alternative to the classical least squares estimator. Given a set of n points in Rd, the objective is to minimize the sum of the smallest 50% squared residuals (or more generally any given fraction). There exist practical heuristics for computing the linear LTS estimator, but they provide no guarantees on the accuracy of the final result. Two results are presented. First, a measure of the numerical condition of a set of points is introduced. Based on this measure, a probabilistic analysis of the accuracy of the best LTS fit resulting from a set of random elemental fits is presented. This analysis shows that as the condition of the point set improves, the accuracy of the resulting fit also increases. Second, a new approximation algorithm for LTS, called Adaptive-LTS, is described. Given bounds on the minimum and maximum slope coefficients, this algorithm returns an approximation to the optimal LTS fit whose slope coefficients lie within the given bounds. Empirical evidence of this algorithm's efficiency and effectiveness is provided for a variety of data sets.
Most cyber security research is focused on detecting network intrusions or anomalies through the use of automated methods, exploratory visual analytics systems, or real-time monitoring using dynamic visual representations. However, there has been minimal investigation of effective decision support systems for cyber analysts. This paper describes the user-centered design and development of a decision support visualization for active network defense. Ocelot helps the cyber analyst assess threats to a network and quarantine affected computers from the healthy parts of a network. The described web-based, functional visualization prototype integrates and visualizes multiple data sources through the use of a hybrid space partitioning tree and node link diagram. We describe our design process for requirements gathering and design feedback which included expert interviews, iterative design, and a user study.
The presented method is a practical, understandable way to monitor single care facilities for chief complaint clusters of concern based on unusually high occurrence of rare or common terms that need not be related to syndromes. Routine implementation requires a human monitor to inspect the relevant CCs make follow-up decisions. Using 7 years of patient records from 15 hospitals, our approach pools CCs into contiguous time blocks and uses a statistical hypothesis test to seek current terms that are anomalous relative to their occurrence in a large sliding baseline. Sets of anomalous terms are then presented for further investigation.
We explore training an automatic modality tagger. Modality is the attitude that a speaker might have toward an event or state. One of the main hurdles for training a linguistic tagger is gathering training data. This is particularly problematic for training a tagger for modality because modality triggers are sparse for the overwhelming majority of sentences. We investigate an approach to automatically training a modality tagger where we first gathered sentences based on a high-recall simple rule-based modality tagger and then provided these sentences to Mechanical Turk annotators for further annotation. We used the resulting set of training data to train a precise modality tagger using a multi-class SVM that delivers good performance.
Large scale data fusion of multiple datasets can often provide insights that individual datasets cannot. However, when these datasets reside in different data centers and cannot be collocated due to technical, administrative, or policy barriers, a unique set of problems arise that hamper querying and data fusion. To address these problems, a system and architecture named Parasol is presented that enables federated queries over graph databases residing in multiple clouds. Parasol's design is flexible and requires only minimal assumptions for client clouds. Query optimization techniques are also described that are compatible with Parasol's lightweight architecture. Experiments on a prototype implementation of Parasol indicate its suitability for cross-cloud federated graph queries.
The linear least trimmed squares (LTS) estimator is a statistical technique for fitting a linear model to a set of points. Given a set of n points in ℝ d and given an integer trimming parameter h ≤ n , LTS involves computing the ( d −1)-dimensional hyperplane that minimizes the sum of the smallest h squared residuals. LTS is a robust estimator with a 50 %-breakdown point, which means that the estimator is insensitive to corruption due to outliers, provided that the outliers constitute less than 50 % of the set. LTS is closely related to the well known LMS estimator, in which the objective is to minimize the median squared residual, and LTA, in which the objective is to minimize the sum of the smallest 50 % absolute residuals. LTS has the advantage of being statistically more efficient than LMS. Unfortunately, the computational complexity of LTS is less understood than LMS. In this paper we present new algorithms, both exact and approximate, for computing the LTS estimator. We also present hardness results for exact and approximate LTS. A number of our results apply to the LTA estimator as well.