Synchrotron beamlines differ in hardware, technique, and workflow, making customized control interfaces necessary; bespoke per-beamline graphical user interfaces (GUIs) do not scale well, one-size-fits-all facility software forces compromises that leave most of the interface unused, and even recent component-library approaches keep per-scientist tweaks on a developer's queue. We present Lightfall, a control platform designed for facility-wide use, whose API-first architecture exposes every panel, device, and scan plan through a single uniform addressable interface. An embedded language-model agent drives experiments through that interface, from a single move-and-read to a Gaussian-process-driven autonomous scan, while beamline staff extend the interface during operation via skills: plugin modules the agent invokes to compose and modify panels in the running application. The result is a closed development loop: a beamline scientist authors a panel change in natural language, the agent emits and applies it, and the commit lands in the beamline's plugin repository as a side effect. The per-iteration cost of a scientist-driven change is then fixed in the scientist's own time rather than in developer hours the facility must supply. Lightfall is in testing at the COSMIC-Scattering beamline at the Advanced Light Source.
Conditional density estimation is complicated by multimodality, heteroscedasticity, and strong non-Gaussianity. Gaussian processes (GPs) provide a principled nonparametric framework with calibrated uncertainty, but standard GP regression is limited by its unimodal Gaussian predictive form. We introduce the Generalized Gaussian Mixture Process (GGMP), a GP-based method for multimodal conditional density estimation in settings where each input may be associated with a complex output distribution rather than a single scalar response. GGMP combines local Gaussian mixture fitting, cross-input component alignment and per-component heteroscedastic GP training to produce a closed-form Gaussian mixture predictive density. The method is tractable, compatible with standard GP solvers and scalable methods, and avoids the exponentially large latent-assignment structure of naive multimodal GP formulations. Empirically, GGMPs improve distributional approximation on synthetic and real-world datasets with pronounced non-Gaussianity and multimodality.
Chiral 2D metal halide perovskites (MHPs) are promising for spin-optoelectronic applications, yet their absorption dissymmetry factor (gabs) exhibits significant variability due to complex, co-dependent structural and experimental factors. We established a data-driven framework using Pearson’s correlation, ANOVA, and Gaussian process regression to identify and model key synthesis “knobs” governing these properties. The analysis revealed that solvent choice is the primary factor driving variability. For acetonitrile-based films, gabs was maximized by optimizing annealing temperature and film thickness. Conversely, films from higher boiling point solvents showed complex dependencies on annealing temperature, excitonic integral intensity, and film texture. These statistical correlations provide a roadmap for the rational design of high-performance chiral MHPs and establish a foundation for future machine learning-driven material exploration.
There are numerous ways in which the battery research community could--but does not--address issues that are important to manufacturers and end users. We will discuss them here. Shifting towards sustainable energy sources requires developing new storage systems and estimating their remaining energy over their lifetime. The remaining energy of these systems depends on many operating parameters, resulting in a large high-dimensional parameter space to explore. Testing cells exhaustively on a dense grid in the parameter space is prohibitively expensive. This is especially true with considerable cell-to-cell variability in performance, even under the same cycling conditions. Here, we develop a framework based on Gaussian processes, equipped with domain knowledge, to implement Bayesian optimization to explore the parameter space efficiently and quantify remaining energy using failure distributions. Bayesian optimization identifies future experiments that maximize information gain and minimize uncertainty. Experimental results show accurate remaining energy predictions with significantly fewer experiments. However, laboratory cycling conditions, including those in literature, may not represent real-world cycling. We propose an approach based on laboratory results to predict remaining energy under real-world cycling condition
IntroductionAdvances in automation and AI/ML offer new opportunities for plant science, including design, modeling, and analysis. This study aimed to develop an automated platform for researching small model plants under axenic conditions and integrate it with AI/ML tools.MethodsThe EcoBOT platform was developed, which consists of sterile containers (EcoFABs) for growing plants and imaging for monitoring plant growth and health. Brachypodium distachyon was grown on the EcoBOT, and its response to nutrient limitation and copper stress was evaluated.ResultsThe results showed that Brachypodium distachyon grown in the EcoBOT maintained sterility and responded to nutrient limitation and copper stress. Analysis of over 6,500 root and shoot images revealed varying sensitivity and response rates to copper. Bayesian Optimization was used to improve model accuracies relating copper concentrations to plant biomass via sequential experiments, resulting in a >30% improvement.DiscussionThe findings of this study demonstrate the potential of the EcoBOT platform for researching plant responses to environmental factors. Future experiments could focus on relating other chemical stresses and microbial interactions to create generalized models of plant responses.
There is a growing focus on new energy sources and storage systems. The challenge with such emerging systems is their need to be warrantied for around 15 years with just a year of early testing. This requires accurate data extrapolation and estimation of the failure distribution. Physics-based approaches can be overwhelmed by the complexity of degradation, and pure data-driven approaches are inherently unable to extrapolate beyond the testing data. Here, we propose a framework for a hybrid approach for technology-agnostic customizations of a Gaussian process for stochastic and domain-knowledge-informed failure-distribution predictions. We equip the Gaussian process with customized non-stationary kernels, heteroscedastic noise models, and prior mean functions to allow for accurate extrapolation with high accuracy. Furthermore, we minimize testing time with an experiment-stopping criterion, which can significantly reduce the required data. Our framework could revolutionize energy-storage testing, enabling the rapid development of new technologies.
Materials Acceleration Platforms (MAPs) - also known as self-driving laboratories- present a new paradigm for materials science and promise an order of magnitude accelerated materials discovery compared to the traditional trial-and-error approach. Metal halide perovskites (MHPs) are an emerging class of materials for optoelectronic applications but are plagued by irreproducible optoelectronic quality, particularly for films fabricated in a humid atmosphere. Here, a machine learning (ML)-guided closed-loop platform is developed with a multimodal data fusion approach to predict synthesis-property relations for the optical quality of MHP thin films in relative humidities (RHs) ranging from 5-55%. The efficiency of this approach is confirmed by the fast-dropping learning rate to 2% after experimentally sampling less than 1% of the possible 5,000+ combinations. The prediction of synthesis-property relations is done by optical and imaging characterizations. In situ photoluminescence characterization revealed the origin of thin film quality variation at different RH. These insights provide an avenue for controlling the MHP crystallization by fine-tuning the synthesis parameters and RH for a given chemistry, thus lifting the need for stringent atmosphere control. The MAP enables an accelerated screening and understanding of the synthesis design space, facilitating rational synthesis recipe choice for a wide range of materials.
The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space-time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.
The increase in energy demand requires developing new storage systems and estimating their remaining energy over their lifetime. The remaining energy of these systems depends on many operating parameters, resulting in a large high-dimensional parameter space to explore. Testing cells exhaustively on a dense grid in the parameter space is prohibitively expensive. This is especially true with considerable cell-to-cell variability in performance, even under the same cycling conditions. Here, we develop a framework based on Gaussian processes, equipped with domain knowledge, and implement Bayesian optimization to explore the parameter space efficiently and quantify remaining energy using failure distributions. Bayesian optimization identifies future experiments that maximize information gain and minimize uncertainty. Experimental results show accurate remaining energy predictions with significantly fewer experiments. However, laboratory cycling conditions, including those in the literature, may not represent real-world cycling. We propose an approach based on laboratory results for predicting remaining energy under real-world cycling conditions.
Aberration correction is an important aspect of modern high-resolution scanning transmission electron microscopy. Most methods of aligning aberration correctors require specialized sample regions and are unsuitable for fine-tuning aberrations without interrupting on-going experiments. Here, we present an automated method of correcting first- and second-order aberrations called BEACON which uses Bayesian optimization of the normalized image variance to efficiently determine the optimal corrector settings. We demonstrate its use on gold nanoparticles and a hafnium dioxide thin film showing its versatility in nano- and atomic-scale experiments. BEACON can correct all first- and second-order aberrations simultaneously to achieve an initial alignment and first- and second-order aberrations independently for fine alignment. Ptychographic reconstructions are used to demonstrate an improvement in probe shape and a reduction in the target aberration.
In recent years, several groups have designed Autonomous Experiment (AE) models with the aim of using them as an alternative method for neutron scattering scanning. In an AE, Gaussian processes (GPs) are most frequently used due to their interpretability, their non-parametric nature, their universal approximation, and their closed-form predictive distribution. GPs have two key components, namely, the model for the likelihood of a neutron count knowing the underlying dynamic structure factor and the acquisition function. In this paper, we investigate the impact, on the quality of an AE, of the likelihood and acquisition function choices, in energy scans and (Q, ω) ones, with respect to the signal-over-noise ratio. While we hypothesized that the quality of GP predictions would decrease when the normal to Poisson likelihood approximation breaks down at low count rates, we found that the use of the correct Poisson likelihood does not improve the quality of the data collected, as well as yields very poor results in (Q, ω) scans at low count rates. In fact, the best results are obtained with a combination of normal likelihood, including the observation noise, and the change in variance acquisition function. In addition, we find that the performance, or quality of the predictive distribution, is a misleading measure of efficiency, that is, of the quality of the data collected.
There is a growing focus on sustainable energy sources and storage systems. The challenge with such emerging systems is their need to be warrantied for around 15 years with just a year of early testing. This requires accurate data extrapolation and estimation of the failure distribution. Physics-based approaches can be overwhelmed by the complexity of degradation, and pure data-driven approaches are inherently unable to extrapolate beyond the testing data. Here, we propose a framework for a hybrid approach for technology-agnostic customizations of a Gaussian process for stochastic and domain-knowledge-informed failure distribution predictions. We equip the Gaussian process with customized non-stationary kernels, heteroscedastic noise models, and prior-mean functions to allow for accurate extrapolation with high accuracy. Furthermore, we minimize testing time with a novel experiment-stopping criterion, which can significantly reduce the required data. Our framework could revolutionize energy-storage testing, enabling the rapid development of new technologies.
The Gaussian process (GP) is a widely used probabilistic machine learning method with implicit uncertainty characterization for stochastic function approximation, stochastic modeling, and analyzing real-world measurements of nonlinear processes. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about ten thousand points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space-time prediction based on more than one million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.
Point defects in two-dimensional materials are of key interest for quantum information science. However, the space of possible defects is immense, making the identification of high-performance quantum defects extremely challenging. Here, we perform high-throughput (HT) first-principles computational screening to search for promising quantum defects within WS$_2$, which present localized levels in the band gap that can lead to bright optical transitions in the visible or telecom regime. Our computed database spans more than 700 charged defects formed through substitution on the tungsten or sulfur site. We found that sulfur substitutions enable the most promising quantum defects. We computationally identify the neutral cobalt substitution to sulfur (Co$_{\rm S}^{0}$) as very promising and fabricate it with scanning tunneling microscopy (STM). The Co$_{\rm S}^{0}$ electronic structure measured by STM agrees with first principles and showcases an attractive new quantum defect. Our work shows how HT computational screening and novel defect synthesis routes can be combined to design new quantum defects.
Machine learning Gaussian Process analysis is applied to search an experimental 4D parameter space that includes temperature, C rate, maximum SOC, and cycle number. Predictions are made for any point in parameter space, together with the uncertainty in each prediction
Autonomous experimentation (AE) is an emerging paradigm that seeks to automate the entire workflow of an experiment, including-crucially-the decision-making step. Beyond mere automation and efficiency, AE aims to liberate scientists to tackle more challenging and complex problems. We describe our recent progress in the application of this concept at synchrotron x-ray scattering beamlines. We automate the measurement instrument, data analysis, and decision-making, and couple them into an autonomous loop. We exploit Gaussian process modeling to compute a surrogate model and associated uncertainty for the experimental problem, and define an objective function exploiting these. We provide example applications of AE to x-ray scattering, including imaging of samples, exploration of physical spaces through combinatorial methods, and coupling to in situ processing platforms These uses demonstrate how autonomous x-ray scattering can enhance efficiency, and discover new materials.
The Gaussian process (GP) is a popular statistical technique for stochastic function approximation and uncertainty quantification from data. GPs have been adopted into the realm of machine learning (ML) in the last two decades because of their superior prediction abilities, especially in data-sparse scenarios, and their inherent ability to provide robust uncertainty estimates. Even so, their performance highly depends on intricate customizations of the core methodology, which often leads to dissatisfaction among practitioners when standard setups and off-the-shelf software tools are being deployed. Arguably, the most important building block of a GP is the kernel function, which assumes the role of a covariance operator. Stationary kernels of the Matérn class are used in the vast majority of applied studies; poor prediction performance and unrealistic uncertainty quantification are often the consequences. Non-stationary kernels show improved performance but are rarely used due to their more complicated functional form and the associated effort and expertise needed to define and tune them optimally. In this perspective, we want to help ML practitioners make sense of some of the most common forms of non-stationarity for Gaussian processes. We show a variety of kernels in action using representative datasets, carefully study their properties, and compare their performances. Based on our findings, we propose a new kernel that combines some of the identified advantages of existing kernels.