Bayesian Optimization (BO) is a powerful tool for optimizing complex non-linear systems. However, its performance degrades in high-dimensional problems with tightly coupled parameters and highly asymmetric objective landscapes, where rewards are sparse. In such needle-in-a-haystack scenarios, even advanced methods like trust-region BO (TurBO) often lead to unsatisfactory results. We propose a domain knowledge guided Bayesian Optimization approach, which leverages physical insight to fundamentally simplify the search problem by transforming coordinates to decouple input features and align the active subspaces with the primary search axes. We demonstrate this approach's efficacy on a challenging 12-dimensional, 6-crystal Split-and-Delay optical system, where conventional approaches, including standard BO, TuRBO and multi-objective BO, consistently led to unsatisfactory results. When combined with an reverse annealing exploration strategy, this approach reliably converges to the global optimum. The coordinate transformation itself is the key to this success, significantly accelerating the search by aligning input co-ordinate axes with the problem's active subspaces. As increasingly complex scientific instruments, from large telescopes to new spectrometers at X-ray Free Electron Lasers are deployed, the demand for robust high-dimensional optimization grows. Our results demonstrate a generalizable paradigm: leveraging physical insight to transform high-dimensional, coupled optimization problems into simpler representations can enable rapid and robust automated tuning for consistent high performance while still retaining current optimization algorithms.
X-ray imaging is a powerful technique to scan samples in a variety of contexts including biological, environmental and materials science, but commonly requires a synchrotron light source to produce X-rays at sufficient intensity. As these facilities are expensive to operate, the available beam time is limited and always in high demand. Particularly if the illuminated samples are sparse, standard raster scanning methods can be time-consuming, with a majority of that time being spent on areas of the image that carry little information. To increase the efficiency and maximize the information gain for a given time budget, we split the scanning process into a series of steps where previous measurements are used to inform the decision making and adapt the exposure distribution at later stages of the sequence. We formulate this task as a reinforcement learning problem where the goal is to produce a sequence of exposure maps that maximize a predefined scalar metric. We demonstrate the potential of this approach in simulations where the adaptive illumination can accelerate the measurement process by up to an order of magnitude compared with standard raster scanning. Finally, we present the first results from deploying the trained agents on an X-ray fluorescence beamline at the Stanford Synchrotron Radiation Lightsource.
Gravitational wave (GW) observatories have used template-based search to detect hundreds of compact binary coalescences (CBCs). However, template-based search cannot detect astrophysical sources that lack accurate, computationally tractable waveform models. Here, we present a novel approach for template-free search using coincident anomaly detection (CoAD). CoAD requires neither labeled training examples nor background-only training sets, instead exploiting the coincidence of events across spatially separated detectors as the training loss itself: two neural networks independently analyze data from each detector and are trained to maximize coincident predictions. Additionally, we show that integrated gradient analysis can localize GW signals from the neural network weights, providing a path toward data-driven template construction of unmodeled sources and further improving precision by frequency matching. Using the Codabench dataset of real LIGO backgrounds with injected simulated CBCs and sine-Gaussian low-frequency bursts, CoAD achieves recall up to 0.91 and 0.85, respectively, at a false-alarm rate of one event per year and achieves recall above 0.5 at signal-to-noise ratios below 10. The fully unsupervised nature of CoAD makes it especially well suited for next-generation detectors with greater sensitivity and associated increases in GW event rates.
Electrons in matter can rearrange extremely quickly under external perturbations, underpinning subsequent structural and chemical transformations. Coulomb interactions between neighbouring electrons often shape this response, giving rise to correlated motion and strongly affecting the distribution of electrons in the system. Here we show that non-resonant hard X-ray scattering can directly access changes in the radial electron-pair density during the rapid rearrangement of core and valence electrons. We do this by studying sulfur hexafluoride molecules undergoing Auger-Meitner decay. We exploit a second-order interaction between the X-ray photons and the molecules to trigger and probe the decay dynamics with a single pulse, capturing the electron loss and redistribution before the molecules dissociate. The experiment shows that changes in electron-pair densities can be isolated and measured on ultrafast timescales, providing insight into the real-space evolution of highly excited and short-lived electronic states.
Reliability is one of the most critical metrics for accelerator operation, especially in user facilities. To reduce costly facility downtime and provide an operational environment where system performance can be reliably predicted in support of scientific studies, we are developing a model-driven approach for prediction and anomaly detection. In this study, we present the application of a model-driven method that employs a linear regression model to predict the future temperature, in real time, of accelerator magnets at the NSLS-II light source. This approach enables proactive identification of magnet-heating issues, facilitating magnet flushing prior to the occurrence of permanent damage without interrupting machine operation. The implementation of this method in the NSLS-II control room is described and the analysis of the online results is presented. The results demonstrate the model's effectiveness in providing early alerts to engineers and improving the reliability of accelerator operations.
This white paper summarizes scientific challenges and AI/ML research opportunities identified through the FAIRS Japan 2024 unconference process. The discussion focuses on three major physics domains: accelerator physics, cosmology and astrophysics, and neutrino physics. Although each domain has distinct scientific goals and experimental constraints, several common technical themes emerge: high-dimensional reconstruction, fast and accurate simulation, uncertainty propagation, simulation-to-data mismatch, anomaly detection, real-time decision-making, and shared infrastructure.
To exploit the thousand-fold increase in spectral brightness of modern light sources, increasingly intricate experiments are being conducted that demand extremely precise beam trajectory. Maintaining the optimal trajectory over several hours of an experiment with the needed precision necessitates active drift control. Here, we outline time varying Bayesian optimization (TVBO) as a data driven approach for robust drift correction, and illustrate its application for a split and delay optical system composed of six crystals and twelve input dimensions. Using numerical simulations, we exhibit the application of TVBO for linear drift, non-smooth temporal drift as well as constrained TVBO for multi-objective control settings, representing real-life operating conditions. This approach can be easily adapted to other X-ray beam conditioning and guidance systems, including multi-crystal monochromators and grazing-incidence mirrors, to maintain sub-micrometer and nanoradian beam stability over the course of an experiment spanning several hours.
We present the first application of simulation-based inference to resonant inelastic X-ray scattering spectroscopy. Using truncated marginal neural ratio estimation to efficiently restrict the prior and conditional flow matching as the joint density estimator, we infer full posteriors with a modest simulation budget for two Ni^2+ compounds—NiPS_3 as a representative covalent case and K_2NiF_4 as a more atomic one. We demonstrate that a vision transformer encoder whose tokenization matches the physical layout of the RIXS map yields better-covered and sharper posteriors than generic image encoders. Applying the validated method to experimental NiPS_3 and K_2NiF_4 data, we recover a joint posterior that reveals parameter correlations invisible to point estimators, and a posterior predictive distribution that closely matches the observed spectrum. The amortized posterior unlocks a class of analyses not previously available to the field such as nuisance-marginalized uncertainty quantification, multi-measurement posterior fusion and active experimental design.
Bayesian optimal experimental design (BOED) seeks to maximize the expected information gain (EIG) of experiments. This requires a likelihood estimate, which in many settings is intractable. Simulation-based inference (SBI) provides powerful tools for this regime. However, existing work explicitly connecting SBI and BOED is restricted to a single contrastive EIG bound. We show that the EIG admits multiple formulations which can directly leverage modern SBI density estimators, encompassing neural posterior, likelihood, and ratio estimation. Building on this perspective, we define a novel EIG estimator using neural likelihood estimation. Further, we identify optimization as a key bottleneck of gradient based EIG maximization and show that a simple multi-start parallel gradient ascent procedure can substantially improve reliability and performance. With these innovations, our SBI-based BOED methods are able to match or outperform by up to 22% existing state-of-the-art approaches across standard BOED benchmarks.
Metastable states and their minimum energy pathways (MEPs) are central to understanding transformations and phase stability in complex materials, yet mapping transition pathways between competing states remains computationally demanding and experimentally challenging. Here, we introduce a hybrid solid-state nudged elastic band (SSNEB) framework that integrates two pretrained machine learning models, EquiformerV2 (eqV2) and the equivariant Smooth Energy Network (eSEN), with DFT for energy, force, and stress evaluations. Applied to three solid-state systems, CsPbI_3, GaN, and TiO_2, our framework achieves up to a 7-fold speedup while converging to the same pathways predicted by first-principles calculations. Moreover, the hybrid SSNEB framework enables systematic benchmarking of existing ML models, providing both efficiency and reliability for predicting MEPs across various materials.
Understanding and manipulating two-dimensional materials for real-world applications remains challenging due to a lack of effective and high-throughput characterization techniques. Soft X-ray time-of-flight photoemission electron microscopy (XPEEM) provides element- and depth-sensitive information of materials and buried interfaces. However, chromatic and spherical aberrations cannot be corrected with electron-lens combinations. These aberrations, combined with astigmatism and space-charge effects, significantly degrade the spatial and energy resolutions. To overcome this limitation, we outline a spatial-attention based deep learning approach to automatically correct for these effects and attain nanometer resolution over the entire field-of-view (FoV). The combination of this corrective algorithm with XPEEM, termed as nanoXPEEM, establishes a new record of 48-nm spatial resolution with a 232-micrometer diameter FoV in the soft x-ray regime (700-1000 eV). nanoXPEEM provides unique spatial mapping of the element-specificity, depth-sensitivity, and local structure on the nanoscale. It can bridge the current gap to achieve angstrom (atomic) scale resolution.
The discovery of a minimum energy pathway (MEP) between metastable states is crucial for scientific tasks including catalyst and biomolecular design. However, the standard nudged elastic band (NEB) algorithm requires hundreds to tens of thousands of compute-intensive simulations, making applications to complex systems prohibitively expensive. We introduce Neural Network Bayesian Algorithm Execution (NN-BAX), a framework that jointly learns the energy landscape and the MEP. NN-BAX sequentially fine-tunes a foundation model by actively selecting samples targeted at improving the MEP. Tested on Lennard-Jones and Embedded Atom Method systems, our approach achieves a one to two order of magnitude reduction in energy and force evaluations with negligible loss in MEP accuracy and demonstrates scalability to >100-dimensional systems. This work is therefore a promising step towards removing the computational barrier for MEP discovery in scientifically relevant systems, suggesting that weeks-long calculations may be achieved in hours or days with minimal loss in accuracy.
X-ray Free Electron Lasers (X-FELs) operate in a wide range of lasing configurations for a broad variety of scientific applications at ultrafast time-scales such as structural biology, materials science, and atomic and molecular physics. Shot-by-shot characterization of the X-FEL pulses is crucial for analysis of many experiments as well as tuning the X-FEL performance. However, for the weak pulses found in advanced configurations, e.g. those needed for coherent, two-pulse studies of quantum materials, there is no current method for reliably resolving pulse profiles. Here we show that a physics-based U-net model can reconstruct the individual pulse power profiles for sub-picosecond pulse separation without the need for simulations. Using experimental data from weak X-FEL pulse pairs, we demonstrate we can learn the pulse characteristics on a shot-by-shot basis when conventional methods fail.
With the increasing brightness of Light sources, including the Diffraction-Limited brightness upgrade of APS and the high-repetition-rate upgrade of LCLS, the proposed experiments therein are becoming increasingly complex. For instance, experiments at LCLS-II-HE will require the X-ray beam to be within a fraction of a micron in diameter, with pointing stability of a few nanoradians, at the end of a kilometer-long electron accelerator, a hundred-meter-long undulator section, and tens of meters long X-ray optics. This enhancement of brightness will increase the data production rate to rival the largest data generators in the world. Without real-time active feedback control and an optimized pipeline to transform measurements to scientific information and insights, researchers will drown in a deluge of mostly useless data, and fail to extract the highly sophisticated insights that the recent brightness upgrades promise. In this article, we outline the strategy we are developing at SLAC to implement Machine Learning driven optimization, automation and real-time knowledge extraction from the electron-injector at the start of the electron accelerator, to the multidimensional X-ray optical systems, and till the experimental endstations and the high readout rate, multi-megapixel detectors at LCLS to deliver the design performance to the users. This is illustrated via examples from Accelerator, Optics and End User applications.
Cryogenic electron microscopy (cryo-EM) has emerged as the method of choice to characterize the structural variability of biomolecules at near-atomic resolution. We present a reconstruction approach that eliminates the need for post-hoc atomic model fitting in 3D maps by deforming a given atomic model along its normal modes directly against the 2D data. This end-to-end approach inherently reduces the risk of error propagation while increasing interpretability of resulting structural ensembles. ### Competing Interest Statement The authors have declared no competing interest. United States Department of Energy, DE-AC02-76SF00515, the LCLS seed grant "From Atomic Models to Noisy Images and Back Again: End-to-end Differentiable SPI simulators for CryoEM and XFELs" (PI: G.W.) SLAC National Accelerator Laboratory, LDRD project "AtomicSPI: Learning atomic scale biomolecular dynamics from single-particle imaging data" (PI: F.P., Y.N.) SLAC National Accelerator Laboratory, Machine Learning Initiative United States Department of Health and Human Services, https://ror.org/033jnv181, National Institutes of Health (NIH) grant No. 1R01GM144965-02
X-ray free electron lasers (X-FELs) produce ultrafast pulses in a wide range of lasing configurations, supporting a wide variety of scientific applications, including structural biology, materials science, and atomic and molecular physics. Shot-by-shot characterization of the X-FEL pulses is crucial for the analysis of experiments as well as for tuning the X-FEL performance. However, for the weak pulses found in advanced configurations, e.g., those needed for monochromatic, two-pulse studies of quantum materials, there is no current method for reliably resolving pulse profiles. Here, we show that an interpretable neural network (NN) model can reconstruct the individual pulse power profiles for sub-picosecond pulse separation without the need for simulations. Using experimental data from low-signal X-FEL pulse pairs, we demonstrate a NN can learn the pulse characteristics on a shot-by-shot basis when conventional methods fail. This new method enables the characterization of weak pulses-a condition expected to dominate future experimental configurations such as at the Linac Coherent Light Source-II-and opens the door to a wide range of new experiments.
Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.
Tuning particle accelerators is a challenging and time-consuming task that can be automated and carried out efficiently using suitable optimization algorithms, such as model-based Bayesian optimization techniques. One of the major advantages of Bayesian algorithms is the ability to incorporate prior information about beam physics and historical behavior into the model used to make control decisions. In this work, we examine incorporating prior accelerator physics information into Bayesian optimization algorithms by utilizing fast executing, neural network models trained on simulated or historical datasets as prior mean functions in Gaussian process models. We show that in ideal cases, this technique substantially increases convergence speed to optimal solutions in high-dimensional tuning parameter spaces. Additionally, we demonstrate that even in non-ideal cases, where prior models of beam dynamics do not exactly match experimental conditions, the use of this technique can still enhance convergence speed. Finally, we demonstrate how these methods can be used to improve optimization in practical applications, such as transferring information gained from beam dynamics simulations to online control of the LCLS injector, and transferring knowledge gained from experimental measurements across different operating modes, such as accelerating different ion species at the ATLAS heavy ion accelerator.
Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F<^>beta, out of analogy to the supervised classification F beta statistic. CoAD uses F<^>beta to train an anomaly detection algorithm on unlabeled data, based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.