Despite the growing application of Large Language Models (LLMs) to theoretical physics, there is little academic exploration into how domain-specific physics reasoning ability develops while training these models. To investigate this, we perform the first academic fine-tuning study of small (7B-parameter) reasoning models dedicated specifically to theoretical physics. Because open-source verifiable training data required to train such capabilities is scarce, we developed a robust data generation pipeline that can both create synthetic problems and make existing human-authored problems suitable for model training. Selecting Quantum Field Theory (QFT) as our primary domain, we generated over 2,500 synthetic problems alongside a curated collection of human-adapted problems sourced from arXiv and standard pedagogical resources. We conduct both Reinforcement Learning (RL) and Supervised Fine-Tuning (SFT) experiments, benchmarking performance gains as well as generalization to other physics domains. We perform an extensive analysis of model chains-of-though before and after fine-tuning, to understand how reasoning errors evolve during RL and SFT. Finally, we publicly release our data pipeline, verifiable QFT training data, and ∼200M tokens of QFT reasoning traces.
For many analyses in cosmology it is necessary to reconstruct the likely distribution of unobserved fields, such as dark matter or baryons, from observed luminous tracers. The dominant approach in cosmology has been to use the so-called halo model, which assumes radially symmetric profiles centered around luminous tracers such as galaxies. More recently, field-level machine learning methods have been proposed that can learn to estimate the unobserved field after being trained on simulations. However, it is unclear whether machine learning methods indeed significantly improve over linear methods or the halo model. In this paper we make a systematic comparison of different approaches to reconstruct dark matter and baryons from galaxy data using the CAMELS simulations. We find the best results using a combined GNN-CNN approach. We also provide a general analysis and visualization of the relationship of matter, baryons, halos and galaxies in these simulations to interpret our results.
Several statistics have been proposed for measuring the kinetic Sunyaev-Zeldovich (kSZ) effect by combining the small-scale CMB with galaxy surveys. We review five such statistics, and show that they are all mathematically equivalent to the optimal bispectrum estimator of type (ggT). Reinterpreting these kSZ statistics as special cases of bispectrum estimation makes many aspects transparent, for example optimally weighting the estimator or incorporating photometric redshift errors. We analyze the information content of the bispectrum and show that there are two observables: the small-scale galaxy-electron power spectrum Pge(kS), and the large-scale galaxy-velocity power spectrum Pgv(k). The cosmological constraining power of the kSZ effect arises from its sensitivity to fluctuations on large length scales, where its effective noise level can be much better than galaxy surveys.
We present a new method to constrain local primordial non-Gaussianity using the large-scale modulation of the local lensing power spectrum. Our work extends our recently proposed pi-field method for primordial non-Gaussianity estimation to spherical coordinates and applies it to galaxy lensing. Our approach is computationally efficient and only requires binned multipole power spectra Cl(z1; z2) on large scales, as well as their covariance. Our method is simpler to implement than a full bispectrum estimator, but still contains the full squeezed-limit information. We validate our model using a suite of N-body simulations and demonstrate its accuracy in recovering the fNL values. We then perform a Fisher forecast for an Legacy Survey of Space and Time-like weak lensing survey, finding 6fNL similar or equal to 44. Our approach readily combines with other fNL-sensitive fields such as kinetic Sunyaev-Zel'dovich velocity reconstruction and clustering-based pi fields, for a future combined fNL estimator using various large-scale galaxy and CMB observables.
We present a method to estimate non-Gaussian power spectrum covariance matrices by directly measuring the response of the small-scale power spectrum to long-wavelength perturbations via bispectrum and trispectrum estimators. Specifically, we derive estimators for the complete non-Gaussian matter power spectrum covariance, including the super-sample contribution, in terms of the squeezed bispectrum and collapsed trispectrum of the underlying density field. We apply these estimators to the Quijote simulations, and recover unbiased estimates of the small-scale (k≳ 0.15 h/ Mpc) matter power spectrum covariance at the percent level using only 25 simulations - comparable to the precision of the sample covariance estimated using 5,000 simulations. This technique significantly reduces the number of simulations needed to estimate power spectrum covariances and opens the possibility of inferring power spectrum covariances directly from survey data, enabling stringent tests of simulations and, potentially, power spectrum analyses that do not rely on external covariance matrices.
It was recently shown that neural networks can be combined with the analytic method of scale-dependent bias to obtain a measurement of local primordial non-Gaussianity, which is optimal in the squeezed limit that dominates the signal-to-noise. The method is robust to nonlinear physics, but also inherits the statistical precision offered by neural networks applied to very nonlinear scales. In prior work, we assumed that the neural network has access to the full matter distribution. In this work, we apply our method to halos. We first describe a novel two-field formalism that is optimal even when the matter distribution is not observed. We show that any N halo fields can be compressed to two fields without losing information, and we obtain loss functions to learn these fields. We then apply the method to high-resolution AbacusSummit and AbacusPNG simulations. In the present work, the two neural networks observe the local population statistics, in particular, the halo mass and concentration distribution in a patch of the sky. While the traditional mass-binned halo analysis is optimal in practice without further halo properties on AbacusPNG, our novel formalism easily allows us to include additional halo properties such as the halo concentration, which can improve fNL constraints by a factor of a few. We also explore whether shot noise can be lowered with machine learning compared to a traditional reconstruction, finding no improvement for our simulation parameters.
Large language models (LLMs) have shown strong capabilities in complex reasoning, and test-time scaling techniques can enhance their performance with comparably low cost. Many of these methods have been developed and evaluated on mathematical reasoning benchmarks such as AIME. This paper investigates whether the lessons learned from these benchmarks generalize to the domain of advanced theoretical physics. We evaluate a range of common test-time scaling methods on the TPBench physics dataset and compare their effectiveness with results on AIME. To better leverage the structure of physics problems, we develop a novel, symbolic weak-verifier framework to improve parallel scaling results. Our empirical results demonstrate that this method significantly outperforms existing test-time scaling approaches on TPBench. We also evaluate our method on AIME, confirming its effectiveness in solving advanced mathematical problems. Our findings highlight the power of step-wise symbolic verification for tackling complex scientific problems.
We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics (TP), focusing on high-energy theory and cosmology. The first iteration of our benchmark consists of 57 problems of varying difficulty, from undergraduate to research level. These problems are novel in the sense that they do not come from public problem collections. We evaluate our data set on various open and closed language models, including o3-mini, o1, DeepSeek-R1, GPT-4o and versions of Llama and Qwen. While we find impressive progress in model performance with the most recent models, our research-level difficulty problems are mostly unsolved. We address challenges of auto-verifiability and grading, and discuss common failure modes. While currently state-of-the art models are still of limited use for researchers, our results show that AI assisted TP research may become possible in the near future. We discuss the main obstacles towards this goal and possible strategies to overcome them. The public problems and solutions, results for various models, and updates to the data set and score distribution, are available on the website of the dataset tpbench.org.
We implement a novel formalism to constrain primordial non-Gaussianity of the local type from the large-scale modulation of the small-scale power spectrum. Our approach combines information about primordial non-Gaussianity contained in the squeezed bispectrum and the collapsed trispectrum of large-scale structure together in a computationally amenable and consistent way, while avoiding the need to model complicated covariances of higher N-point functions. This work generalizes our recent work, which used a neural network estimate of local power, to the more conventional local power spectrum statistics, and explores using both matter field and halo catalogues from the Quijote simulations. We find that higher N-point functions of the matter field can provide strong constraints on f_NL, but higher N-point functions of the halo field, at the halo density of Quijote, only marginally improve constraints from the two-point function.
We perform kinetic Sunyaev-Zel'dovich (kSZ) velocity reconstruction on data from ACT DR6 and DESI-LS DR9. To estimate the cross-power between kSZ velocity reconstruction and galaxy density, we make use of a novel quadratic maximum likelihood QML power spectrum estimator implementation in red-shift binned spherical coordinates. We find a detection of the kSZ signal from the cross-correlation between the estimated velocity field and the large-scale galaxy field of 11.7 σ. We estimate an amplitude A=0.39 ± 0.04 of the kSZ signal with respect to a halo model prediction, possibly indicating a high feedback in massive halos, in agreement with previous studies. Our result demonstrates the feasibility of an optimal QML pipeline at the resolution required for this analysis, and will be a powerful tool for kSZ cosmology with upcoming high-resolution surveys.
We present a novel implementation for the quadratic maximum likelihood (QML) power spectrum estimator for multiple correlated scalar fields on the sphere. Our estimator supports arbitrary binning in redshift and multipoles ℓ and includes cross-correlations of redshift bins. It implements a fully optimal analysis with a pixel-wise covariance model. We implement a number of optimizations which make the estimator and associated covariance matrix computationally tractable for a low-ℓ analysis, suitable for example for kSZ velocity reconstruction or primordial non-Gaussianity from scale-dependent bias analyses. We validate our estimator extensively on simulations and compare its features and precision with the common pseudo-C_ℓ method, showing significant gains at large scales. We make our code publicly available. In a companion paper, we apply the estimator to kSZ velocity reconstruction using data from ACT and DESI Legacy Survey and construct full set of QML estimators on 40 correlated fields up to N_side= 32 in timescale of an hour on a single 24-core CPU requiring <256 Gb RAM, demonstrating the performance of the code.
The existence, properties, and dynamics of the dark sectors of our universe pose fundamental challenges to our current model of physics, and large-scale astronomical surveys may be our only hope to unravel these long-standing mysteries. In this white paper, we describe the science motivation, instrumentation, and survey plan for the next-generation spectroscopic observatory, the Stage-5 Spectroscopic Experiment (Spec-S5). Spec-S5 is a new all-sky spectroscopic instrument optimized to efficiently carry out cosmological surveys of unprecedented scale and precision. The baseline plan for Spec-S5 involves upgrading two existing 4-m telescopes to new 6-m wide-field facilities, each with a highly multiplexed spectroscopic instrument capable of simultaneously measuring the spectra of 13,000 astronomical targets. Spec-S5, which builds and improves on the hardware used for previous cosmology experiments, represents a cost-effective and rapid approach to realizing a more than 10× gain in spectroscopic capability compared to the current state-of-the-art represented by the Dark Energy Spectroscopic Instrument project (DESI). Spec-S5 will provide a critical scientific capability in the post-Rubin and post-DESI era for advancing cosmology, fundamental physics, and astrophysics in the 2030s.
We develop an optimization-based maximum likelihood approach to analyze the cross-correlation of the Cosmic Microwave Background (CMB) and large-scale structure induced by the kinetic Sunyaev-Zeldovich (kSZ) effect. Our main goal is to reconstruct the radial velocity field of the universe. While the existing quadratic estimator (QE) is statistically optimal for current and near-term experiments, the likelihood can extract more signal-to-noise in the future. Our likelihood formulation has further advantages over the QE, such as the possibility of jointly fitting cosmological and astrophysical parameters and the possibility of unifying several different kSZ analyses. We implement an auto-differentiable likelihood pipeline in JAX, which is computationally tractable for a realistic survey size and resolution, and evaluate it on the Agora simulation. We also implement a machine learning-based estimate of the electron density given an observed galaxy distribution, which can increase the signal-to-noise for both the QE and the likelihood method.
We develop a hybrid GNN-CNN architecture for the reconstruction of 3-dimensional continuous cosmological matter fields from discrete point clouds, provided by observed galaxy catalogs. Using the CAMELS hydrodynamical cosmological simulations we demonstrate that the proposed architecture allows for an accurate reconstruction of both the dark matter and electron density given observed galaxies and their features. Our approach includes a learned grid assignment scheme that improves over the traditional cloud-in-cell method. Our method can improve cosmological analyses in situations where non-luminous (and thus unobservable) continuous fields need to be estimated from luminous (observable) discrete point cloud tracers.
It was recently shown that neural networks can be combined with the analytic method of scale-dependent bias to obtain a measurement of local primordial non-Gaussianity, which is optimal in the squeezed limit that dominates the signal-to-noise. The method is robust to non-linear physics, but also inherits the statistical precision offered by neural networks applied to very non-linear scales. In prior work, we assumed that the neural network has access to the full matter distribution. In this work, we apply our method to halos. We first describe a novel two-field formalism that is optimal even when the matter distribution is not observed. We show that any N halo fields can be compressed to two fields without losing information, and obtain optimal loss functions to learn these fields. We then apply the method to high-resolution AbacusSummit and AbacusPNG simulations. In the present work, the two neural networks observe the local population statistics, in particular the halo mass and concentration distribution in a patch of the sky. While the traditional mass-binned halo analysis is optimal in practice without further halo properties on AbacusPNG, our novel formalism easily allows to include additional halo properties such as the halo concentration, which can improve f_NL constraints by a factor of a few. We also explore whether shot noise can be lowered with machine learning compared to a traditional reconstruction, finding no improvement for our simulation parameters.
Fast radio bursts (FRBs) are brief, energetic, typically extragalactic flashes of radio emission whose progenitors are largely unknown. Although studying the FRB population is essential for understanding how these astrophysical phenomena occur, such studies have been difficult to conduct without large numbers of FRBs and characterizable observational biases. Using the recently released catalog of 536 FRBs published by the Canadian Hydrogen Intensity Mapping Experiment/Fast Radio Burst (CHIME/FRB) collaboration, we present a study of the FRB population that also calibrates for selection effects. Assuming a Schechter function, we infer a characteristic energy cut-off of E char = 2.38 − 1.64 + 5.35 × 10 41 erg and a differential power-law index of γ = − 1.3 − 0.4 + 0.7 . Simultaneously, we infer a volumetric rate of [ 7.3 − 3.8 + 8.8 (stat.) − 1.8 + 2.0 ( sys . ) ] × 10 4 Gpc −3 yr −1 above a pivot energy of 10 39 erg and below a scattering timescale of 10 ms at 600 MHz, and find we cannot significantly constrain the cosmic evolution of the FRB population with star-formation rate. Modeling the host’s dispersion measure (DM) contribution as a log-normal distribution and assuming a total Galactic contribution of 80 pc cm −3 , we find a median value of DM host = 84 − 49 + 69 pc cm −3 , comparable with values typically used in the literature. Proposed models for FRB progenitors should be consistent with the energetics and abundances of the full FRB population predicted by our results. Finally, we infer the redshift distribution of FRBs detected with CHIME, which will be tested with the localizations and redshifts enabled by the upcoming CHIME/FRB Outriggers project.
Here we present a deep learning-based image analysis platform (DLAP), tailored to autonomously quantify cell numbers, and fluorescence signals within cellular compartments, derived from RNAscope or immunohistochemistry. We utilised DLAP to analyse subtypes of tyrosine hydroxylase (TH)-positive dopaminergic midbrain neurons in mouse and human brain-sections. These neurons modulate complex behaviour, and are differentially affected in Parkinson’s and other diseases. DLAP allows the analysis of large cell numbers, and facilitates the identification of small cellular subpopulations. Using DLAP, we identified a small subpopulation of TH-positive neurons (~5%), mainly located in the very lateral Substantia nigra (SN), that was immunofluorescence-negative for the plasmalemmal dopamine transporter (DAT), with ~40% smaller cell bodies. These neurons were negative for aldehyde dehydrogenase 1A1, with a lower co-expression rate for dopamine-D2-autoreceptors, but a ~7-fold higher likelihood of calbindin-d28k co-expression (~70%). These results have important implications, as DAT is crucial for dopamine signalling, and is commonly used as a marker for dopaminergic SN neurons.
The bright millisecond-duration radio burst from the Galactic magnetar SGR 1935+2154 in 2020 April was a landmark event, demonstrating that at least some fast radio burst (FRB) sources could be magnetars. The two-component burst was temporally coincident with peaks observed within a contemporaneous short X-ray burst envelope, marking the first instance where FRB-like bursts were observed to coincide with X-ray counterparts. In this study, we detail five new radio burst detections from SGR 1935+2154, observed by the CHIME/FRB instrument between October 2020 and December 2022. We develop a fast and efficient Bayesian inference pipeline that incorporates state-of-the-art Markov chain Monte Carlo techniques and use it to model the intensity data of these bursts under a flexible burst model. We revisit the 2020 April burst and corroborate that both the radio sub-components lead the corresponding peaks in their high-energy counterparts. For a burst observed in 2022 October, we find that our estimated radio pulse arrival time is contemporaneous with a short X-ray burst detected by GECAM and HEBS, and Konus-Wind and is consistent with the arrival time of a radio burst detected by GBT. We present flux and fluence estimates for all five bursts, employing an improved estimator for bursts detected in the side-lobes. We also present upper limits on radio emission for X-ray emission sources which were within CHIME/FRB's field-of-view at trigger time. Finally, we present our exposure and sensitivity analysis and estimate the Poisson rate for FRB-like events from SGR 1935+2154 to be $0.005^{+0.082}_{-0.004}$ events/day above a fluence of $10~\mathrm{kJy~ms}$ during the interval from 28 August 2018 to 1 December 2022, although we note this was measured during a time of great X-ray activity from the source.
The CHIME/FRB Collaboration, Mandana Amiri , Bridget C. Andersen , Kevin Bandura , Sabrina Berger , Mohit Bhardwaj , Michelle M. Boyce , P. J. Boyle , Charanjot Brar , Daniela Breitman , Tomas Cassanelli , Pragya Chawla , Tianyue Chen , J.-F. Cliche , Amanda Cook , Davor Cubranic , Alice P. Curtin , Meiling Deng , Matt Dobbs , Fengqiu (Adam) Dong , Gwendolyn Eadie , Mateus Fandino , Emmanuel Fonseca , B. M. Gaensler , Utkarsh Giri , Deborah C. Good , Mark Halpern , Alex S. Hill , Gary Hinshaw , Alexander Josephy , Jane F. Kaczmarek , Zarif Kader , Joseph W. Kania , Victoria M. Kaspi , T. L. Landecker , Dustin Lang , Calvin Leung , Dongzi Li , Hsiu-Hsien Lin , Kiyoshi W. Masui , Ryan Mckinven , Juan Mena-Parra , Marcus Merryfield , Bradley W. Meyers , Daniele Michilli , Nikola Milutinovic , Arash Mirhosseini , Moritz Münchmeyer , Arun Naidu , Laura Newburgh , Cherry Ng , Chitrang Patel , Ue-Li Pen , Emily Petroff , Tristan Pinsonneault-Marotte , Ziggy Pleunis , Masoud Rafiei-Ravandi , Mubdi Rahman , Scott M. Ransom , Andre Renard , Pranav Sanghavi , Paul Scholz , J. Richard Shaw , Kaitlyn Shin , Seth R. Siegel , Andrew E. Sikora , Saurabh Singh , Kendrick M. Smith , Ingrid Stairs , Chia Min Tan , S. P. Tendulkar , Keith Vanderlinde , Haochen Wang , Dallas Wulf , and A. V. Zwaniga 1 Department of Physics and Astronomy, University of British Columbia, 6224 Agricultural Road, Vancouver, BC V6T 1Z1, Canada 2 Department of Physics, McGill University, 3600 rue University, Montréal, QC H3A 2T8, Canada 3 McGill Space Institute, McGill University, 3550 rue University, Montréal, QC H3A 2A7, Canada 4 Lane Department of Computer Science and Electrical Engineering, 1220 Evansdale Drive, PO Box 6109, Morgantown, WV 26506, USA 5 Center for Gravitational Waves and Cosmology, West Virginia University, Chestnut Ridge Research Building, Morgantown, WV 26505, USA 6 Department of Physics and Astronomy, University of Manitoba, Winnipeg, MB R3T 2N2, Canada 7 Department of Physics, University of Toronto, 60 St. George Street, Toronto, ON M5S 1A7, Canada 8 Dunlap Institute for Astronomy & Astrophysics, University of Toronto, 50 St. George Street, Toronto, ON M5S 3H4, Canada 9 David A. Dunlap Department of Astronomy & Astrophysics, University of Toronto, 50 St. George Street, Toronto, ON M5S 3H4, Canada 10 MIT Kavli Institute for Astrophysics and Space Research, Massachusetts Institute of Technology, 77 Massachusetts Ave., Cambridge, MA 02139, USA 11 Perimeter Institute for Theoretical Physics, 31 Caroline Street N, Waterloo, ON N25 2YL, Canada 12 Dominion Radio Astrophysical Observatory, Herzberg Research Centre for Astronomy and Astrophysics, National Research Council Canada, PO Box 248,