Stellar spectra emulators often rely on large grids and tend to reach a plateau in emulation accuracy, leading to significant systematic errors when inferring stellar properties. Our study explores the use of Transformer models to capture long-range information in spectra, comparing their performance to the Payne emulator (a fully connected multilayer perceptron), an expanded version of The Payne, and a convolutional-based emulator. We tested these models on synthetic spectral grids, evaluating their performance by analyzing emulation residuals and assessing the quality of spectral parameter inference. The newly introduced TransformerPayne emulator outperformed all other tested models, achieving a mean absolute error (MAE) of approximately 0.15% when trained on the full grid. The most significant improvements were observed in grids containing between 1000 and 10,000 spectra, with TransformerPayne showing 2–5 times better performance than the scaled-up version of The Payne. Additionally, TransformerPayne demonstrated superior fine-tuning capabilities, allowing for pretraining on one spectral model grid before transferring to another. This fine-tuning approach enabled up to a 10-fold reduction in training grid size compared to models trained from scratch. Analysis of TransformerPayne's attention maps revealed that they encode interpretable features common across many spectral lines of chosen elements. While scaling up The Payne to a larger network reduced its MAE from 1.2% to 0.3% when trained on the full data set, TransformerPayne consistently achieved the lowest MAE across all tests. The inductive biases of the TransformerPayne emulator enhance accuracy, data efficiency, and interpretability for spectral emulation compared to existing methods.
We present the MULTIMODAL UNIVERSE, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, the MULTIMODAL UNIVERSE contains hundreds of millions of astronomical observations, constituting 100 TB of multi-channel and hyper-spectral images, spectra, multivariate time series, as well as a wide variety of associated scientific measurements and "metadata". In addition, we include a range of benchmark tasks representative of standard practices for machine learning methods in astrophysics. This massive dataset will enable the development of large multi-modal models specifically targeted towards scientific applications. All codes used to compile the MULTIMODAL UNIVERSE and a description of how to access the data is available at https://github.com/MultimodalUniverse/MultimodalUniverse
We explore the potential of enhancing LLM performance in astronomy-focused question-answering through targeted, continual pre-training. By employing a compact 7B-parameter LLaMA-2 model and focusing exclusively on a curated set of astronomy corpora—comprising abstracts, introductions, and conclusions—we achieve notable improvements in specialized topic comprehension. While general LLMs like GPT-4 excel in broader question-answering scenarios due to superior reasoning capabilities, our findings suggest that continual pre-training with limited resources can still enhance model performance on specialized topics. Additionally, we present an extension of AstroLLaMA: the fine-tuning of the 7B LLaMA model on a domain-specific conversational data set, culminating in the release of the chat-enabled AstroLLaMA for community use. Comprehensive quantitative benchmarking is currently in progress and will be detailed in an upcoming full paper. The model, AstroLLaMA-Chat, is now available at https://huggingface.co/universeTBD , providing the first open-source conversational AI tool tailored for the astronomy community.
Abstract We present the Multimodal Universe, a new framework collating over 100 TB of multimodal astronomical data for its first release, spanning images, spectra, time series, tabular and hyper-spectral data. This unified collection enables a wide variety of machine learning (ML) applications and research across astronomical domains. The dataset brings together observations from multiple surveys, facilities, and wavelength regimes, providing standardized access to diverse data types. By providing uniform access to this diverse data, the Multimodal Universe aims to accelerate the development of ML methods for observational astronomy that can work across the large differences in astronomical datasets. The framework is actively supported and is designed to be extended whilst enforcing minimal self consistent conventions making contributing data as simple and practical as possible.
The exponential growth of astronomical literature poses significant challenges for researchers navigating and synthesizing general insights or even domain-specific knowledge. We present pathfinder, a machine learning framework designed to enable literature review and knowledge discovery in astronomy, focusing on semantic searching with natural language instead of syntactic searches with keywords. Utilizing state-of-the-art large language models (LLMs) and a corpus of 385,166 peer-reviewed papers from the Astrophysics Data System, pathfinder offers an innovative approach to scientific inquiry and literature exploration. Our framework couples advanced retrieval techniques with LLM-based synthesis to search astronomical literature by semantic context as a complement to currently existing methods that use keywords or citation graphs. It addresses complexities of jargon, named entities, and temporal aspects through time-based and citation-based weighting schemes. We demonstrate the tool's versatility through case studies, showcasing its application in various research scenarios. The system's performance is evaluated using custom benchmarks, including single-paper and multipaper tasks. Beyond literature review, pathfinder offers unique capabilities for reformatting answers in ways that are accessible to various audiences (e.g., in a different language or as simplified text), visualizing research landscapes, and tracking the impact of observatories and methodologies. This tool represents a significant advancement in applying artificial intelligence to astronomical research, aiding researchers at all career stages in navigating modern astronomy literature.
Zoobot is a Python package for measuring the detailed appearance of galaxies in telescope images using deep learning.Zoobot is aimed at astronomers who want to solve a galaxy image task such as finding merging galaxies or counting spiral arms.Astronomers can use Zoobot to adapt (finetune) pretrained deep learning models to solve their task.These finetuned models perform better and require far fewer new labels than training from scratch (Walmsley, Slijepcevic, et al., 2022).
Astrophysical explorations are underpinned by large-scale stellar spectroscopy surveys, necessitating a paradigm shift in spectral fitting techniques. Our study proposes three enhancements to transcend the limitations of the current spectral emulation models. We implement an attention-based emulator, adept at unveiling long-range information between wavelength pixels. We leverage a domain-specific fine-tuning strategy where the model is pre-trained on spectra with fixed stellar parameters and variable elemental abundances, followed by fine-tuning on the entire domain. Moreover, by treating wavelength as an autonomous model parameter, akin to neural radiance fields, the model can generate spectra on any wavelength grid. In the case with a training set of O(1000), our approach exceeds current leading methods by a factor of 5-10 across all metrics.
Large language models excel in many human-language tasks but often falter in highly specialized domains like scholarly astronomy. To bridge this gap, we introduce AstroLLaMA, a 7-billion-parameter model fine-tuned from LLaMA-2 using over 300,000 astronomy abstracts from arXiv. Optimized for traditional causal language modeling, AstroLLaMA achieves a 30% lower perplexity than Llama-2, showing marked domain adaptation. Our model generates more insightful and scientifically relevant text completions and embedding extraction than state-of-the-arts foundation models despite having significantly fewer parameters. AstroLLaMA serves as a robust, domain-specific model with broad fine-tuning potential. Its public release aims to spur astronomy-focused research, including automatic paper summarization and conversational agent development.
FU Orionis objects (FUors) are eruptive young stars, which exhibit outbursts that last from decades to a century. Due to the duration of their outbursts, and to the fact that only about two dozens of such sources are known, information on the end of their outbursts is limited. Here we analyse follow-up photometry and spectroscopy of Gaia21elv, a young stellar object, which had a several decades long outburst. It was reported as a Gaia science alert due to its recent fading by more than a magnitude. To study the fading of the source and look for signatures characteristic of FUors, we have obtained follow-up near infrared (NIR) spectra using Gemini South/IGRINS, and both optical and NIR spectra using VLT/X-SHOOTER. The spectra at both epochs show typical FUor signatures, such as a triangular shaped $H$-band continuum, absorption-line dominated spectrum, and P Cygni profiles. In addition to the typical FUor signatures, [OI], [FeII], and [SII] were detected, suggesting the presence of a jet or disk wind. Fitting the spectral energy distributions with an accretion disc model suggests a decrease of the accretion rate between the brightest and faintest states. The rapid fading of the source in 2021 was most likely dominated by an increase of circumstellar extinction. The spectroscopy presented here confirms that Gaia21elv is a classical FUor, the third such object discovered among the Gaia science alerts.
ABSTRACT We use near-infrared photometry and astrometry from the VISTA Variables in the Via Lactea (VVV) survey to analyse microlensing events containing annual microlensing parallax information. These events are located in highly extincted and low-latitude regions of the Galactic bulge typically off-limits to optical microlensing surveys. We fit a catalogue of 1959 events previously found in the VVV and extract 21 microlensing parallax candidates. The fitting is done using nested sampling to automatically characterize the multimodal and degenerate posterior distributions of the annual microlensing parallax signal. We compute the probability density in lens mass-distance using the source proper motion and a Galactic model of disc and bulge deflectors. By comparing the expected flux from a main sequence lens to the baseline magnitude and blending parameter, we identify four candidates which have probability >50 per cent that the lens is dark. The strongest candidate corresponds to a nearby (≈0.78 kpc), medium-mass ($1.46^{+1.13}_{-0.71} \ M_{\odot }$) dark remnant as lens. In the next strongest, the lens is located at heliocentric distance ≈5.3 kpc. It is a dark remnant with a mass of $1.63^{+1.15}_{-0.70} \ M_{\odot }$. Both of those candidates are most likely neutron stars, though possibly high-mass white dwarfs. The last two events may also be caused by dark remnants, though we are unable to rule out other possibilities because of limitations in the data. We are also demonstrating future possibilities of studying similar events with the Roman Space Telescopeby modelling a mock dataset of Roman photometry and astrometry for an event resembling our strongest candidate.
Massive galactic lenses with large Einstein Radii should cause a measurable astrometric microlensing effect, i.e. the light centroid shift due to the motion of the two images. Such a shift in the position of a background star due to microlensing was not included in the $Gaia$ astrometric model, therefore significant deviation should cause $Gaia$'s astrometric parameters to be determined incorrectly. Here we studied one of the photometric microlensing events reported in the $Gaia$ DR3, GaiaDR3-ULENS-001, for which poor goodness of $Gaia$ fit and erroneous parallax could indicate the presence of the astrometric microlensing signal. Based on the photometric microlensing model, we simulated $Gaia$ astrometric time-series with the astrometric microlensing effect added. We found that including microlensing with the angular Einstein Radius of $\theta_{\rm E}$ = $2.60^{+0.21}_{-0.24}$ mas ($2.47^{+0.28}_{-0.24}$ mas) assuming positive (negative) impact parameter $u_0$ reproduces well the astrometric quantitie reported by $Gaia$. We estimate the mass of the lens to $1.00^{+0.23}_{-0.18}$ $M_\odot$ ($0.70^{+0.17}_{-0.13}$ $M_\odot$) and its distance to $0.90^{+0.14}_{-0.11}$ kpc ($0.69^{+0.13}_{-0.09}$ kpc), proposing the lens could be a nearby isolated white dwarf.