A significant number of Lyman-break galaxies (LBGs) with redshifts 3 less than or similar to z less than or similar to 5 are expected to be observed by the upcoming Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). This will enable us to probe the Universe at higher redshifts than is currently possible with cosmological galaxy clustering and weak lensing surveys. However, accurate inference of cosmological parameters requires precise knowledge of the redshift distributions of selected galaxies, where the number of faint objects expected from LSST alone will make spectroscopic based methods of determining these distributions extremely challenging. To overcome this difficulty, it may be possible to leverage the information in the large volume of photometric data alone to precisely infer these distributions. This could be facilitated using forward models, where in this paper we use stellar population synthesis (SPS) to estimate uncertainties on LBG redshift distributions for a 10 yr LSST (LSSTY10) survey. We characterize some of the modelling uncertainties inherent to SPS by introducing a flexible parametrization of the galaxy population prior, informed by observations of the galaxy stellar mass function (GSMF) and cosmic star formation rate density (CSFRD). These uncertainties are subsequently marginalised over and propagated to cosmological constraints in a Fisher forecast, leveraging galaxy clustering and lensing of the cosmic microwave background (CMB). Assuming a known dust attenuation model for LBGs, we forecast constraints on the sigma(8) parameter comparable to Planck CMB constraints.
We study star formation over similar to 12 Gyr using pop-cosmos , a generative model trained on 26-band photometry of similar to 420 000 COSMOS2020 galaxies ( Spitzer IRAC Ch.1 < 26 ). The model learns distributions over 16 stellar population synthesis parameters via score-based diffusion, matching observed colours and magnitudes. We use pop-cosmos to compute the cosmic star formation rate density (SFRD) to z = 3 . 5 by directly integrating individual galaxy SFRs. The SFRD peaks at z = 1.3 +/- 0.1 , Delta z similar or equal to 0.6 later than previous canonical estimates, with peak value 0.08 +/- 0.01 M-circle dot yr(-1) Mpc(-3). We classify star-forming (SF) and quiescent (Q) galaxies using specific SFR (sSFR) < 10(-11) yr(-1), comparing with NUVrJ colour selection. The sSFR criterion yields up to 20 per cent smaller Q fractions across 0 < z < 3.5 , with NUVrJ-selected samples contaminated by galaxies with sSFR up to 10(-9) yr(-1). Our sSFR-selected stellar mass function shows a negligible number density of low-mass (less than or similar to 10(9) . 5 M-circle dot) Q galaxies at z similar or equal to 1 , where colour-selection shows a prominent increase. Non-parametric star formation histories around the SFRD peak reveal distinct patterns: SF galaxies show gradually weakening correlations between their recent and earlier SFRs, implying increasingly stochastic star formation towards early epochs. Q galaxies exhibit full correlation ( r > 0 . 95 ) during the most recent similar to 300 Myr, then sharp decorrelation with earlier SF epochs, marking clear quenching transitions. Massive (10 < log(10)(M-*/M-circle dot) < 11 ) galaxies quench on a time-scale of similar to 1 Gyr, with mass assembly concentrated in their first 3.5 Gyr. Finally, active galactic nucleus (AGN) activity (infrared torus luminosity fraction) peaks as massive (similar to 10(10)(.5 )M(circle dot)) galaxies approach the transition between SF and Q states, declining sharply once quiescence is established. This provides evidence that AGN feedback operates in a critical regime during the similar to 1 Gyr quenching transition.
With the next generation of both electromagnetic and gravitational wave observatories beginning to come online, rapid analysis methods for kilonova data are becoming increasingly important in astronomy. Traditional Bayesian parameter estimation using Markov chain Monte Carlo (MCMC) is time-consuming and relies on explicit likelihood approximations that can break down when modeling uncertainties are significant. We develop a simulation-based inference (SBI) framework for kilonova parameter estimation using density-estimation likelihood-free inference. The framework uses a Gaussian process emulator trained on ∼1300 radiative transfer simulations generated with the POSSIS code. We demonstrate that SBI provides a rapid alternative to MCMC for inference with emulators or approximate likelihoods that is robust to emulator uncertainty and likelihood misspecification. On simulated data, the SBI method accurately recovers injected parameters and produces posterior predictive light curves consistent with the data, but the MCMC posterior recovery suffers from systematic bias caused by likelihood misspecification. When analyzing AT2017gfo, the SBI and MCMC methods yield similar light-curve predictions but different posterior distributions, with a subset of the MCMC posteriors piling up at prior boundaries. The likelihood in the MCMC fails to capture the non-Gaussian, correlated structure of the emulator uncertainty, but SBI learns the posterior directly from forward simulations that include the full predictive distribution. Once trained, the SBI framework generates ∼2×10^4 posterior samples in seconds.
This paper presents a search for high redshift galaxies from the Euclid Early Release Observations program `Magnifying Lens.' The 1.5\,$ area covered by the twin Abell lensing cluster fields is comparable in size to the few other deep near-infrared surveys such as COSMOS, and so provides an opportunity to significantly increase known samples of rare UV-bright galaxies at $z UV Beyond their still uncertain role in reionisation, these UV-bright galaxies are ideal laboratories from which to study galaxy formation and constrain the bright-end of the UV luminosity function. Of the sources detected from a combined and NISP detection image, 168 do not have any appreciable VIS/ flux. These objects span a range in spectral colours, separated into two classes: 139 extremely red sources; and 29 Lyman-break galaxy candidates. Best-fit redshifts and spectral templates suggest the former is composed of both $z dusty star-forming galaxies and $z quiescent systems. The latter is composed of more homogeneous Lyman-break galaxies at $z In both cases, contamination by L- and T-type dwarfs cannot be ruled out with images alone. Additional contamination from instrumental persistence is investigated using a novel time series analysis. This work lays the foundation for future searches within the Euclid Deep Fields, where thousands more $z Lyman-break systems and extremely red sources will be identified.
We present the first analysis of the Early Release Observations (ERO) program that targets fields around two lensing clusters, Abell 2390 and Abell 2764. We use imaging data from the Visible instrument (VIS) and the Near-Infrared Spectrometer and Photometer (NISP) to produce photometric catalogs for a total of $ 500\,000$ objects. The imaging data reach a typical depth of $5\ in the range 25.1--25.4 AB in the NISP bands and 27.1--27.3 AB in the VIS band. Using the Lyman-break method in combination with photometric redshifts, we searched for high-redshift galaxies. We identified $30$ Lyman-break galaxy (LBG) candidates at $z>6$ and 139 extremely red sources (ERSs), most of which likely lie at lower redshift. The VIS imaging is deeper than the NISP imaging, which means that we can routinely identify high-redshift Lyman-break galaxies at about a magnitude of 3, which reduces contamination by brown dwarf stars and low-redshift galaxies. The difficulty of spatially resolving most of these sources in 0 $ imaging means that it is difficult to distinguish between galaxies and quasars. Spectroscopic follow-up campaigns of these bright sources will help us to constrain the bright end of the ultraviolet galaxy luminosity function and the quasar luminosity function at $z>6$, and it will constrain the physical nature of these objects. Additionally, we performed a combined strong- and weak-lensing analysis of A2390, and we show that will contribute to constraining the virial mass of galaxy clusters better. We also identify optical and near-infrared counterparts of known $z>0.6$ clusters in these data. These counterparts exhibit strong-lensing features. This establishes that can characterize high-redshift clusters. Finally, we provide a glimpse of the ability of to map the intracluster light out to larger radii than current facilities, which enables us to understand the cluster assembly history better and to map the dark matter distribution. This initial dataset illustrates the diverse spectrum of legacy science that is possible with the survey.
Upcoming surveys such as Euclid, the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST) and the Nancy Grace Roman Telescope (Roman) will detect hundreds of high-redshift (z ≳ 7) quasars, but distinguishing them from the billions of other sources in these catalogues represents a significant data analysis challenge. We address this problem by extending existing selection methods by using both i) Bayesian model comparison on measured fluxes and ii) a likelihood-based goodness-of-fit test on images, which are then combined using the Fβ statistic (where β is a parameter which can be tuned to prioritise completeness). The result is an automated, reproduceable and objective high-redshift quasar selection pipeline. We test this on both simulations and real data from the cross-matched Sloan Digital Sky Survey (SDSS) and UKIRT Infrared Deep Sky Survey (UKIDSS) catalogues. On this cross-matched dataset we achieve an area under the curve (AUC) score of up to 0.81 and an F3 score of up to 0.79 ; or, if the completeness is fixed to be 0.9 then we can obtain an efficiency of 0.15. This is sufficient to be applied to the Euclid, LSST and Roman data when available.
We present an extension of the pop-cosmos model for the evolving galaxy population up to redshift z ∼ 6. The model is trained on distributions of observed colors and magnitudes, from 26-band photometry of ∼420,000 galaxies in the COSMOS2020 catalog with Spitzer IRAC Channel 1 < 26 mag. The generative model includes a flexible distribution over 16 stellar population synthesis (SPS) parameters, and a depth-dependent photometric uncertainty model, both represented using score-based diffusion models. We use the trained model to predict scaling relationships for the galaxy population, such as the stellar mass function, star-forming main sequence, and gas phase and stellar metallicity versus mass relations, demonstrating reasonable to excellent agreement with previously published results. We explore the connection between mid-infrared emission from active galactic nuclei (AGN) and star formation rate, finding high AGN activity for galaxies above the star-forming main sequence at 1 ≲ z ≲ 2. Using the trained population model as a prior distribution, we perform inference of the redshifts and SPS parameters for 429,669 COSMOS2020 galaxies, including 39,588 with publicly available spectroscopic redshifts. The resulting redshift estimates exhibit minimal bias (median[Δ _z ] = −8 × 10 ^−4 ), scatter ( σ _MAD = 0.0132), and outlier fraction (6.19%) for the full 0 < z < 6 spectroscopic compilation. These results establish that pop-cosmos can achieve the accuracy and realism needed to forward model modern wide, deep surveys for Stage IV cosmology. We publicly release pop-cosmos software, mock galaxy catalogs, and COSMOS2020 redshift and SPS parameter posteriors.
The current standard model of cosmology successfully describes a variety of measurements, but the nature of its main ingredients, dark matter and dark energy, remains unknown. Euclid is a medium-class mission in the Cosmic Vision 2015-2025 programme of the European Space Agency (ESA) that will provide high-resolution optical imaging, as well as near-infrared imaging and spectroscopy, over about 14,000 deg^2 of extragalactic sky. In addition to accurate weak lensing and clustering measurements that probe structure formation over half of the age of the Universe, its primary probes for cosmology, these exquisite data will enable a wide range of science. This paper provides a high-level overview of the mission, summarising the survey characteristics, the various data-processing steps, and data products. We also highlight the main science objectives and expected performance.
Model mis-specification (e.g. the presence of outliers) is commonly encountered in astronomical analyses, often requiring the use of ad hoc algorithms which are sensitive to arbitrary thresholds (e.g. sigma-clipping). For any given data set, the optimal approach will be to develop a bespoke statistical model of the data generation and measurement processes, but these come with a development cost; there is hence utility in having generic modelling approaches that are both principled and robust to model mis-specification. Here we develop and implement a generic Bayesian approach to linear regression, based on Student's t-distributions, that is robust to outliers and mis-specification of the noise model. Our method is validated using simulated data sets with various degrees of model mis-specification; the derived constraints are shown to be systematically less biased than those from a similar model using normal distributions. We demonstrate that, for a data set without outliers, a worst-case inference using t-distributions would give unbiased results with $\lesssim \!\!10$ per cent increase in the reported parameter uncertainties. We also compare with existing analyses of real-world data sets, finding qualitatively different results where normal distributions have been used and agreement where more robust methods have been applied. A Python implementation of this model, t-cup, is made available for others to use.
We present a simple method for assessing the predictive performance of high-dimensional models directly in data space when only samples are available. Our approach is to compare the quantiles of observables predicted by a model to those of the observables themselves. In cases where the dimensionality of the observables is large (e.g., multiband galaxy photometry), we advocate that the comparison is made after projection onto a set of principal axes to reduce the dimensionality. We demonstrate our method on a series of two-dimensional examples. We then apply it to results from a state-of-the-art generative model for galaxy photometry (pop-cosmos) that generates predictions of colors and magnitudes by forward simulating from a 16-dimensional distribution of physical parameters represented by a score-based diffusion model. We validate the predictive performance of this model directly in a space of nine broadband colors. Although motivated by this specific example, we expect that the techniques we present will be broadly useful for evaluating the performance of flexible, nonparametric population models of this kind, and other settings where two sets of samples are to be compared.
Gravitational-wave (GW) observations of neutron star-black hole (NSBH) mergers are sensitive to the nuclear equation of state (EOS). We present a new methodology for EOS inference with nonparametric Gaussian process priors, enabling direct constraints on the pressure at specific densities and the length-scale of correlations on the EOS. Using realistic simulations of NSBH mergers, incorporating both GW and electromagnetic selection to ensure sample purity, we find that a GW detector network operating at O5 sensitivities will constrain the radius of a 1.4M circle dot NS and the maximum NS mass with 1.6% and 13% precision, respectively. With the same sample, the projected constraint on the length-scale of correlations in the EOS is >= 3.2 MeV fm-3. These results demonstrate strong potential for insights into the nuclear EOS from NSBH systems, provided they are robustly identified.
We present pop-cosmos: a comprehensive model characterizing the galaxy population, calibrated to $140,938$ ($r<25$ selected) galaxies from the Cosmic Evolution Survey (COSMOS) with photometry in $26$ bands from the ultra-violet to the infra-red. We construct a detailed forward model for the COSMOS data, comprising: a population model describing the joint distribution of galaxy characteristics and its evolution (parameterized by a flexible score-based diffusion model); a state-of-the-art stellar population synthesis (SPS) model connecting galaxies' instrinsic properties to their photometry; and a data-model for the observation, calibration and selection processes. By minimizing the optimal transport distance between synthetic and real data we are able to jointly fit the population- and data-models, leading to robustly calibrated population-level inferences that account for parameter degeneracies, photometric noise and calibration, and selection. We present a number of key predictions from our model of interest for cosmology and galaxy evolution, including the mass function and redshift distribution; the mass-metallicity-redshift and fundamental metallicity relations; the star-forming sequence; the relation between dust attenuation and stellar mass, star formation rate and attenuation-law index; and the relation between gas-ionization and star formation. Our model encodes a comprehensive picture of galaxy evolution that faithfully predicts galaxy colors across a broad redshift ($z<4$) and wavelength range.
This paper presents a search for high redshift galaxies from the Euclid Early Release Observations program "Magnifying Lens." The 1.5 deg$^2$ area covered by the twin Abell lensing cluster fields is comparable in size to the few other deep near-infrared surveys such as COSMOS, and so provides an opportunity to significantly increase known samples of rare UV-bright galaxies at $z\approx6-8$ ($M_{\rm UV}\lesssim-22$). Beyond their still uncertain role in reionisation, these UV-bright galaxies are ideal laboratories from which to study galaxy formation and constrain the bright-end of the UV luminosity function. Of the 501994 sources detected from a combined $Y_{\rm E}$, $J_{\rm E}$, and $H_{\rm E}$ NISP detection image, 168 do not have any appreciable VIS/$I_{\rm E}$ flux. These objects span a range in spectral colours, separated into two classes: 139 extremely red sources; and 29 Lyman-break galaxy candidates. Best-fit redshifts and spectral templates suggest the former is composed of both $z\gtrsim5$ dusty star-forming galaxies and $z\approx1-3$ quiescent systems. The latter is composed of more homogeneous Lyman break galaxies at $z\approx6-8$. In both cases, contamination by L- and T-type dwarfs cannot be ruled out with Euclid images alone. Additional contamination from instrumental persistence is investigated using a novel time series analysis. This work lays the foundation for future searches within the Euclid Deep Fields, where thousands more $z\gtrsim6$ Lyman break systems and extremely red sources will be identified.
We present an efficient Bayesian method for estimating individual photometric redshifts and galaxy properties under a pretrained population model (pop-cosmos) that was calibrated using purely photometric data. This model specifies a prior distribution over 16 stellar population synthesis (SPS) parameters using a score-based diffusion model, and includes a data model with detailed treatment of nebular emission. We use a GPU-accelerated affine-invariant ensemble sampler to achieve fast posterior sampling under this model for 292,300 individual galaxies in the COSMOS2020 catalog, leveraging a neural network emulator (Speculator) to speed up the SPS calculations. We apply both the pop-cosmos population model and a baseline prior inspired by Prospector-alpha, and compare these results to published COSMOS2020 redshift estimates from the widely used EAZY and LePhare codes. For the similar to 12,000 galaxies with spectroscopic redshifts, we find that pop-cosmos yields redshift estimates that have minimal bias (similar to 10(-4)), high accuracy (sigma(MAD) = 7 x 10(-3)), and a low outlier rate (1.6%). We show that the pop-cosmos population model generalizes well to galaxies fainter than its r < 25 mag training set. The sample we have analyzed is greater than or similar to 3x larger than has previously been possible via posterior sampling with a full SPS model, with average throughput of 15 GPU-sec per galaxy under the pop-cosmos prior, and 0.6 GPU-sec per galaxy under the Prospector prior. This paves the way for principled modeling of the huge catalogs expected from upcoming Stage IV galaxy surveys.
We present a Bayesian hierarchical framework to analyze photometric galaxy survey data with stellar population synthesis (SPS) models. Our method couples robust modeling of spectral energy distributions with a population model and a noise model to characterize the statistical properties of the galaxy populations and real observations, respectively. By self-consistently inferring all model parameters, from high-level hyper-parameters to SPS parameters of individual galaxies, one can separate sources of bias and uncertainty in the data.We demonstrate the strengths and flexibility of this approach by deriving accurate photometric redshifts for a sample of spectroscopically-confirmed galaxies in the COSMOS field, all with 26-band photometry and spectroscopic redshifts. We achieve a performance competitive with publicly-released photometric redshift catalogs based on the same data. Prior to this work, this approach was computationally intractable in practice due to the heavy computational load of SPS model calls; we overcome this challenge using with neural emulators. We find that the largest photometric residuals are associated with poor calibration for emission line luminosities and thus build a framework to mitigate these effects. This combination of physics-based modeling accelerated with machine learning paves the path towards meeting the stringent requirements on the accuracy of photometric redshift estimation imposed by upcoming cosmological surveys. The approach also has the potential to create new links between cosmology and galaxy evolution through the analysis of photometric datasets.
We construct a field-based Bayesian Hierarchical Model for cosmic shear that includes, for the first time, the important astrophysical systematics of intrinsic alignments and baryon feedback, in addition to a gravity model. We add to the BORG-WL framework the tidal alignment and tidal torquing model (TATT) for intrinsic alignments and compare them with the non-linear alignment (NLA) model. With synthetic data, we have shown that adding intrinsic alignments and sampling the TATT parameters does not reduce the constraining power of the method and the field-based approach lifts the weak lensing degeneracy. We add baryon effects at the field level using the enthalpy gradient descent (EGD) model. This model displaces the dark matter particles without knowing whether they belong to a halo and allows for self-calibration of the model parameters, which are inferred from the data. We have also illustrated the effects of model misspecification for the baryons. The resulting model now contains the most important physical effects and is suitable for application to data.
We present a forward-modeling framework for estimating galaxy redshift distributions from photometric surveys. Our forward model is composed of: a detailed population model describing the intrinsic distribution of the physical characteristics of galaxies, encoding galaxy evolution physics; a stellar population synthesis model connecting the physical properties of galaxies to their photometry; a data model characterizing the observation and calibration processes for a given survey; and explicit treatment of selection cuts, both into the main analysis sample and for the subsequent sorting into tomographic redshift bins. This approach has the appeal that it does not rely on spectroscopic calibration data, provides explicit control over modeling assumptions and builds a direct bridge between photo- z inference and galaxy evolution physics. In addition to redshift distributions, forward modeling provides a framework for drawing robust inferences about the statistical properties of the galaxy population more generally. We demonstrate the utility of forward modeling by estimating the redshift distributions for the Galaxy And Mass Assembly (GAMA) survey and the Vimos VLT Deep Survey (VVDS), validating against their spectroscopic redshifts. Our baseline model is able to predict tomographic redshift distributions for GAMA and VVDS with respective biases of Δ z ≲ 0.003 and Δ z ≃ 0.01 on the mean redshift—comfortably accurate enough for Stage III cosmological surveys—without any hyperparameter tuning (i.e., prior to doing any fitting to those data). We anticipate that with additional hyperparameter fitting and modeling improvements, forward modeling will provide a path to accurate redshift distribution inference for Stage IV surveys.
This report is the result of a joint discussion between the Rubin and Euclid scientific communities. The work presented in this report was focused on designing and recommending an initial set of Derived Data products (DDPs) that could realize the science goals enabled by joint processing. All interested Rubin and Euclid data rights holders were invited to contribute via an online discussion forum and a series of virtual meetings. Strong interest in enhancing science with joint DDPs emerged from across a wide range of astrophysical domains: Solar System, the Galaxy, the Local Volume, from the nearby to the primaeval Universe, and cosmology.
Received 13 June 2022DOI:https://doi.org/10.1103/PhysRevD.105.129904Published by the American Physical Society under the terms of the Creative Commons Attribution 4.0 International license. Further distribution of this work must maintain attribution to the author(s) and the published article’s title, journal citation, and DOI.Published by the American Physical SocietyPhysics Subject Headings (PhySH)Research AreasCosmic ray sourcesCosmic rays & astroparticlesTechniquesNeutrino detectorsGravitation, Cosmology & Astrophysics
Several tentative associations between high-energy neutrinos and astrophysical sources have been recently reported, but a conclusive identification of these potential neutrino emitters remains challenging. We explore the use of Monte Carlo simulations of source populations to gain deeper insight into the physical implications of proposed individual source–neutrino associations. In particular, we focus on the IC170922A–TXS 0506+056 observation. Assuming a null model, we find a 7.6% chance of mistakenly identifying coincidences between γ -ray flares from blazars and neutrino alerts in 10-year surveys. We confirm that a blazar–neutrino connection based on the γ -ray flux is required to find a low chance coincidence probability and, therefore, a significant IC170922A–TXS 0506+056 association. We then assume this blazar–neutrino connection for the whole population and find that the ratio of neutrino to γ -ray fluxes must be ≲10 −2 in order not to overproduce the total number of neutrino alerts seen by IceCube. For the IC170922A–TXS 0506+056 association to make sense, we must either accept this low flux ratio or suppose that only some rare sub-population of blazars is capable of high-energy neutrino production. For example, if we consider neutrino production only in blazar flares, we expect the flux ratio of between 10 −3 and 10 −1 to be consistent with a single coincident observation of a neutrino alert and flaring γ -ray blazar. These constraints should be interpreted in the context of the likelihood models used to find the IC170922A–TXS 0506+056 association, which assumes a fixed power-law neutrino spectrum of E −2.13 for all blazars.