The link between age and migration propensity is long established, but existing models of country-level net migration ignore the effect of population age distribution on past and projected migration rates. We propose a method to estimate and forecast international net migration rates for the 200 most populous countries, taking account of changes in population age structure. We use age-standardized estimates of country-level net migration rates and in-migration rates over quinquennial periods from 1990 through 2020 to decompose past net migration rates into in-migration rates and out-migration rates. We then recalculate historic migration rates on a scale that removes the influence of the population age distribution. This is done by scaling past and projected migration rates in terms of a reference population and period. We show that this can be done very simply, using a quantity we call the migration age structure index (MASI). We use a Bayesian hierarchical model to generate joint probabilistic forecasts of total and age- and sex- specific net migration rates over five-year periods for all countries from 2020 through 2100. We find that accounting for population age structure in historic and forecast net migration rates leads to narrower prediction intervals by the end of the century for most countries. Also, applying a Rogers Castro-like migration age schedule to migration outflows reduces uncertainty in population pyramid forecasts. Finally, accounting for population age structure leads to less out-migration among countries with rapidly aging populations that are forecast to contract most rapidly by the end of the century. This leads to less drastic population declines than are forecast without accounting for population age structure.
Bayesian variable selection requires sampling from a posterior distribution that combines discrete model indicators with continuously varying parameters, a challenge often addressed through reversible jump Markov chain Monte Carlo (RJMCMC). Despite its generality, RJMCMC is widely regarded as difficult to design and implement correctly. We present mixtures of mutually singular (MoMS) distributions as a transparent alternative in which competing models are represented within a single fixed-dimensional parameter space partitioned into mutually singular subspaces. We show that this formulation reproduces the exact spike-and-slab interpretation of Bayesian variable selection and that, under appropriate constructions, MoMS and RJMCMC share the same Metropolis–Hastings acceptance probability. On a benchmark dataset with ten predictors, both methods recover posterior inclusion probabilities that match full enumeration, while MoMS achieves comparable or superior effective sample size per second relative to a carefully engineered RJMCMC scheme. We further illustrate the approach in a mixed-effects logistic regression for a sleep-and-memory experiment and in factor-loading selection for a multidimensional generalized partial credit model. Together, these results show that Bayesian variable selection can be carried out within standard fixed-dimensional Markov chain Monte Carlo methodology – without regret.
The Transient Climate Response to cumulative CO2 Emissions (TCRE) is a key metric for linking greenhouse gas emissions to global temperature change and informing climate policy. However, extant estimates of the TCRE often depend on subjective model selection or assumed sensitivity ranges, with limited validation against observed data. We develop a fully statistical data-driven approach using a Bayesian Model Averaging (BMA) approach to estimate the TCRE. This uses 37 climate models from the Coupled Model Intercomparison Project phase 6 (CMIP6), weighted according to their consistency with observed temperature data. Compared to the Intergovernmental Panel on Climate Change (IPCC)'s Sixth Assessment Report (AR6) TCRE estimate, our BMA approach yields a very likely range (90
Model uncertainty is a central challenge in statistical models for binary outcomes such as logistic regression, arising when it is unclear which predictors should be included in the model. Many methods have been proposed to address this issue for logistic regression, but their relative performance under realistic conditions remains poorly understood. We therefore conducted a preregistered, simulation-based comparison of 28 established methods for variable selection and inference under model uncertainty, using 11 empirical datasets spanning a range of sample sizes and numbers of predictors, in cases both with and without separation. We found that Bayesian model averaging (BMA) methods based on [Formula: see text]-priors, particularly [Formula: see text], show the strongest overall performance when separation is absent. When separation occurs, penalized likelihood approaches, especially the LASSO, provide the most stable results, while BMA with the local empirical Bayes (EB-local) prior is competitive in both situations. These findings offer practical guidance for applied researchers on how to effectively address model uncertainty in logistic regression in modern empirical and machine learning research.
We consider the problem of estimating a high-dimensional covariance matrix from a small number of observations when covariates on pairs of variables are available and the variables can have spatial structure. This is motivated by the problem arising in demography of estimating the covariance matrix of the total fertility rate (TFR) of 195 different countries when only 11 observations are available. We construct an estimator for high-dimensional covariance matrices by exploiting information about pairwise covariates, such as whether pairs of variables belong to the same cluster, or spatial structure of the variables, and interactions between the covariates. We reformulate the problem in terms of a mixed effects model. This requires the estimation of only a small number of parameters, which are easy to interpret and which can be selected using standard procedures. The estimator is consistent under general conditions, and asymptotically normal. It works if the mean and variance structure of the data is already specified or if some of the data are missing. We assess its performance under our model assumptions, as well as under model misspecification, using simulations. We find that it outperforms several popular alternatives. We apply it to the TFR dataset and draw some conclusions.
Responsible data sharing anchors research reproducibility and promotes the integrity of scientific research. Motivated by Canadian Scleroderma Research Group (CSRG) patient registry data, we present a risk-based method to produce privacy-preserved and high-utility synthetic datasets, which also simultaneously imputes missing data of mixed continuous and categorical types in the original dataset. This method divides all individuals into different subgroups, based on their reidentification risks, and provides tailored synthesis strategies targeted for each risk subgroup, through the associated tuning mechanisms. Under our setting, our risk-based method reduced the number of patients at risk from 198 to four, among the 691 CSRG patients who have no missing values in any of the quasi-identifying variables, while preserving all correct inferential conclusions in the target analysis. The 95% confidence intervals (CIs) have 92.6% overlap, on average, with the CIs constructed using the unperturbed imputation-completed datasets. These findings suggest that our risk-based method makes it possible to release complete synthetic datasets for research reproducibility while ensuring that the reidentification risks are acceptably low. In contrast, the existing one-size-fits-all synthesis strategies that do not take account of different risk levels can lead to unnecessary information loss and possibly incorrect scientific conclusions.
International comparisons of hierarchical time series data sets based on survey data, such as annual country-level estimates of school enrollment rates, can suffer from large amounts of missing data due to differing coverage of surveys across countries and across times. A popular approach to handling missing data in these settings is through multiple imputation, which can be especially effective when there is an auxiliary variable that is strongly predictive of and has a smaller amount of missing data than the variable of interest. However, standard methods for multiple imputation of hierarchical time series data can perform poorly when the auxiliary variable and the variable of interest have a nonlinear relationship. Performance can also suffer if the multiple imputations are used to estimate an analysis model that makes different assumptions about the data compared to the imputation model, leading to uncongeniality between analysis and imputation models. We propose a Bayesian method for multiple imputation of hierarchical nonlinear time series data that uses a sequential decomposition of the joint distribution and incorporates smoothing splines to account for nonlinear relationships between variables. We compare the proposed method with existing multiple imputation methods through a simulation study and an application to secondary school enrollment data. We find that the proposed method can lead to substantial performance increases for estimation of parameters in uncongenial analysis models and for prediction of individual missing values.
Estimates of future migration patterns are of broad interest in demography. Forced migration, including refugee and asylum seekers, plays an important role in overall migration patterns but is notoriously difficult to forecast. Focusing on refugees and asylum seekers, we propose a modeling pipeline based on Bayesian hierarchical time-series modeling for projecting refugee population official statistics by country of origin using data from the United Nations High Commissioner for Refugees. Our approach is based on a conceptual model of refugee and asylum seeker populations following growth and decline phases, separated by a peak. The growth and decline phases are modeled by logistic growth and decline through an interrupted logistic process model. We evaluate our method through a set of validation exercises that show it has good performance for forecasts at 1-, 5-, and 10-year horizons, and we present projections for 35 countries of origin of large refugee and asylum seeker populations.
BACKGROUND:Non-pharmaceutical interventions (NPIs) in response to the COVID-19 pandemic necessitated a trade-off between the health impacts of viral spread and the social and economic costs of restrictions. Navigating this trade-off proved consequential, contentious, and challenging for decision-makers. METHODS:We conduct a cost-effectiveness analysis of NPIs enacted at the state level in the United States (US) in 2020. We combine data on COVID-19 cases, deaths, policies, and the social, economic, and health consequences of infections and interventions within an epidemiological model. We estimate SARS-CoV-2 prevalence, transmission rates, effects of interventions, and costs associated to infections and NPIs in each US state. We use these estimates to quantitatively evaluate the efficacy and gross impacts of the policy schedules implemented during the pandemic. We also derive optimal cost-effective strategies that minimize aggregate costs to society. RESULTS:We find that NPIs were effective in substantially reducing SARS-CoV-2 transmission, averting 860,000 (95% CI: 560,000-1,190,000) COVID-19 deaths in the US in 2020. Although school closures reduced transmission, their social impact in terms of student learning loss was too costly, depriving the nation of $2 trillion in 2020 US dollars (USD2020), conservatively, in future Gross Domestic Product (GDP). Moreover, this marginal trade-off between school closure and COVID-19 deaths was not inescapable: a combination of other measures would have been enough to maintain similar or lower mortality rates without incurring such profound learning loss. Optimal policies involve consistent implementation of mask mandates, public test availability, contact tracing, social distancing orders, and reactive workplace closures, with no closure of schools. Their use would have reduced the gross impact of the pandemic in the US in 2020 from $4.6 trillion to $1.9 trillion and, with high probability, saved over 100,000 lives. CONCLUSIONS:US COVID-19 school closure was not cost-effective, but other measures were. While our study focuses on COVID-19 in the US prior to vaccines, our methodological contributions and findings about the cost-effectiveness and optimal structure of NPI policies have implications for the response to future epidemics and in other countries. Our results also highlight the need to address the substantial global learning deficit incurred during the pandemic.
We present a new version of the truncated harmonic mean estimator (THAMES) for univariate or multivariate mixture models. The estimator computes the marginal likelihood from Markov chain Monte Carlo (MCMC) samples, is consistent, asymptotically normal and of finite variance. In addition, it is invariant to label switching, does not require posterior samples from hidden allocation vectors, and is easily approximated, even for an arbitrarily high number of components. Its computational efficiency is based on an asymptotically optimal ordering of the parameter space, which can in turn be used to provide useful visualisations. We test it in simulation settings where the true marginal likelihood is available analytically. It performs well against state-of-the-art competitors, even in multivariate settings with a high number of components. We demonstrate its utility for inference and model selection on univariate and multivariate data sets.
Projecting future climate change is important for implementing the 2015 Paris Agreement, which aims to limit greenhouse gas emissions to a level that would keep the global average temperature increase to 2100 below 2 °C. The Intergovernmental Panel on Climate Change uses emissions scenarios for projecting climate change, but since 2017, an alternative fully statistical Bayesian probabilistic approach has been developed. Both approaches rely on an equation that expresses emissions as the product of population, Gross Domestic Product (GDP) per capita, and carbon intensity, namely carbon emissions per unit of GDP. Here, we use data on these quantities for 2015-2024 to probabilistically assess the changes in climate change prospects associated with post-Paris emissions. These show that carbon intensity declined (i.e., improved) substantially over that period, but that overall carbon emissions rose, due to the rapid rise in world GDP, which more than canceled out the progress made. We found that the projected temperature increase to 2100 declined only slightly, from 2.6° C to 2.4 °C. Meanwhile, the chance of staying below 2 °C remained low, at 17%. However, the chance of the most catastrophic climate change, above 3 °C, has gone down substantially, from 26% to 9%.
Multidimensional scaling (MDS) is a widely used approach to representing high-dimensional, dependent data. MDS works by assigning each observation a location on a low-dimensional geometric manifold, with distance on the manifold representing similarity. We propose a Bayesian approach to multidimensional scaling when the low-dimensional manifold is hyperbolic. Using hyperbolic space facilitates representing tree-like structures common in many settings (e.g. text or genetic data with hierarchical structure). A Bayesian approach provides regularization that minimizes the impact of measurement error in the observed data and assesses uncertainty. We also propose a case-control likelihood approximation that allows for efficient sampling from the posterior distribution in larger data settings, reducing computational complexity from approximately $O(n^2)$ to $O(n)$. We evaluate the proposed method against state-of-the-art alternatives using simulations, canonical reference datasets, Indian village network data, and human gene expression data.
BACKGROUND Population projections for all countries are published by the United Nations Population Division (UNPD) every two years as part of the World Population Prospects (WPP). Since 2015, probabilistic population projections have been published as part of WPP, produced using Bayesian statistical models. Central to this methodological change was a team of statisticians at the University of Washington, led by Professor Adrian Raftery. OBJECTIVE This interview with Adrian Raftery details the history of the UNPD WPP probabilistic population projections, including how the project started, the methodological challenges, main takeaways and lessons, and priorities for future research. CONTRIBUTION This interview contributes to the record of scientific thought and the advancement of methodology in demographic research. It demonstrates the evolution of a successful scientific project with large scientific impact and a broader influence on the field of Bayesian demography.
In this chapter, we present a review of latent position models for networks. We review the recent literature in this area and illustrate the basic aspects and properties of this modeling framework. Through several illustrative examples we highlight how the latent position model is able to capture important features of observed networks. We emphasize how the canonical design of this model has made it popular thanks to its ability to provide interpretable visualizations of complex network interactions. We outline the main extensions that have been introduced to this model, illustrating its flexibility and applicability.
We propose an easily computed estimator of the marginal likelihood from posterior simulation output, via reciprocal importance sampling, combining earlier proposals of DiCiccio et al (1997) and Robert and Wraith (2009). This involves only the unnormalized posterior densities from the sampled parameter values, and does not involve additional simulations beyond the main posterior simulation, or additional complicated calculations, provided that the parameter space is unconstrained. Even if this is not the case, the estimator is easily adjusted by a simple Monte Carlo approximation. It is unbiased for the reciprocal of the marginal likelihood, consistent, has finite variance, and is asymptotically normal. It involves one user-specified control parameter, and we derive an optimal way of specifying this. We illustrate it with several numerical examples.
Women's educational attainment and contraceptive prevalence are two mechanisms identified as having an accelerating effect on fertility decline and that can be directly impacted by policy. Quantifying the potential accelerating effect of education and family planning policies on fertility decline in a probabilistic way is of interest to policymakers, particularly in highfertility countries. We propose a conditional Bayesian hierarchical model for projecting fertility, given education and family planning policy interventions. To illustrate the effect policy changes could have on future fertility, we create probabilistic projections of fertility that condition on scenarios such as achieving the sustainable development goals (SDGs) for universal secondary education and universal access to family planning by 2030.
This paper proposes two new approaches to improve the estimation of the coefficients of the multivariate HAR (MHAR) model with the primary purpose of improving forecast performance. A robust estimator of the covariance matrix is adopted to replace the realized covariance matrix while estimating the MHAR model. The robustness to outliers of the new estimator makes the OLS estimation scheme for the MHAR model more reliable. In addition, a robust estimation scheme is developed for the MHAR model, which is based on the multivariate least-trimmed squares method. Both approaches provide significant improvements in forecasting performance based on both statistical loss and portfolio outcomes. The forecast performance of the multivariate HARQ model can also be improved with the proposed approaches, as evidenced by robustness checks.
The bayesTFR package for R provides a set of functions to produce probabilistic projections of the total fertility rates (TFR) for all countries, and is widely used, including as part of the basis for the UN's official population projections for all countries. Liu and Raftery (2020) extended the theoretical model by adding a layer that accounts for the past TFR estimation uncertainty. A major update of bayesTFR implements the new extension. Moreover, a new feature of producing annual TFR estimation and projections extends the existing functionality of estimating and projecting for five-year time periods. An additional autoregressive component has been developed in order to account for the larger autocorrelation in the annual version of the model. This article summarizes the updated model, describes the basic steps to generate probabilistic estimation and projections under different settings, compares performance, and provides instructions on how to summarize, visualize and diagnose the model results.
Population projections provide predictions of future population sizes for an area. Historically, most population projections have been produced using deterministic or scenario-based approaches and have not assessed uncertainty about future population change. Starting in 2015, however, the United Nations (UN) has produced probabilistic population projections for all countries using a Bayesian approach. There is also considerable interest in subnational probabilistic population projections, but the UN's national approach cannot be used directly for this purpose, because within-country correlations in fertility and mortality are generally larger than between-country ones, migration is not constrained in the same way, and there is a need to account for college and other special populations, particularly at the county level. We propose a Bayesian method for producing subnational population projections, including migration and accounting for college populations, by building on but modifying the UN approach. We illustrate our approach by applying it to the counties of Washington State and comparing the results with extant deterministic projections produced by Washington State demographers. Out-of-sample experiments show that our method gives accurate and well-calibrated forecasts and forecast intervals. In most cases, our intervals were narrower than the growth-based intervals issued by the state, particularly for shorter time horizons.
Clustering is the task of automatically gathering observations into homogeneous groups, where the number of groups is unknown. Through its basis in a statistical modeling framework, model-based clustering provides a principled and reproducible approach to clustering. In contrast to heuristic approaches, model-based clustering allows for robust approaches to parameter estimation and objective inference on the number of clusters, while providing a clustering solution that accounts for uncertainty in cluster membership. The aim of this article is to provide a review of the theory underpinning model-based clustering, to outline associated inferential approaches, and to highlight recent methodological developments that facilitate the use of model-based clustering for a broad array of data types. Since its emergence six decades ago, the literature on model-based clustering has grown rapidly, and as such, this review provides only a selection of the bibliography in this dynamic and impactful field.