Estimates of future migration patterns are of broad interest in demography. Forced migration, including refugee and asylum seekers, plays an important role in overall migration patterns but is notoriously difficult to forecast. Focusing on refugees and asylum seekers, we propose a modeling pipeline based on Bayesian hierarchical time-series modeling for projecting refugee population official statistics by country of origin using data from the United Nations High Commissioner for Refugees. Our approach is based on a conceptual model of refugee and asylum seeker populations following growth and decline phases, separated by a peak. The growth and decline phases are modeled by logistic growth and decline through an interrupted logistic process model. We evaluate our method through a set of validation exercises that show it has good performance for forecasts at 1-, 5-, and 10-year horizons, and we present projections for 35 countries of origin of large refugee and asylum seeker populations.
Estimating quantiles of an outcome conditional on covariates is of fundamental interest in statistics with broad application in probabilistic prediction and forecasting. We propose an ensemble method for conditional quantile estimation, Quantile Super Learning, that combines predictions from multiple candidate algorithms based on their empirical performance measured with respect to a cross-validated empirical risk of the quantile loss function. We present theoretical guarantees for both iid and online data scenarios. The performance of our approach for quantile estimation and in forming prediction intervals is tested in simulation studies. Two case studies related to solar energy are used to illustrate Quantile Super Learning: in an iid setting, we predict the physical properties of perovskite materials for photovoltaic cells, and in an online setting we forecast ground solar irradiance based on output from dynamic weather ensemble models.
Several demographic and health indicators, including the total fertility rate (TFR) and modern contraceptive use rate (mCPR), evolve similarly over time, characterized by a transition between stable states. Existing approaches for estimation or projection of transitions in multiple populations have successfully used parametric functions to capture the relation between the rate of change of an indicator and its level. However, incorrect parametric forms may result in bias or incorrect coverage in long-term projections. We propose a new class of models to capture demographic transitions in multiple populations. Our proposal, the B-spline Transition Model (BTM), models the relationship between the rate of change of an indicator and its level using B-splines, allowing for data-adaptive estimation of transition functions. Bayesian hierarchical models are used to share information on the transition function between populations. We apply the BTM to estimate and project country-level TFR and mCPR and compare the results against those from extant parametric models. For TFR, BTM projections have generally lower error than the comparison model. For mCPR, while results are comparable between BTM and a parametric approach, the B-spline model generally improves out-of-sample predictions. The case studies suggest that the BTM may be considered for demographic applications.
Statistical models are used to produce estimates of demographic and global health indicators in populations with limited data. Such models integrate multiple data sources to produce estimates and forecasts with uncertainty based on model assumptions. Model assumptions can be divided into assumptions that describe latent trends in the indicator of interest versus assumptions on the data generating process of the observed data, conditional on the latent process value. Focusing on the latter, we introduce a class of data models that can be used to combine data from multiple sources with various reporting issues. The proposed data model accounts for sampling errors and differences in observational uncertainty based on survey characteristics. In addition, the data model employs horseshoe priors to produce estimates that are robust to outlying observations. We refer to the data model class as the normal-with-optional-shrinkage (NOS) set up. We illustrate the use of the NOS data model for the estimation of modern contraceptive use and other family planning indicators at the national level for countries globally, using survey data.
Demographic and health indicators may exhibit short or large short-term shocks; for example, armed conflicts, epidemics, or famines may cause shocks in period measures of life expectancy. Statistical models for estimating historical trends and generating future projections of these indicators for a large number of populations may be biased or not well probabilistically calibrated if they do not account for the presence of shocks. We propose a flexible method for modeling shocks when producing estimates and projections for multiple populations. The proposed approach makes no assumptions about the shape or duration of a shock, and requires no prior knowledge of when shocks may have occurred. Our approach is based on the modeling of shocks in level of the indicator of interest. We use Bayesian shrinkage priors such that shock terms are shrunk to zero unless the data suggest otherwise. The method is demonstrated in a model for male period life expectancy at birth. We use as a starting point an existing projection model and expand it by including the shock terms, modeled by the Bayesian shrinkage priors. Out-of-sample validation exercises find that including shocks in the model results in sharper uncertainty intervals without sacrificing empirical coverage or prediction error.
The Family Planning Estimation Tool (FPET) is used in low- and middle-income countries to produce estimates and short-term forecasts of family planning indicators, such as modern contraceptive use and unmet need for contraceptives. Estimates are obtained via a Bayesian statistical model that is fitted to country-specific data from surveys and service statistics data. The model has evolved over the last decade based on user inputs. In this paper we summarize the main features of the statistical model used in FPET and introduce recent updates related to capturing contraceptive transitions, fitting to survey data that may be error prone, and the use of service statistics data. We assess model performance through a validation exercise and find that FPET is reasonably well calibrated. We use our experience with FPET to briefly discuss lessons learned and open challenges related to the broader field of statistical modeling for monitoring of demographic and global health indicators.
Aggregate measures of family planning are used to monitor demand for and usage of contraceptive methods in populations globally, for example as part of the FP2030 initiative. Family planning measures for low- and middle-income countries are typically based on data collected through cross-sectional household surveys. Recently proposed measures account for sexual activity through assessment of the distribution of time-between-sex (TBS) in the population of interest. In this paper, we propose a statistical approach to estimate the distribution of TBS using data typically available in low- and middle-income countries, while addressing two major challenges. The first challenge is that timing of sex information is typically limited to women's time-since-last-sex (TSLS) data collected in the cross-sectional survey. In our proposed approach, we adopt the current duration method to estimate the distribution of TBS using the available TSLS data, from which the frequency of sex at the population level can be derived. Furthermore, the observed TSLS data are subject to reporting issues because they can be reported in different units and may be rounded off. To apply the current duration approach and account for these data reporting issues, we develop a flexible Bayesian model, and provide a detailed technical description of the proposed modeling approach.
BACKGROUND:Per- and polyfluoroalkyl substances (PFASs) are common industrial and consumer product chemicals with widespread human exposures that have been linked to adverse health effects. PFASs are commonly detected in foods and food-contact materials (FCMs), including fast food packaging and microwave popcorn bags. OBJECTIVES:Our goal was to investigate associations between serum PFASs and consumption of restaurant food and popcorn in a representative sample of Americans. METHODS:We analyzed 2003-2014 serum PFAS and dietary recall data from the National Health and Nutrition Examination Survey (NHANES). We used multivariable linear regressions to investigate relationships between consumption of fast food, restaurant food, food eaten at home, and microwave popcorn and serum levels of perfluorooctanoic acid (PFOA), perfluorononanoic acid (PFNA), perfluorodecanoic acid (PFDA), perfluorohexanesulfonic acid (PFHxS), and perfluorooctanesulfonic acid (PFOS). RESULTS:Calories of food eaten at home in the past 24 h had significant inverse associations with serum levels of all five PFASs; these associations were stronger in women. Consumption of meals from fast food/pizza restaurants and other restaurants was generally associated with higher serum PFAS concentrations, based on 24-h and 7-d recall, with limited statistical significance. Consumption of popcorn was associated with significantly higher serum levels of PFOA, PFNA, PFDA, and PFOS, based on 24-h and 12-month recall, up to a 63% (95% CI: 34, 99) increase in PFDA among those who ate popcorn daily over the last 12 months. CONCLUSIONS:Associations between serum PFAS and popcorn consumption may be a consequence of PFAS migration from microwave popcorn bags. Inverse associations between serum PFAS and food eaten at home-primarily from grocery stores-is consistent with less contact between home-prepared food and FCMs, some of which contain PFASs. The potential for FCMs to contribute to PFAS exposure, coupled with concerns about toxicity and persistence, support the use of alternatives to PFASs in FCMs. https://doi.org/10.1289/EHP4092.
Two of the principle tasks of causal inference are to define and estimate the effect of a treatment on an outcome of interest. Formally, such treatment effects are defined as a possibly functional summary of the data generating distribution, and are referred to as target parameters. Estimation of the target parameter can be difficult, especially when it is high-dimensional. Marginal Structural Models (MSMs) provide a way to summarize such target parameters in terms of a lower dimensional working model. We introduce the semi-parametric efficiency bound for estimating MSM parameters in a general setting. We then present a frequentist estimator that achieves this bound based on Targeted Minimum Loss-Based Estimation. Our results are derived in a general context, and can be easily adapted to specific data structures and target parameters. We then describe a novel targeted Bayesian estimator and provide a Bernstein von-Mises type result analyzing its asymptotic behavior. We propose a universal algorithm that uses automatic differentiation to put the estimator into practice for arbitrary choice of working model. The frequentist and Bayesian estimators have been implemented in the Julia software package TargetedMSM.jl. Finally, we illustrate our proposed methods by investigating the effect of interventions on family planning behavior using data from a randomized field experiment conducted in Malawi.
Family planning measures for unmarried women are based on contraceptive demand and use among sexually active women. Sexual activity status is commonly defined based on comparing reported time-since-last-sex to a cutoff time, with women defined to be sexually active if their most recent sex was within the last four weeks. While easy to understand and compute, this approach to constructing family planning measures results in a limited understanding of family planning and exposure to unintended pregnancy because it cannot comprehensively capture the frequency of sex at the population level. We propose a new statistical approach to quantify sexual activity, using reported time-since-last-sex data. Based on estimated frequencies of sex among users and nonusers in need of family planning, we propose new family planning measures, including the ratio of protected exposure over all women's exposure to risk of unintended pregnancy.
Several e ff ect size and con fi dence interval estimates from the 24-hour dietary recall models were incorrectly reported due to an error in the equation used to derive the estimates. The estimates in the published manuscript were calculated by ð exp ð b Þ -1 Þ × a × 100 % with a 95% con fi dence interval (CI) given by ð exp ð b ± critical value× SE Þ − 1 Þ × a × 100 % , where b is the coe ffi cient from the linearregression model and a was a scaling factor used to transform the scale of the estimate (e.g., a = 100 was used to transform from percent increase in serum PFAS per additional 1 kcal = day to percent increase in serum PFAS per additional 100 kcal = day). This formula was incorrect, as the raw coe ffi cient should have been scaled prior to exponentiation. There were also mistakes in the presentations of these formulae, leading to the equations present in the text appearing as “ ð exp ð b Þ − 1 Þ × 100 % ” and “ ð exp ð b ± critical value× SE Þ − 1 Þ × 100 % . ” The estimates were updated to follow the correct formula, ð exp ð b × a Þ − 1 Þ × 100 % with a 95% CI given by ð exp ðð b ± critical value× SE Þ × a Þ − 1 Þ × 100 % . The p -values associated with these estimates were accurate as reported and were not a ff ected by the error. This error a ff ects text in the “ Results
Conformal Inference (CI) is a popular approach for generating finite sample prediction intervals based on the output of any point prediction method when data are exchangeable. Adaptive Conformal Inference (ACI) algorithms extend CI to the case of sequentially observed data, such as time series, and exhibit strong theoretical guarantees without having to assume exchangeability of the observed data. The common thread that unites algorithms in the ACI family is that they adaptively adjust the width of the generated prediction intervals in response to the observed data. We provide a detailed description of five ACI algorithms and their theoretical guarantees, and test their performance in simulation studies. We then present a case study of producing prediction intervals for influenza incidence in the United States based on black-box point forecasts. Implementations of all the algorithms are released as an open-source R package, AdaptiveConformal, which also includes tools for visualizing and summarizing conformal prediction intervals.
There is growing interest in producing estimates of demographic and global health indicators in populations with limited data. Statistical models are needed to combine data from multiple data sources into estimates and projections with uncertainty. Diverse modelling approaches have been applied to this problem, making comparisons between models difficult. We propose a model class, Temporal Models for Multiple Populations (TMMPs), to facilitate both documentation of model assumptions in a standardised way and comparison across models. The class makes a distinction between the process model, which describes latent trends in the indicator interest, and the data model, which describes the data generating process of the observed data. We provide a general notation for the process model that encompasses many popular temporal modelling techniques, and we show how existing models for a variety of indicators can be written using this notation. We end with a discussion of outstanding questions and future directions.
Background: Study participants want to receive their biomonitoring results for environmental chemicals, and ethics guidelines encourage reporting back. However, few studies have quantitively assessed participants’ responses to individual exposure reports, and digital methods have not been evaluated. Objectives: We isolated effects of receiving personal results vs. only study-wide findings and investigated whether effects differed for Black participants. Methods: We randomly assigned a subset of 295 women from the Child Health and Development Studies, half of whom were Black, to receive a report with personal environmental chemical results or only study-wide (aggregate) findings. Reports included results for 42 chemicals and lipids and were prepared using the Digital Exposure Report-Back Interface (DERBI). Women were interviewed before and after viewing their report. We analyzed differences in website activity, emotional responses, and intentions to participate in future research by report type and race using Wilcoxon rank sum tests, Wilcoxon-Pratt signed ranks tests, and multiple regression. Results: The personal report group spent approximately twice as much time on their reports as the aggregate group before the post-report-back interview. Among personal-report participants (n=93), 84% (78) viewed chemical group information for at least one personal result highlighted on their home page; among aggregate-report participants (n=94), 66% (62) viewed any chemical group page. Both groups reported strong positive feelings (curious, informed, interested, respected) about receiving results before and after report-back and mild negative feelings (helpless, scared, worried). Although most participants remained unworried after report-back, worry increased by a small amount in both groups. Among Black participants, higher post report-back worry was associated with having high levels of chemicals. Conclusions: Participants were motivated by their personal results to access online information about chemical sources and potential health effects. Report-back was associated with a small increase in worry, which could motivate appropriate action. Personal report-back increased engagement with exposure reports among Black participants. https://doi.org/10.1289/EHP9072
Nearly all Americans have detectable concentrations of endocrine disrupting chemicals from consumer products in their bodies, and expert panels recommend reducing exposures. To inform exposure reduction, we investigated whether consumers who are trying to avoid certain chemicals in consumer products have lower exposures than those who are not. We also aimed to make exposure biomonitoring more widely available. We enrolled 726 participants in a crowdsourced biomonitoring study. We targeted phenolic compounds-specifically parabens, bisphenol A (BPA) and analogs bisphenol F (BPF) and bisphenol S (BPS), the UV filter benzophenone-3, the anti-microbial triclosan, 2,4-dichlorophenol, and 2,5-dichlorophenol-and collected survey data on consumer products, cleaning habits, and efforts to avoid related chemicals. We investigated associations between 68 self-reported exposure behaviors and urine concentrations of ten chemicals, and evaluated whether associations were modified by intention to avoid exposures. A large majority (87%) of participants reported taking steps to limit exposure to specific chemicals, and, overall, participants achieved lower concentrations than the general U.S. population for parabens, BPA, triclosan, and benzophenone-3 but not BPF and BPS. Participants who reported avoiding all four ingredient groups-parabens, triclosan, bisphenols, and fragrances-were twice as likely as others to be in the lowest quartile of cumulative exposure. Avoiding certain products and reading ingredient labels to avoid chemicals was most effective for parabens, triclosan, and benzophenone-3. Avoiding BPA was not effective for reducing bisphenol exposures. Avoiding certain chemicals in products was generally associated with reduced exposure for chemicals listed on labels. Greater ingredient transparency will help consumers who read labels to reduce their exposure to a wider range of potentially harmful chemicals. In order to more equitably address public health, labeling policies should be complemented by regulations that exclude harmful chemicals from consumer products.
We sought to identify sources of twelve environmental phenols in a population highly motivated to learn about their personal chemical exposures. Over 300 people, most of whom were already connected to an environmental health network, signed up for Detox Me Action Kit as part of an online crowdfunding campaign in 2017. To our knowledge this is the first biomonitoring cohort to use this recruitment strategy. Participants collected two urine samples at home and completed an online questionnaire about exposure-related behaviors. Samples were returned frozen via overnight mail, and then composited and analyzed for parabens, bisphenols, chlorinated phenols, antimicrobials, and a UV filter. Participants received their personal exposure results as an interactive web-report. Over half the participants who completed the questionnaire reported avoiding products containing BPA, triclosan, and parabens, and measured urinary concentrations were generally lower than those reported in the National Health and Nutrition Examination Survey (NHANES). These results suggest that participants were already aware of environmental chemicals and taking steps to reduce exposures. However, intentions to avoid ingredients did not always translate to dependable behavior or to lower exposure levels. After checking the labels of their products, over a quarter of the 112 female participants who reported avoiding products containing parabens found that they used at least one personal care product where parabens was a listed ingredient. In the case of bisphenols, while participants had lower levels of the widely-scrutinized chemical BPA—in line with self-reported intentions—levels of BPF in our cohort were higher compared to levels from the 2013-2014 NHANES cycle. In this example of regrettable substitution, manufacturers may be substituting closely-related chemicals in "BPA-free" products, or in products more broadly. We explored these and other relationships in this cohort of engaged participants.
Summary:Researchers and clinicians in environmental health and medicine increasingly show respect for participants and patients by involving them in decision-making. In this context, the return of personal results to study participants is becoming ethical best practice, and many participants now expect to see their data. However, researchers often lack the time and expertise required for report-back, especially as studies measure greater numbers of analytes, including many without clear health guidelines. In this article, our goal is to demonstrate how a prototype digital method, the Digital Exposure Report-Back Interface (DERBI), can reduce practical barriers to high-quality report-back. DERBI uses decision rules to automate the production of personalized summaries of notable results and generates graphs of individual results with comparisons to the study group and benchmark populations. Reports discuss potential sources of chemical exposure, what is known and unknown about health effects, strategies for exposure reduction, and study-wide findings. Researcher tools promote discovery by drawing attention to patterns of high exposure and offer novel ways to increase participant engagement. DERBI reports have been field tested in two studies. Digital methods like DERBI reduce practical barriers to report-back thus enabling researchers to meet their ethical obligations and participants to get knowledge they can use to make informed choices.
Researchers conduct biomonitoring studies to characterize the prevalence of environmental chemicals in participants’ bodies by testing their blood, urine, or other media. These test results are returned to participants in a process called “report-back.” Designing effective report-back is complicated by several uncertainties related to interpreting personal chemical results. Chemical levels do not inform participants about the sources of their exposure, the health implications of exposure, or if and how they can reduce exposure. In this position paper we investigate how an interactive expert system can help participants identify potential sources of exposure and exposure reduction strategies. Author
There is a lack of robust statistical analyses for random effects linear models. In practice, statistical analyses, including estimation, prediction and inference, are not reliable when data are unbalanced, of small size, contain outliers, or not normally distributed. It is fortunate that rank-based regression analysis is a robust nonparametric alternative to likelihood and least squares analysis. We propose an R package that calculates rank-based statistical analyses for two-and three-level random effects nested designs. In this package, a new algorithm which recursively obtains robust predictions for both scale and random effects is used, along with three rank-based fitting methods.