Accurate estimation and forecasts for neonatal mortality rates (NMRs) in low- and middle-income countries is an urgent problem. Much of child mortality is preventable, and understanding temporal trends is of great interest when evaluating past performance and planning future policy or programming. In countries without robust vital registration, we rely on modeled estimates based on survey data to understand trends. A toolkit of compelling temporal models exists, but these methods have not been comprehensively evaluated for their application for the estimation of the NMR in low- and middle-income countries using household survey data. Using Demographic and Health Surveys (DHS) and Multiple Indicator Cluster Surveys (MICS) data from 41 countries in sub-Saharan Africa, we estimate and forecast the national-level NMR for 1970-2030 separately with random walk, auto-regressive, penalized spline, natural spline, and logit-linear latent temporal models. We examine the statistical behavior of these temporal models with both an out-of-sample analysis using the DHS and MICS data and a simulation study. We find that the second-order random walk and the penalized spline have the least bias, and short-term forecasts from the penalized spline tend to have narrower intervals with better out-of-sample performance. From the analysis of the NMR in sub-Saharan Africa, we estimate that 6 or fewer of the 41 countries included are on track to achieve the Sustainable Development Goals target of 12 neonatal deaths per 1000 live births by 2030.
Understanding the prevalence of key demographic and health indicators in small geographic areas and domains is of global interest, especially in low- and middle-income countries (LMICs), where vital registration data is sparse and household surveys are the primary source of information. Recent advances in computation and the increasing availability of spatially detailed datasets have led to much progress in sophisticated statistical modeling of prevalence. As a result, high-resolution prevalence maps for many indicators are routinely produced in the literature. However, statistical and practical guidance for producing prevalence maps in LMICs has been largely lacking. In particular, advice in choosing and evaluating models and interpreting results is needed, especially when data is limited. Software and analysis tools are also usually inaccessible to researchers in low-resource settings to conduct their own analysis or reproduce findings in the literature. In this paper, we propose a general workflow for prevalence mapping using household survey data. We consider all stages of the analysis pipeline, with particular emphasis on model choice and interpretation. We illustrate the proposed workflow using a case study mapping the proportion of pregnant women who had at least four antenatal care visits in Kenya. The workflow is implemented using the R package surveyPrev, and all reproducible code is provided in the Supplementary Materials. It can be readily extended to a wide range of indicators.
Small area estimation using survey data can be achieved by using either a design-based or a model-based inferential approach. Design-based direct estimators are generally preferable because of their consistency, asymptotic normality, and reliance on fewer assumptions. However, when data are sparse at the desired area level, as is often the case when measuring rare events, these direct estimators can have extremely large uncertainty, making a model-based approach preferable. A model-based approach with a random spatial effect borrows information from surrounding areas at the cost of inducing shrinkage. As a result, estimates may be over-smoothed and inconsistent with design-based estimates at higher area levels when aggregated. We propose two unit-level Bayesian models for small area estimation of rare event prevalence which use design-based direct estimates at a higher area level to increase consistency in aggregation. This model framework is designed to accommodate sparse data obtained from two-stage stratified cluster sampling, which is particularly relevant to applications in low- and middle-income countries. After introducing the model framework and its implementation, we conduct a simulation study to evaluate its properties and apply it to the estimation of the neonatal mortality rate in Zambia, using 2014 Demographic Health Surveys data.
Accurate fertility estimates at fine spatial resolution are essential for localized public health planning, particularly in low- and middle-income countries (LMICs). While national-level indicators such as age-specific fertility rates (ASFR) and total fertility rate (TFR) are often reported through official statistics, they lack the spatial granularity needed to guide targeted interventions. To address this, we develop a framework for subnational fertility estimation using small-area estimation (SAE) techniques applied to birth history data from household surveys, in particular Demographic and Health Surveys (DHS). Disaggregation by geographic area, time period, and maternal age group leads to significant data sparsity, limiting the reliability of direct estimates at fine scales. To overcome this, we propose a suite of methods, including direct estimators, area-level and unit-level Bayesian hierarchical models, to produce accurate estimates across varying spatial resolutions. The model-based approaches incorporate spatiotemporal smoothing and integrate covariates such as maternal education, contraceptive use and urbanicity. Using data from the 2021 Madagascar DHS, we generate district-level ASFR and TFR estimates and evaluate model performance through cross-validation.
Multilevel regression and poststratification (MRP) is a computationally efficient indirect estimation method that can quickly produce improved population-adjusted estimates with limited data. Recent computational advancements allow efficient, relatively simple, and quick approximate Bayesian estimation for MRP. As population health outcomes of interest including vaccination uptake are known to have spatial structure, precision may be gained by including space in the model. We test a recently proposed spatial MRP method that includes a BYM2 spatial term that smooths across demographics and geographic areas using a large, unrepresentative survey. We produce California county-level estimates of first-dose COVID-19 vaccination up to June 2021 using classic and spatial MRP models, and poststratify using data from the American Community Survey (US Census Bureau). We assess validity using reported first-dose vaccination counts from the Centers for Disease Control (CDC). Neither classic nor spatial MRP models performed well, highlighting: 1. spatial MRP may be most appropriate for richer data contexts, 2. some demographics in the survey data are over-sampled and -aggregated, producing model over-smoothing, and 3. a need for survey producers to share user-representative metrics to better benchmark estimates.
Prediction is a classic challenge in spatial statistics and the inclusion of spatial covariates can greatly improve predictive performance when incorporated into a model with latent spatial effects. It is desirable to develop flexible regression models that allow for nonlinearities and interactions in the covariate specification. Existing machine learning approaches that allow for spatial dependence in the residuals fail to provide reliable uncertainty estimates. In this paper, we investigate the combination of a Gaussian process spatial model with a Bayesian Additive Regression Tree (BART) model. The computational burden of the approach is reduced by combining Markov chain Monte Carlo (MCMC) with the Integrated Nested Laplace Approximation (INLA) technique. We study the performance of the method first via simulation. We then use the model to predict anthropometric responses in Kenya, with the data collected via a complex sampling design. In particular, household survey data are collected via stratified two-stage unequal probability cluster sampling, which requires special care when modeled.
Accurate subnational estimation of health indicators is critical for public health planning, particularly in low- and middle-income countries (LMICs), where data and analytic tools are often limited. sae4health is an open-access Shiny application (https://rsc.stat.washington.edu/sae4health/) that generates small area estimates for more than 150 demographic and health indicators, based on over 150 Demographic and Health Surveys (DHS) from 60 countries. The platform offers both area- and unit-level models with spatial random effects, implemented through fast Bayesian inference using Integrated Nested Laplace Approximation (INLA). The app is fully browser-based and requires no data input, programming skills, or statistical modeling expertise, making advanced methods accessible to a wide range of users. Estimates are processed in real time and presented as interactive maps, tables, and downloadable reports. A companion website (https://sae4health.stat.uw.edu) provides documentation and methodological background to support the app. Together, these resources enhance access to subnational health data and facilitate the use of DHS surveys for evidence-based decision making.
The under-5 mortality rate (U5MR), a critical health indicator, is typically estimated from household surveys in lower and middle income countries. Spatio-temporal disaggregation of household survey data can lead to highly variable estimates of U5MR, necessitating the usage of smoothing models which borrow information across space and time. The assumptions of common smoothing models may be unrealistic when certain time periods or regions are expected to have shocks in mortality relative to their neighbors, which can lead to oversmoothing of U5MR estimates. In this paper, we develop a spatial and temporal smoothing approach based on Gaussian Markov random field models which incorporate knowledge of these expected shocks in mortality. We demonstrate the potential for these models to improve upon alternatives not incorporating knowledge of expected shocks in a simulation study. We apply these models to estimate U5MR in Rwanda at the national level from 1985 to 2019, a time period which includes the Rwandan civil war and genocide.
In small area estimation, it is sometimes necessary to use model-based methods to produce estimates in areas with little or no data. In official statistics, we often require that aggregates of small area estimates agree with national estimates for internal consistency purposes. Enforcing this agreement is referred to as benchmarking, and while methods currently exist to perform benchmarking, few are ideal for applications with non-normal outcomes and benchmarks with uncertainty. Fully Bayesian benchmarking is a theoretically appealing approach insofar as we can obtain posterior distributions conditional on a benchmarking constraint. However, existing implementations may be computationally prohibitive. In this paper, we critically review benchmarking methods in the context of small area estimation in low- and middle-income countries with binary outcomes and uncertain benchmarks, and propose a novel approach in which posterior samples of small area characteristics from an unbenchmarked model can be combined with a rejection sampler or Metropolis-Hastings algorithm to produce benchmarked posterior distributions in a computationally efficient way. To illustrate the flexibility and efficiency of our approach, we provide comparisons to an existing benchmarking approach in a simulation, and applications to HIV prevalence and under-5 mortality estimation. Code implementing our methodology is available in the R package stbench.
Small area estimates of population are necessary for many epidemiological studies, yet their quality and accuracy are often not assessed. In the United States, small area estimates of population counts are published by the United States Census Bureau (USCB) in the form of the Decennial census counts, Intercensal population projections (PEP), and American Community Survey (ACS) estimates. Although there are significant relationships between these data sources, there are important contrasts in data collection and processing methodologies, such that each set of estimates may be subject to different sources and magnitudes of error. Additionally, these data sources do not report identical small area population counts due to post-survey adjustments specific to each data source. Resulting small area disease/mortality rates may differ depending on which data source is used for population counts (denominator data). To accurately capture annual small area population counts, and associated uncertainties, we present a Bayesian population model (B-Pop), which fuses information from all three USCB sources, accounting for data source specific methodologies and associated errors. The main features of our framework are: 1) a single model integrating multiple data sources, 2) accounting for data source specific data generating mechanisms, and specifically accounting for data source specific errors, and 3) prediction of estimates for years without USCB reported data. We focus our study on the 159 counties of Georgia, and produce estimates for years 2005-2021.
Subnational estimates of under-five mortality rates (U5MRs) are a vital statistic for the United Nations to reduce mortality inequalities between high-income and Low-and-Middle Income Countries (LMICs). Current methods of modelling U5MR in LMICs smooth across trends in age and year of death, but not birth-cohort, to reduce uncertainty in estimates caused by data-sparsity. Using survey data from Kenya, we innovatively apply an Age-Period-Cohort model which accounts for spatial trends and the complex survey design of the data to estimate subnational U5MRs in Kenya. After validating our results against current methods, the inclusion of cohort can provide new insights into U5MRs. We ensure our method is flexible and can be applied to other LMICs.
Doubly-stochastic point processes model the occurrence of events over a spatial domain as an inhomogeneous Poisson process conditioned on the realization of a random intensity function. They are flexible tools for capturing spatial heterogeneity and correlation. However, existing implementations of doubly-stochastic spatial models are computationally demanding, often have limited theoretical guarantee, and/or rely on restrictive assumptions. We propose a penalized regression method for estimating covariate effects in doubly-stochastic point processes that is computationally efficient and does not require a parametric form or stationarity of the underlying intensity. Our approach is based on an approximate (discrete and deterministic) formulation of the true (continuous and stochastic) intensity function. We show that consistency and asymptotic normality of the covariate effect estimates can be achieved despite the model misspecification, and develop a covariance estimator that leads to a valid, albeit conservative, statistical inference procedure. A simulation study shows the validity of our approach under less restrictive assumptions on the data generating mechanism, and an application to Seattle crime data demonstrates better prediction accuracy compared with existing alternatives.
We consider methods for model-based small area estimation when the number of areas with sampled data is a small fraction of the total areas for which estimates are required. Abundant auxiliary information is available from the survey for all the sampled areas. Further, through an external source, there is information for all areas. The goal is to use auxiliary variables to predict the outcome of interest for all areas. We compare areal-level random forests and LASSO approaches to a frequentist forward variable selection approach and a Bayesian shrinkage method using a horseshoe prior. Further, to measure the uncertainty of estimates obtained from random forests and the LASSO, we propose a modification of the split conformal procedure that relaxes the assumption of exchangeable data. We show that the proposed method yields intervals with the correct coverage rate and this is confirmed through a simulation study. This work is motivated by Ghanaian data available from the sixth Ghana Living Standards Survey (GLSS) and the 2010 Population and Housing Census, in the Greater Accra Metropolitan Area (GAMA) region, which comprises eight districts that are further divided into enumeration areas (EAs). We estimate the areal mean household log consumption using both datasets. The outcome variable is measured only in the GLSS for 3 percent of all the EAs (136 out of 5019) and 174 potential covariates are available in both datasets. In the application, among the four modeling methods considered, the Bayesian shrinkage performed the best in terms of bias, mean squared error (MSE), and prediction interval coverages and scores, as assessed through a cross-validation study. We find substantial between-area variation with the estimated log consumption showing a 1.3-fold variation across the GAMA region. The western areas are the poorest while the Accra Metropolitan Area district has the richest areas.
In low- and middle-income countries, household surveys are the most reliable data source to examine health and demographic indicators at the subnational level, an exercise in small area estimation. Model-based unit-level models are favored in producing the subnational estimates at fine scale, such as the admin-2 level. Typically, the surveys employ stratified two-stage cluster sampling with strata consisting of an urban/rural designation crossed with administrative regions. To avoid bias and increase predictive precision, the stratification should be acknowledged in the analysis. To move from the cluster to the area requires an aggregation step in which the prevalence surface is averaged with respect to population density. This requires estimating a partition of the study area into its urban and rural components, and to do this we experiment with a variety of classification algorithms, including logistic regression, Bayesian additive regression trees and gradient boosted trees. Pixel-level covariate surfaces are used to improve prediction. We estimate spatial HIV prevalence in women of age 15-49 in Malawi using the stratification/aggregation method we propose.
The problem of estimating the size of a population based on a subset of individuals observed across multiple data sources is often referred to as capture-recapture or multiple-systems estimation. This is fundamentally a missing data problem, where the number of unobserved individuals represents the missing data. As with any missing data problem, multiple-systems estimation requires users to make an untestable identifying assumption in order to estimate the population size from the observed data. If an appropriate identifying assumption cannot be found for a data set, no estimate of the population size should be produced based on that data set, as models with different identifying assumptions can produce arbitrarily different population size estimates-even with identical observed data fits. Approaches to multiple-systems estimation often do not explicitly specify identifying assumptions. This makes it difficult to decouple the specification of the model for the observed data from the identifying assumption and to provide justification for the identifying assumption. We present a re-framing of the multiple-systems estimation problem that leads to an approach that decouples the specification of the observed-data model from the identifying assumption, and discuss how common models fit into this framing. This approach takes advantage of existing software and facilitates various sensitivity analyses. We demonstrate our approach in a case study estimating the number of civilian casualties in the Kosovo war.
Estimating the mortality associated with a specific mortality crisis event (for example, a pandemic, natural disaster, or conflict) is clearly an important public health undertaking. In many situations, deaths may be directly or indirectly attributable to the mortality crisis event, and both contributions may be of interest. The totality of the mortality impact on the population (direct and indirect deaths) includes the knock-on effects of the event, such as a breakdown of the health care system, or increased mortality due to shortages of resources. Unfortunately, estimating the deaths directly attributable to the event is frequently problematic. Hence, the excess mortality, defined as the difference between the observed mortality and that which would have occurred in the absence of the crisis event, is an estimation target. If the region of interest contains a functioning vital registration system, so that the mortality is fully observed and reliable, then the only modeling required is to produce the expected deaths counts, but this is a nontrivial exercise. In low- and middle-income countries it is common for there to be incomplete (or nonexistent) mortality data, and one must then use additional data and/or modeling, including predicting mortality using auxiliary variables. We describe and review each of these aspects, give examples of excess mortality studies, and provide a case study on excess mortality across states of the United States during the COVID-19 pandemic.
This article is part of the Research Topic ‘ Health Systems Recovery in the Context of COVID-19 and Protracted Conflict ' Introduction After the World Health Organization declared COVID-19 a pandemic, more than 184 million cases and 4 million deaths had been recorded worldwide by July 2021. These are likely to be underestimates and do not distinguish between direct and indirect deaths resulting from disruptions in health care services. The purpose of our research was to assess the early impact of COVID-19 in 2020 and early 2021 on maternal and child healthcare service delivery at the district level in Mozambique using routine health information system data, and estimate associated excess maternal and child deaths. Methods Using data from Mozambique's routine health information system (SISMA, Sistema de Informação em Saúde para Monitoria e Avaliação), we conducted a time-series analysis to assess changes in nine selected indicators representing the continuum of maternal and child health care service provision in 159 districts in Mozambique. The dataset was extracted as counts of services provided from January 2017 to March 2021. Descriptive statistics were used for district comparisons, and district-specific time-series plots were produced. We used absolute differences or ratios for comparisons between observed data and modeled predictions as a measure of the magnitude of loss in service provision. Mortality estimates were performed using the Lives Saved Tool (LiST). Results All maternal and child health care service indicators that we assessed demonstrated service delivery disruptions (below 10% of the expected counts), with the number of new users of family planing and malaria treatment with Coartem (number of children under five treated) experiencing the largest disruptions. Immediate losses were observed in April 2020 for all indicators, with the exception of treatment of malaria with Coartem. The number of excess deaths estimated in 2020 due to loss of health service delivery were 11,337 (12.8%) children under five, 5,705 (11.3%) neonates, and 387 (7.6%) mothers. Conclusion Findings from our study support existing research showing the negative impact of COVID-19 on maternal and child health services utilization in sub-Saharan Africa. This study offers subnational and granular estimates of service loss that can be useful for health system recovery planning. To our knowledge, it is the first study on the early impacts of COVID-19 on maternal and child health care service utilization conducted in an African Portuguese-speaking country.
Accurate and precise estimates of under-5 mortality rates (U5MR) are an important health summary for countries. Full survival curves are additionally of interest to better understand the pattern of mortality in children under 5. Modern demographic methods for estimating a full mortality schedule for children have been developed for countries with good vital registration and reliable census data, but perform poorly in many low- and middle-income countries. In these countries, the need to utilize nationally representative surveys to estimate U5MR requires additional statistical care to mitigate potential biases in survey data, acknowledge the survey design, and handle aspects of survival data (i.e., censoring and truncation). In this paper, we develop parametric and non-parametric pseudo-likelihood approaches to estimating under-5 mortality across time from complex survey data. We argue that the parametric approach is particularly useful in scenarios where data are sparse and estimation may require stronger assumptions. The nonparametric approach provides an aid to model validation. We compare a variety of parametric models to three existing methods for obtaining a full survival curve for children under the age of 5, and argue that a parametric pseudo-likelihood approach is advantageous in low- and middle-income countries. We apply our proposed approaches to survey data from Burkina Faso, Malawi, Senegal, and Namibia. All code for fitting the models described in this paper is available in the R package pssst.
Monitoring subnational healthcare quality is important for identifying and addressing geographic inequities. Yet, health facility surveys are rarely powered to support the generation of estimates at more local levels. With this study, we propose an analytical approach for estimating both temporal and subnational patterns of healthcare quality indicators from health facility survey data. This method uses random effects to account for differences between survey instruments; space-time processes to leverage correlations in space and time; and covariates to incorporate auxiliary information. We applied this method for three countries in which at least four health facility surveys had been conducted since 1999 – Kenya, Senegal, and Tanzania – and estimated measures of sick-child care quality per WHO Service Availability and Readiness Assessment (SARA) guidelines at programmatic subnational level, between 1999 and 2020. Model performance metrics indicated good out-of-sample predictive validity, illustrating the potential utility of geospatial statistical models for health facility data. This method offers a way to jointly estimate indicators of healthcare quality over space and time, which could then provide insights to decision-makers and health service program managers.
In countries where population census data are limited, generating accurate subnational estimates of health and demographic indicators is challenging. Existing model-based geostatistical methods leverage covariate information and spatial smoothing to reduce the variability of estimates but often ignore the survey design, while traditional small area estimation approaches may not incorporate both unit-level covariate information and spatial smoothing in a design consistent way. We propose a smoothed model-assisted estimator that accounts for survey design and leverages both unit-level covariates and spatial smoothing. Under certain regularity assumptions, this estimator is both design consistent and model consistent. We compare it with existing design-based and model-based estimators using real and simulated data.