Small area estimation and disease mapping increasingly rely on areal data where reporting boundaries change over time. We develop a computationally efficient spatio-temporal disaggregation method to recover high-resolution risk surfaces from observed counts under changing boundaries. Our approach extends the spatially aggregated log-Gaussian Cox process and uses the Extended Latent Gaussian Model framework for fast approximate posterior inference. We replace standard lognormal polygon-specific effects with gamma-distributed overdispersion which yields a marginal negative binomial likelihood, and removes one latent variable per polygon-time pair. We illustrate the approach by mapping mortality risk across shifting NUTS-3 boundaries in Belgium and the Netherlands. For the purpose of dissemination we use Codex to leverage the methodology presented in this paper for the analysis of a separate data set concerning the city of Manchester. The methodology is implemented in the open-source R package DAST.
One in four of all global suicide deaths occurs in India, yet the epidemiology of suicide mortality in India remains largely undocumented. We analyzed over 20,000 suicides among 829,000 deaths collected from 2001-2019 within a nationally representative random sample of about 1% of Indian homes using lay field reporting with dual central medical adjudication of causes. We applied suicide death proportions to national demographic totals to estimate death rates and used proportional mortality to examine risk factors. While suicide death rates fell by 1.5% annually, the absolute number of suicides remained constant at around 200,000 annually (or 3.8 million from 2001 to 2019) due to rising population. Over this period, 44% and 50% of all suicides occurred at ages 15-29 and 30-69 years, respectively. Suicide death rates declined fastest in women aged 15-29 years, particularly after 2015. Suicide death rates from poisoning, mostly organophosphate pesticides, fell but those from hanging rose. Suicide death risks at ages 15-69 years were about eight times higher in selected southern high-burden states than in selected northern states. Individual suicide risk was highest among sons or daughters' in-law, rural residents, and men who drank alcohol. Suicide is preventable, but requires accelerating prevention policies to address acute social stress, ongoing efforts to reduce lethality of attempts and reliable epidemiological monitoring.
The accurate quantification of the impact of COVID-19 pandemic on both public health and the economy is essential for informed policy-making. However, the true scope of the pandemic remains challenging to ascertain due to undetected cases, particularly when relying on reported cases, which rely heavily on test availability and strategies. To accurately quantify COVID-19 cases in British Columbia (BC), we develop a Susceptible-Infectious-Recovered multi-event capture-recapture (SIRMECR) model to capture the dynamics of COVID-19. Specifically, we present a time-varying Markov model to estimate the number of undetected COVID-19 cases in five Health Authority Regions in BC, Canada, during the year 2020. We utilize individual-level information available from Population Data BC database to estimate the case detection probability, infection probability, survival probability, and recovery probability by incorporating testing volumes as covariates that improve the estimate of our parameters. We develop a Markov chain Monte Carlo (MCMC) algorithm to estimate SIRMECR model parameters. However, analyzing this big COVID-19 data set prompts a discussion on the computational challenges encountered. Therefore, we developed divide-and-conquer strategies to address the challenges. Our application provides an estimate of the total COVID-19 burden in year 2020 and found the percentage of undetected varying from 77.4 % to 84.0 %. More specifically, we validate our results through a simulation study and N-mixture model for Northern Health Authority Region of BC.
Verbal autopsies (VAs) collect information on deaths in low and middle-income countries occurring outside healthcare facilities to estimate causes of death (CODs) for use in epidemiological or planning studies. Physician coding of VAs focused on the narrative of deaths and past symptoms is current best practice. Large language models (LLM) such as GPT-5 enable possible use of the narrative portion of VAs to assign CODs. However, there are few if any robust comparisons of LLMs to physician coding. We analyzed 6,939 VA records from a random sample of deaths in Sierra Leone (2019–2022) to compare five models: three LLMs (GPT-3.5, GPT-4, GPT-5) and two based on symptom algorithms (InterVA-5, InSilicoVA), against physician-assigned CODs. GPT models used narratives, whereas InterVA-5 and InSilicoVA relied on questionnaires. CODs were grouped into 19, 10, and 7 categories for adult, child, and neonatal deaths. We used cause specific mortality fraction (CSMF) accuracy and partial chance corrected concordance (PCCC) to assess population and individual-level agreement respectively, compared to the standard of physician coding. We stratified analyses by age group as CODs vary among neonates, children and adults. Overall, GPT-5 outperformed all models (PCCC = 0.71), followed by GPT-4 (0.61), GPT-3.5 (0.56), InSilicoVA (0.44), and InterVA-5 (0.44). GPT-5 achieved the highest performance for adult (0.68), child (0.71), and neonatal (0.65) deaths. Across ages, performance increased from 1 month to 14 years and declined from 15 to 69 years. GPT-5, GPT-4, GPT-3.5, and InSilicoVA achieved the highest PCCC in 14, 7, 7, and 2 of the 30 CODs, respectively. At the population level, GPT-5 achieved the highest CSMF accuracy (0.9), while all other models had comparable performance (0.74–0.79). GPT models and InSilicoVA showed greater performance for specific CODs at the individual-level. GPT models demonstrated improvements over InterVA-5 and InSilicoVA models. This study provides foundational evidence for integrating LLM and algorithmic models with physician coding to improve the quality of VA data.
MDD is a leading cause of disability worldwide, yet its associations with mortality and hospitalization remain unclear. We aimed to quantify the risks of cause-specific mortality and hospitalization associated with a lifetime history of MDD, while accounting for potential reverse causality and residual confounding. We conducted a retrospective cohort study using data from 113,196 UK Biobank participants (62,822 women and 50,374 men), aged 40–69 years at recruitment, who were assessed between 2008 and 2010. Participants with pre-existing physical health conditions or mental or behavioural disorders other than MDD were excluded. Cause-specific Cox proportional hazards models estimated associations between MDD and mortality and hospitalization due to suicide, cerebrovascular and cardiac disease, respiratory diseases, and cancer. Models were sequentially adjusted for age, behavioural (smoking behaviour, alcohol use, body mass index, level of physical activity) and social (family status, education level, socioeconomic status, family history of MDD) factors. We focused on reductions in the log of the hazard ratios to assess potential residual confounding. MDD was associated with significantly increased risks of suicide mortality (HR = 8.52, 95
Adult survival among non-Hispanic white (‘white’) people in the USA has stagnated in recent decades, driven by rising mortality among adults with low educational attainment. We examined national mortality and population data to quantify the contributions of smoking-attributable diseases and selected other causes—notably opioids—to survival between ages 30 and 79 years from 1992 to 2019, stratified by three education levels: ≤11 years (low), 12 years (middle) and ≥13 years (high). We found that the absolute gap in the 50-year risk of death between low- and high-education white adults doubled among men and tripled among women. Smoking-attributable mortality increased markedly among low-education white adults but decreased in the two higher-education groups. Among 3.2 million excess premature deaths in low- and middle-education adults, smoking accounted for 63%. At ages 30–64 years, opioids accounted for 3% of excess deaths, rising to 8% during 2010–2019. Smoking-attributable mortality is the leading contributor to stagnating survival among white adults in the USA. Analysis of US mortality data from 1992 to 2019 shows widening educational inequalities in survival among non-Hispanic white adults, with smoking accounting for nearly two-thirds of excess premature deaths in lower-educated groups and outweighing the contributions of opioids.
Adult survival among non-Hispanic Whites (“Whites”) in the United States (US) has stagnated in recent decades, particularly among Whites with lower levels of education. We examined national mortality and population data to quantify the impact of smoking-attributable diseases, opioids, and other causes on survival between ages 30 and 79 from 1989 to 2023, stratified by education levels: ≤11 years (low), 12 years (middle), and ≥13 years (high). Absolute mortality rates widened sharply between the low- and middle-education groups and the high-education group. By 2023, the probability of death at 30-79 years was 30% for high-education Whites, compared to 78% for low-education Whites, accentuated by the COVID pandemic from 2020-22. From 1989 to 2019, smoking-attributable mortality rose substantially among low-education Whites but declined in middle- and high-education groups. Opioid mortality surged across all education levels at ages 30-64, especially after 2010. Among the 5.3 million excess premature deaths observed in low- and middle-education groups from 1989 to 2019, 58% were attributable to smoking. At ages 30-64, opioids caused 4% of the excess deaths from 1989 to 2009, rising to 9% from 2010 to 2019. Smoking remains a primary driver of stagnating survival among US Whites. Public health action on smoking, opioids and other diseases is achievable.
Quasi-periodicity refers to a pattern in a function where it appears periodic at a certain frequency but exhibits evolving amplitudes over time. This is often the case in practical settings such as the modeling of case counts of infectious disease or the population dynamics of species over time. In this paper, we consider a class of Gaussian processes, called seasonal Gaussian Processes (sGP), for model-based inference of such quasi-periodic behavior. We illustrate that the exact sGP can be efficiently fitted using its state space representation for equally spaced time points. However, for large datasets with irregular spacing, the exact approach becomes computationally inefficient and unstable. To address this, we develop a continuous finite dimensional approximation for sGP using the seasonal B-spline (sB-spline) basis constructed by damping B-splines with sinusoidal functions. We prove the covariance convergence rate of the proposed approximation to the true sGP as the number of basis functions increases, and show its superior approximation quality through numerical studies. We also provide a unified and interpretable way to define priors for the sGP, based on the notion of predictive standard deviation. Finally, we implement the proposed inference method on several real data examples to illustrate its practical usage.
We describe a polar coordinate transformation of the anisotropy parameters of the Mat & eacute;rn covariance function, which provides two benefits over the standard parameterization. First, it identifies a single point (the origin) with the special case of isotropy. Second, the posterior distribution of the transformed anisotropic angle and ratio is approximately bell-shaped and unimodal even in the case of isotropy. This has advantages for parameter inference and density estimation. We also apply a transformation to the standard deviation and range such that they are approximately orthogonal. We demonstrate this parameter transformation through two simulated and two real data sets, and conclude by considering possible extensions, such as implementing this transformation for approximate Bayesian inference methods. Les auteurs de cet article pr & eacute;sentent une transformation en coordonn & eacute;es polaires des param & egrave;tres d'anisotropie de la fonction de covariance de Mat & eacute;rn. Cette approche offre deux avantages par rapport & agrave; la param & eacute;trisation standard. Le premier avantage est l'association d'un point unique (l'origine) au cas particulier de l'isotropie. Le second est que la loi a posteriori de l'angle anisotrope transform & eacute; et du rapport prend une forme approximativement gaussienne et unimodale, m & ecirc;me en cas d'isotropie. Ces propri & eacute;t & eacute;s facilitent l'inf & eacute;rence des param & egrave;tres et l'estimation de densit & eacute;. Une transformation additionnelle de l'& eacute;cart-type et de la port & eacute;e permet de les rendre approximativement orthogonaux. L'efficacit & eacute; de cette transformation est d & eacute;montr & eacute;e sur deux jeux de donn & eacute;es simul & eacute;es et deux jeux de donn & eacute;es r & eacute;elles. Les perspectives incluent l'implantation de cette transformation dans des m & eacute;thodes d'inf & eacute;rence bay & eacute;sienne approximative.
Introduction:In the U.S., race and education have proven to be critical determinants of health outcomes inequalities, influenced by both location and time. Lung cancer (primarily linked to smoking), along with drug overdoses, alcohol poisoning, and suicide, emerge as important contributors to mortality risk. The purpose of this study is to enhance the understanding of geographic disparities in mortality related to lung cancer and external causes of death in the U.S. This study focuses on a finer geographic scale while examining the components of variation in mortality rates, stratifying by both education and race. Methods:This study used an aggregated spatial model and Bayesian inference to create continuous maps of risk. The study analyzes mortality counts from 1999 to 2015. This approach not only allows one to analyze and combine data sources at different spatial resolutions in a shorter time, but also facilitates standardized comparisons across age, sex, education, and spatial location. Results:Education emerged as the primary source of variation in both types of mortality, with growing disparities since 1999. Spatially, mortality risk varies more among the less educated and less among the more educated. Lung cancer rates have declined for the most educated but stayed steady for the less educated, while external causes of death rates rose for the less educated but remain unchanged for the more educated ones. Racial differences are significant for external causes of death but minimal for lung cancer. Conclusions:These findings highlight the need to go beyond education as a simple predictor variable in mortality studies, while also accounting for the differences in the nature of spatial variation. While the impact of education on mortality rates is unsurprising, the fact that Black Americans on average encounter inferior socioeconomic conditions compared to White Americans, may explain why this group is usually recognized at higher risk. This investigation suggests that addressing racial disparities in education levels could help to reduce tobacco-related mortality discrepancies effectively.
We propose a novel set of Poisson Cluster Process (PCP) models to detect Ultra-Diffuse Galaxies (UDGs), a class of extremely faint, enigmatic galaxies of substantial interest in modern astrophysics. We model the unobserved UDG locations as parent points in a PCP, and infer their positions based on the observed spatial point patterns of their old star cluster systems. Many UDGs have somewhere from a few to hundreds of these old star clusters, which we treat as offspring points in our models. We also present a new framework to construct a marked PCP model using the marks of star clusters. The marked PCP model may enhance the detection of UDGs and offers broad applicability to problems in other disciplines. To assess the overall model performance, we design an innovative assessment tool for spatial prediction problems where only point-referenced ground truth is available, overcoming the limitation of standard ROC analyses where spatial Boolean reference maps are required. We construct a bespoke blocked Gibbs adaptive spatial birth-death-move MCMC algorithm to infer the locations of UDGs using real data from a \textit{Hubble Space Telescope} imaging survey. Based on our performance assessment tool, our novel models significantly outperform existing approaches using the Log-Gaussian Cox Process. We also obtained preliminary evidence that the marked PCP model improves UDG detection performance compared to the model without marks. Furthermore, we find evidence of a potential new ``dark galaxy'' that was not detected by previous methods.
Our study explores the roles of precipitation and temperature in snakebite fatalities in India, with a focus on short-term effects and different lagged exposures. We propose the use of a spatial case-crossover model that accounts for spatially varying coefficients to assess these environmental exposures. While the spatial case-crossover model has primarily been applied to small area data, we extend its use to continuous spatial fields, allowing for more detailed regional analysis. The spatial model is implemented using MCMC (Markov Chain Monte Carlo) methods, allowing us to capture regional variations in the impacts of environmental factors on snakebite mortality. Our findings indicate that snakebite fatalities are primarily influenced by seasonality rather than precipitation or temperature, with notable spatial heterogeneity in these effects. This emphasizes the importance of spatially explicit models in understanding snakebite-related fatalities and the complexities of this public health challenge.
Since the beginning of the Covid-19 pandemic, public health authorities across the globe have implemented policies, such as lockdowns, in an attempt to reduce population mobility and, consequently, person-to-person contacts. It is well known that lockdowns reduce mobility, but to what extent does this reduction in mobility lead to lower infection rates? In this paper we extend the endemic-epidemic modeling framework in a principled manner, incorporating temporally changing mobility network data and quantifying the risk associated with travelling throughout the first year of the pandemic in two Spanish communities.
In Ethiopia, the reduction in perinatal mortality rates is still falling short of national and global targets set for 2030. Additionally, accurate recording is challenging, as many births occur at home. This study aimed to assess the trends and determinants of perinatal mortality using population-based longitudinal data from 2009 to 2016 across three Health and Demographic Surveillance Systems (HDSS) in Ethiopia: Gelgel-Gibe, Dabat, and Kilite-Awlaelo. Data on vital events and pregnancies were continuously collected at these HDSS sites. The study utilized follow-up data from prospective linked pregnancy and birth cohorts from January 2009 to December 31, 2016. Perinatal mortality was defined as deaths occurring from 28 weeks of gestation until six days after birth, measured per 1000 live births. Relevant health, demographic, and socioeconomic data were included in the analysis. Poisson regression was employed to assess factors associated with perinatal mortality. Out of 38,691 pregnancies that led to births, there were 1214 perinatal deaths (456 stillbirths and 758 early neonatal deaths), resulting in a perinatal mortality rate of 31 deaths per 1000 total births. The early neonatal death rate was higher, at 19.6 deaths per 1000 total births, compared to the stillbirth rate of 11.8 per 1000 total births. The perinatal mortality rate declined from 40.6 in 2009 to 29.1 per 1000 total births in 2016, reflecting an average annual rate reduction of 2.4%. Determinants of perinatal mortality included being a male newborn, multiple births, first-time pregnancies (primi-gravidity), lack of antenatal care visits, absence of delivery services, and residing in tropical zones. The primary causes of death were asphyxia, sepsis, and preterm birth. Overall, perinatal mortality rates were high in the three HDSS sites, with slow reductions over time and significant variations between them. Addressing the issue of stillbirths and improving the availability and quality of emergency obstetric care are crucial. Continuous home visits in rural communities to prevent stillbirths and newborn deaths, are also essential.
The majority of attempts to enumerate the homeless population rely on point-in-time or shelter counts, which can be costly and inaccurate. As an alternative, we use electronic health records from the Vancouver Island Health Authority, British Columbia, Canada from 2013 to 2022 to identify adults contending with homelessness based on their self-reported housing status. We estimate the annual population size of this population using a flexible open-population capture-recapture model that takes into account (1) the age and gender structure of the population, including aging across detection occasions, (2) annual recruitment into the population, (3) behavioural-response, and (4) apparent survival in the population, including emigration and incorporating known deaths. With this model, we demonstrate how to perform model selection for the inclusion of covariates. We then compare our estimates of annual population size with reported point-in-time counts of homeless populations on Vancouver Island over the same time period, and find that using data extracts from electronic health records gives comparable estimates. We find similarly comparable results using only a subset of interaction data, when using only ER interactions, suggesting that even if cross-continuum data is not available, reasonable estimates of population size can still be found using our method.
While SARS-CoV-2 infection appears to have spread widely throughout Africa, documentation of associated mortality is limited. We implemented a representative serosurvey in one city of Sierra Leone in Western Africa, paired with nationally representative mortality and selected death registration data. Cumulative seroincidence using high quality SARS-CoV-2 serological assays was 69% by July 2021, rising to 84% by April 2022, mostly preceding SARS-CoV-2 vaccination. About half of infections showed evidence of neutralizing antibodies. However, excess death rates were low, and were concentrated at older ages. During the peak weeks of viral activity, excess mortality rates were 22% for individuals aged 30-69 years and 70% for those over 70. Based on electronic verbal autopsy with dual independent physician assignment of causes, excess deaths during viral peaks from respiratory infections were notable. Excess deaths differed little across specific causes that, a priori, are associated with COVID, and the pattern was consistent among adults with or without chronic disease risk factors. The overall 6% excess of deaths at ages ≥30 from 2020-2022 in Sierra Leone is markedly lower than reported from South Africa, India, and Latin America. Thus, while SARS-CoV-2 infection was widespread, our study highlights as yet unidentified mechanisms of heterogeneity in susceptibility to severe disease in parts of Africa.