In many leishmaniasis foci, reservoirs that maintain infection remain unknown. Here, we developed a field-applicable toolkit based on the analysis of individual blood fed sand flies (IBF) to identify reservoirs. Sand flies were given a Leishmania donovani-infected first blood meal (iBM1) by feeding artificially on a membrane or naturally on a clinically ill animal followed by two subsequent uninfected blood meals (BMS+). Bulk-RNAseq was used to identify two target parasite genes, sherp and a novel hypothetical gene (HPB), which exhibited a significantly higher expression in BMS+ compared to iBM1 sand flies. DNA and RNA were co-extracted from IBF. DNA was used to detect Leishmania infection and the blood meal source; RNA was used to assess expression of target genes by qRT-PCR. Linear discriminant analysis (LDA) of target gene expression classified sand fly specimens based on their iBM1 or BMS+ status. Co-extraction yielded a mean of >800ng per IBF for DNA and RNA. We detected 31 parasite/s by kDNA qPCR and ssu rRNA RT-qPCR. LDA identified iBM1 parasites with a predictive accuracy of ~87% and ~82%, in membrane or naturally fed sand flies, respectively. This toolkit provides an innovative approach to identification of leishmaniasis reservoirs informing targeted control strategies.
A bstract Serological surveys measure the presence of antibodies in a population to infer past exposure to an infectious pathogen. If study participants’ ages are known, serocatalytic models can be used to retrace the historical transmission strength of a pathogen within that population, quantified by the force of infection (FOI). These models rely on age information as a key variable since infection risks are interpreted in relation to how long individuals have been at risk. However, due to data constraints, participants’ ages may be provided only within “age bins”. A common approach is then to assign individuals’ ages to midpoints of their respective age bins, ignoring uncertainty in this quantity. In this study, we quantify the bias introduced by this midpoint approach and develop a Bayesian framework that explicitly accounts for uncertainty in age. By comparing inference under constant, age-dependent, and time-dependent FOI scenarios, we show that the proposed binned model yields more reliable FOI estimates without sacrificing computational complexity, whereas the midpoint approach can underestimate FOI under constant transmission and introduce further biases as age bin width increases. These improvements support the interpretation of serological data and inform public health decisions, such as estimating disease burden and identifying targeted vaccination groups.
A bstract Serological assays remain the standard experimental approach for estimating the cumulative incidence of a pathogen and monitoring population immunity. The predominant approach for analysing serum titration data from virus neutralisation assays uses a nearly century-old interpolation-based method which neglects inherent imperfections in the assay and produces estimates with no measure of uncertainty. We introduce a two-part Bayesian modelling framework to estimate the underlying antibody concentrations in the raw serum samples taken from serosurveyed individuals, to improve the interpretation of serological data over age. First, we develop a mechanistic Bayesian model for serum antibody titration data that estimates latent antibody concentrations while accounting for assay variability and quantifying uncertainty. Second, we propagate this uncertainty into an age-structured serocatalytic model by integrating over posterior draws of individual antibody concentrations, allowing joint inference on latent serostate membership, force of infection, and serological waning rate. We use this framework to explore the dynamics of infection and immunity for three enterovirus serotypes: enteroviruses A71 (EV-A71) and D68 (EV-D68) and coxsackievirus A6 (CVA6). These serotypes are leading causes of outbreaks of severe respiratory illness and hand, foot, and mouth disease. Applying these approaches to three cross-sectional serosurveys, we estimated consistently higher and more persistent antibody concentrations throughout life for EV-D68 compared to EV-A71 and CVA6. Our analysis suggests that the proportion of recently infected individuals (i.e. individuals with high estimated antibody concentration levels given their age) peaks around 25% by age 7 years for both EV-A71 and CVA6 before gradually declining with age. In contrast, for EV-D68 the inferred proportion of the population in the infected state exceeds 50% by age 9 years and continues to grow with age. We also estimate that EV-D68 antibody concentration levels are higher than those of the other two serotypes, with the force of infection estimated to be highest in early childhood and declining more gradually with age than for EV-A71 and CVA6. These estimates are different to previous estimates found in the literature. Our inferential framework uncovers the wide-ranging variation in antibody levels that are often obscured by conventional endpoint titre estimation methods. We demonstrate that our framework can infer infection rates without relying on predetermined seropositivity cut-offs and without making explicit assumptions of virus-specific infection mechanisms. Author summary Serological tests measure antibody levels in blood to show how widely a virus has spread and how well populations are protected. Titre-based tests dilute blood samples in steps, mix these dilutions with virus, and add the mixture to living cells; the titre is the highest dilution where antibodies still protect cells from infection. Traditional analyses overlook test imperfections. We present a new two-part Bayesian framework to estimate antibody levels and track age-related exposure to infection. First, we estimate underlying antibody concentrations while accounting for uncertainty, then use these estimates in another model to infer age-specific transmission of three common viruses – EV-A71, EV-D68, and CVA6. Our results show that EV-D68 infections may be more common, especially in children, compared to the other viruses. This new approach provides a clearer picture of the dynamics of seroconversion, without relying on arbitrary thresholds, helping to improve public health monitoring and responses.
Due to its ability to summarise 'real-time' epidemic behaviour, the time-dependent reproduction number, Rt, is a useful metric for tracking pathogen transmission and quantifying the effects of interventions during infectious disease outbreaks. The predominant models underlying inferred Rt trajectories are renewal equations, their success owing in part to the relatively few assumptions they require. One necessary assumption is the generation time distribution, which summarises the time periods between infections in infector-infectee transmission pairs. This distribution is typically assumed to be the same across all members of a population. In reality, however, it may vary systematically between population groups. In this study, we consider two Rt inference frameworks based on renewal equation models: one for a single, homogeneous group and another accounting for a structured population. We compare the estimates of Rt generated by the two models and investigate, both analytically and through simulations, under which conditions the conclusions drawn from these modelling paradigms differ. We also demonstrate a methodology for selecting the generation time for the one-group model that correctly encapsulates variations between different population groups; this allows us to use a renewal framework for a one-group model to infer Rt when, in fact, the population is structured. Finally, we use real epidemic data to demonstrate that practical Rt estimates can differ depending on whether the underlying model is the one-group model or the multi-group model. Our results motivate the need for rigorous collection of detailed epidemic data and consideration of differences between population groups to improve the accuracy of Rt estimates that are used to guide public health policy responses.
Madariaga virus (MADV) has recently been associated with severe human disease in Panama, where the closely related Venezuelan equine encephalitis virus (VEEV) also circulates. In June 2017, a fatal MADV infection was confirmed in a community of Darien Province. We conducted a cross-sectional outbreak investigation with human and mosquito collections in July 2017, where sera were tested for alphavirus antibodies and viral RNA. In addition, by applying a catalytic, force-of-infection (FOI) statistical model to two serosurveys from Darien Province in 2012 and 2017, we investigated whether endemic or epidemic alphavirus transmission occurred historically. In 2017, MADV and VEEV IgM seroprevalences were 1.6% and 4.4%, respectively; IgG antibody prevalences were MADV: 13.2%, VEEV: 16.8%, Una virus (UNAV): 16.0%, and Mayaro virus: 1.1%. Active viral circulation was not detected. Evidence of MADV and UNAV infection was found near households, raising questions about its vectors and enzootic transmission cycles. Insomnia was associated with MADV and VEEV infections, depression symptoms were associated with MADV, and dizziness with VEEV and UNAV. Force-of-infection analyses suggest endemic alphavirus transmission historically, with recent increased human exposure to MADV and VEEV in Aruza and Mercadeo, respectively. The lack of additional neurological cases suggests that severe MADV and VEEV infections occur only rarely. Our results indicate that over the past five decades, alphavirus infections have occurred at low levels in eastern Panama, but that MADV and VEEV infections have recently increased-potentially during the past decade. Endemic infections and outbreaks of MADV and VEEV appear to differ spatially in some locations of eastern Panama.
When performing Bayesian inference, we frequently need to work with conditional probability densities. For example, the posterior function is the conditional density of the parameters given the data. Some might worry that conditional densities are ill-defined, considering that for a continuous random variable Y, the event {Y=y} has probability zero, meaning the formula ℙ(A|B)=ℙ(A∩ B)/ℙ(B) is inapplicable. In reality, when we work with conditional densities, we never condition directly on the zero-probability event {Y=y}; rather, we first condition on the random variable Y, and then we may plug in an observed value y. The first purpose of our article is to provide an exposition on conditional densities that elaborates on this point. While we have aimed to make this explanation accessible, we follow it with a roadmap of the measure theory needed to make it rigorous. A recent preprint (arXiv:2411.13570) has expressed the concern that probability densities are ill-defined and that as a result Bayes' theorem cannot be used, and they provide examples that allegedly demonstrate inconsistencies in the Bayesian framework. The second purpose of our article is to investigate their claims. We contend that the examples given in their work do not demonstrate any inconsistencies; we find that there are mathematical errors and that they deviate significantly from the Bayesian framework.
Introduction We retrospectively evaluated the impact of COVID-19 testing among residents and staff in social care homes in England.Methods We obtained 80 million reported PCR and lateral flow device (LFD) test results, from 14 805 care homes (residents and staff) in England, conducted between October 2020 and March 2022. These testing data were then linked to care home characteristics, test costs and 24 500 COVID-19-related deaths of residents. We decomposed the mechanism of outbreak mitigation into outbreak discovery and outbreak control and used Poisson regressions to investigate how reported testing intensity was associated with the size of outbreak discovered and to uncover its association with outbreak control. We used negative binomial regressions to determine the factors influencing COVID-19-related deaths subsequent to outbreaks. We performed a cost-effectiveness analysis of the impact of testing on preventing COVID-19-related deaths of residents.Results Reported testing intensity generally reflected changes in testing policy over time, although there was considerable heterogeneity among care homes. Client type was the strongest determinant of whether COVID-19-related deaths in residents occurred subsequent to testing positive. Higher staff-to-resident ratios were associated with larger outbreak sizes but rapid outbreak control and a decreased risk of COVID-19-related deaths. Assuming our regression estimates represent causal effects, care home testing in England was cost-effective at preventing COVID-19-related deaths among residents during the pandemic and approximately 3.5 times more cost-effective prior to the vaccine rollout.Conclusions PCR and LFD testing was likely an impactful intervention for detecting and controlling COVID-19 outbreaks in care homes in England and cost-effective for preventing COVID-19-related deaths among residents. In future pandemics, testing must be prioritised for care homes, especially if severe illness and death particularly affect older people or individuals with characteristics similar to care home residents, and an efficacious vaccine is unavailable.
Entomological surveillance is an important component of mosquito-borne disease control. Mosquito abundance, infection prevalence and the entomological inoculation rate are the most widely reported entomological metrics, although these data are notoriously noisy and difficult to interpret. For many infections, only older mosquitoes are infectious, which is why, in part, vector control tools that reduce mosquito life expectancy have been so successful. The age structure of wild mosquitoes has been proposed as a metric to assess the effectiveness of interventions that kill adult mosquitoes, and age grading tools are becoming increasingly advanced. Mosquito populations show seasonal dynamics with temporal fluctuations. How seasonal changes in adult mosquito emergence and vector control could affect the mosquito age distribution or other important metrics is unclear. We develop stochastic mathematical models of mosquito population dynamics to show how variability in mosquito emergence causes substantial heterogeneity in the mosquito age distribution, with low frequency, positively autocorrelated changes in emergence being the most important driver of this variability. Fitting a population model to mosquito abundance data collected in experimental hut trials indicates these dynamics are likely to exist in wild Anopheles gambiae populations. Incorporating age structuring into an established compartmental model of mosquito dynamics and vector control, indicates that the use of mosquito age as a metric to assess the efficacy of vector-control tools will require an understanding of underlying variability in mosquito ages, with the mean age and other entomological metrics affected by short-term and seasonal fluctuations in mosquito emergence.
Mathematical models play a crucial role in understanding the spread of infectious disease outbreaks and influencing policy decisions. These models have aided pandemic preparedness by predicting outcomes under hypothetical scenarios and identifying weaknesses in existing frameworks; however, their accuracy, utility, and comparability are being scrutinised. Agent-based models (ABMs) have emerged as a valuable tool, capturing population heterogeneity and spatial effects, particularly when assessing potential intervention strategies. Here we present EpiGeoPop, a user-friendly tool for rapidly preparing spatially accurate population configurations of entire countries. EpiGeoPop helps to address the problem of complex and time-consuming model set-up in ABMs, specifically improving the integration of real-world spatial detail. We subsequently demonstrate the importance of accurate spatial detail in ABM simulations of disease outbreaks using Epiabm, an ABM based on Imperial College London's CovidSim with improved modularity, documentation and testing. Our simulations present a number of possible applications of ABMs where including spatially accurate data is crucial, highlighting the potential impact of EpiGeoPop in facilitating this process using multiple international data sources.
The injury rate is a common measure of injury occurrence in epidemiological surveillance and is used to express the incidence of injuries as a function of both the population at risk as well as at-risk exposure time. Traditional approaches to surveillance-based injury rates use a frequentist perspective; here, we discuss the Bayesian perspective and present a practical framework on how to apply a Bayesian analysis to estimate injury rates. We estimated finescale injury rates across a broad range of categories for men’s and women’s soccer, applying a Bayesian methodology and using injury surveillance data captured within the National Collegiate Athletic Association Injury Surveillance Program from 2014/15–2018/19. Through an iterative process of assessing model fidelity, we found that a negative binomial model was an effective choice for modeling surveillance-based injury rates. We also found differences between schools to be a key driver of variation in injury rates. Our findings indicate that the Bayesian framework naturally characterizes injury rates by modeling injury counts as outcomes of an underlying data-generation process that explicitly incorporates inherent uncertainty, complementing traditional frequentist approaches. Key benefits of the Bayesian approach in this context are the ability to test model suitability in a variety of methods, and to be able to generate plausible estimates with sparse data.
The time-varying reproduction number ($R_t$) gives an indication of the trajectory of an infectious disease outbreak. Commonly used frameworks for inferring $R_t$ from epidemiological time series include those based on compartmental models (such as the SEIR model) and renewal equation models. These inference methods are usually validated using synthetic data generated from a simple model, often from the same class of model as the inference framework. However, in a real outbreak the transmission processes, and thus the infection data collected, are much more complex. The performance of common $R_t$ inference methods on data with similar complexity to real world scenarios has been subject to less comprehensive validation. We therefore propose evaluating these inference methods on outbreak data generated from a sophisticated, geographically accurate agent-based model. We illustrate this proposed method by generating synthetic data for two outbreaks in Northern Ireland: one with minimal spatial heterogeneity, and one with additional heterogeneity. We find that the simple SEIR model struggles with the greater heterogeneity, while the renewal equation model demonstrates greater robustness to spatial heterogeneity, though is sensitive to the accuracy of the generation time distribution used in inference. Our approach represents a principled way to benchmark epidemiological inference tools and is built upon an open-source software platform for reproducible epidemic simulation and inference.
Malaria vector control tools currently focus on insecticide treated nets (ITNs) and indoor residual spraying in malaria-endemic locations, but additional preventative strategies are needed to address protection gaps. Larval source management (LSM) includes larvicide application to aquatic habitat and an array of alternative forms of environmental efforts. An individual-based transmission model for falciparum malaria is used to demonstrate the theoretical benefit of suppressing malaria adult mosquito vector densities through LSM. The model simulates results of epidemiological trials from Western Kenya (a hilly area with papyrus swamps adjacent to human settlements and moderate to high perennial malaria transmission) and Côte d'Ivoire (an area with Sudanese climate, reducing vegetation cover and high transmission) that applied larvicide alongside ITNs, and investigates whether estimated changes in adult density can be used to project changes in human malaria. In the Western Kenya setting generalised linear models estimate 82% (90% credible intervals: 64% - 92%) and 88% (79% - 94%) reductions in the proportion of adult Anopheles funestus and Anopheles gambiae complex mosquitoes respectively as measured by CDC light traps. In Côte d'Ivoire, an 82% (56% - 93%) reduction of the dominant An. gambiae vector was estimated using standard window trap and pyrethrum spray catch. Both studies had variable village-level impacts. The transmission dynamics model predicted that these entomological impacts would result in a reduction in malaria prevalence in children of 6-months to 10-years of age of 48 - 72% in Kenya, and a 11 - 78% reduction in all-age clinical incidence across villages in Côte d'Ivoire, which are broadly consistent with the empirically observed outcomes. High heterogeneity between villages within the same study indicate that the relative or absolute reductions in mosquito adult density observed in these trials cannot be simply extrapolated to other regions. The LSM strategy adopted, unit area covered, and multiple environmental covariates all contribute to differences in indicators that could be used to assess entomological impacts and the corresponding epidemiological outcomes. This important malaria control tool was impactful across all sites examined, though further work is needed to understand how best to use this tool in the fight against malaria.
COVID-19 data exhibit various biases, not least a significant weekly periodic oscillation observed globally in case and death data. There has been significant debate over whether this may be attributed to weekly socialising and working patterns, or is due to underlying biases in the reporting process. We characterise the weekly biases globally and demonstrate that equivalent biases also occur in the current cholera outbreak in Haiti. By comparing published COVID-19 time series to retrospective datasets from the United Kingdom (UK) that are not subject to the same reporting biases, we demonstrate that this dataset does not contain any weekly periodicity, and hence the weekly trends observed both in the UK and globally may be fully explained by biases in the testing and reporting processes. These conclusions play an important role in forecasting healthcare demand and determining suitable interventions for future infectious disease outbreaks.
During infectious disease outbreaks, delays in case reporting mean that the time series of cases is unreliable, particularly for those cases occurring most recently. This means that real-time estimates of the time-varying reproduction number, R t , are often made using a time series of cases only up until a time period sufficiently far in the past that there is some confidence in the case counts. This means that the most recent R t estimates are usually out of date, inducing lags in the response of public health authorities. Here, we introduce an R t estimation method, which makes use of the retrospective updates to case time series which happen as more cases that occurred historically enter the health system; these data encode within them information about the reporting delays, which our method also estimates. These estimates, in turn, allow us to estimate the true count of cases occurring most recently allowing up-to-date estimates of R t . Our method simultaneously estimates the reporting delays, true historical case counts and R t in a single Bayesian framework, allowing the uncertainty in each of these quantities to be accounted for. We apply our method to both simulated and real outbreak data, which shows that the method substantially improves upon naive estimates of R t which do not account for reporting delays. Our method is available in an open-source fully tested R package, incidenceinflation . Our research highlights the value of keeping historical time series of cases since changes to these data can help to characterize nuisance processes, such as reporting delays, which allow these to be accounted for when estimating key epidemic quantities. This article is part of the theme issue ‘Uncertainty quantification for healthcare and biological systems (Part 1)’.
The time-dependent reproduction number Rt can be used to track pathogen transmission and to assess the efficacy of interventions. This quantity can be estimated by fitting renewal equation models to time series of infectious disease case counts. These models almost invariably assume a homogeneous population. Individuals are assumed not to differ systematically in the rates at which they come into contact with others. It is also assumed that the typical time that elapses between one case and those it causes (known as the generation time distribution) does not differ across groups. But contact patterns are known to widely differ by age and according to other demographic groupings, and infection risk and transmission rates have been shown to vary across groups for a range of directly transmitted diseases. Here, we derive from first principles a renewal equation framework which accounts for these differences in transmission across groups. We use a generalisation of the classic McKendrick-von Foerster equation to handle populations structured into interacting groups. This system of partial differential equations allows us to derive a simple analytical expression for Rt which involves only group-level contact patterns and infection risks. We show that the same expression emerges from both deterministic and stochastic discrete-time versions of the model and demonstrate via simulations that our Rt expression governs the long-run fate of epidemics. Our renewal equation model provides a basis from which to account for more realistic, diverse populations in epidemiological models and opens the door to inferential approaches which use known group characteristics to estimate Rt.
During infectious disease outbreaks, estimates of time-varying pathogen transmissibility, such as the instantaneous reproduction number Rt$$ R(t) $$ or epidemic growth rate rt$$ {r}_t $$, are used to inform decision-making by public health authorities. For directly transmitted infectious diseases, the renewal equation framework is a widely used method for measuring time-varying transmissibility. The framework uses information on the typical time elapsing between an infection and the offspring infections (quantified by the generation time distribution), and Rt$$ R(t) $$, to describe the rate at which currently infected individuals generate new infections, if conditions affecting human transmission were to remain the same as at time t$$ t $$. For diseases with transmission cycles involving hosts and vectors, however, renewal equation models have been far less used. This is likely due to difficulties in mechanistically defining generation times that can capture the complexity of multi-stage, human-vector relationships. Here, using dengue as an example, we provide general renewal equations that are derived from first principles using age-structured systems of coupled partial differential equations across human and vector sub-populations. Our framework tracks the multi-stage transmission cycle over calendar time and across stage-specific ages, resulting in governing renewal equations that quantify how the rate at which new infections are generated from existing infections depends on stage-specific processes. In a numerical application to real-world temperature data, we show how the generation time distribution depends on both current and historical conditions. The framework provides a foundation on which to base inferential frameworks for estimating Rt$$ R(t) $$ and rt$$ {r}_t $$ for infectious diseases with multiple stages in the transmission cycle.
Deciding on when to initiate or relax an intervention in response to an emerging infectious disease is both difficult and important. Uncertainties from noise in epidemiological surveillance data must be hedged against the potentially unknown and variable costs of false alarms and delayed actions. Here, we clarify and quantify how case under-reporting and latencies in case ascertainment, which are predominant surveillance noise sources, can restrict the timeliness of decision-making. Decisions are modelled as binary choices between responding or not that are informed by reported case curves or transmissibility estimates from those curves. Optimal responses are triggered by thresholds on case numbers or estimated confidence levels, with thresholds set by the costs of the various choices. We show that, for growing epidemics, both noise sources induce additive delays on hitting any case-based thresholds and multiplicative reductions in our confidence in estimated reproduction numbers or growth rates. However, for declining epidemics, these noise sources have counteracting effects on case data and limited cumulative impact on transmissibility estimates. We find that this asymmetry persists even if more sophisticated feedback control algorithms that consider the longer-term effects of interventions are employed. Standard surveillance data, therefore, provide substantially weaker support for deciding when to initiate a control action or intervention than for determining when to relax it. This information bottleneck during epidemic growth may justify proactive intervention choices.
Serocatalytic models are powerful tools which can be used to infer historical infection patterns from age-structured serological surveys. These surveys are especially useful when disease surveillance is limited and have an important role to play in providing a ground truth gauge of infection burden. In this tutorial, we consider a wide range of serocatalytic models to generate epidemiological insights. With mathematical analysis, we explore the properties and intuition behind these models and include applications to real data for a range of pathogens and epidemiological scenarios. We also include practical steps and code in R and Stan for interested learners to build experience with this modeling framework. Our work highlights the usefulness of serocatalytic models and shows that accounting for the epidemiological context is crucial when using these models to understand infectious disease epidemiology.
Mosquito infection experiments that characterise how sporogony changes with temperature are increasingly being used to parameterise malaria transmission models. In these experiments, mosquitoes are exposed to a range of temperatures, with each group experiencing a single temperature. Diurnal temperature variation can, however, affect the sporogonic cycle of Plasmodium parasites. Mosquito dissection data is not available for all temperature profiles, so we investigate whether mathematical models of mosquito infection parameterised with constant temperature thermal performance curves can predict the effects of diurnal temperature variation. We use this model to predict two key parameters governing disease transmission: the human-to-mosquito transmission probability and extrinsic incubation period - and, embed this model into a malaria transmission model to simulate sporozoite prevalence with and without the effects of diurnal and seasonal temperature variation for a single site in Burkina Faso. Simulations incorporating diurnal temperature variation better predict changes in sporogony in laboratory mosquitoes, indicating that constant temperature experiments can be used to predict the effects of fluctuating temperatures. Including the effects of diurnal temperature variation, however, did not substantially improve the predictive ability of the transmission model to predict changes in sporozoite prevalence in wild mosquitoes, indicating further research is needed in more settings.