Abstract As global temperatures rise, there is growing evidence that extreme heat will have a detrimental impact on the lives and livelihoods of workers. Despite adaptations being implemented to protect workers, there remain ascertainment challenges in identifying optimal solutions. We need higher spatial and temporal resolution of heat exposure data that is measured longitudinally─day and night, home and workplace, and across seasons and years. Additionally, we need to improve our understanding of how heat impacts individuals; thresholds for heat are dependent on the outcome of interest and the individuals exposed. Through technological advances in environmental sensors and wearable physiological monitors, it is now possible to pursue personalized occupational heat exposure research in situ and at scale, by measuring the longitudinal physiological and health impacts of varying heat exposures. We propose the development of a Personalized Occupational Heat Exposure Index (POHEI) that combines individualized exposures with physiological impact, enabling in situ evaluation of adaptation interventions to protect the lives and livelihoods of workers most at risk of rising heat.
BACKGROUND:Digital data sources such as mobile phone call detail records (CDRs) are increasingly being used to estimate population mobility fluxes and to predict the spatiotemporal dynamics of infectious disease outbreaks. Differences in mobile phone operators' geographic coverage, however, may result in biased mobility estimates. METHODS:We leverage a unique dataset consisting of CDRs from three mobile phone operators in Bangladesh and digital trace data from Meta's Data for Good program to compare mobility patterns across these sources. We use a metapopulation model to compare the sources' effects on simulated outbreak trajectories, and compare results with a benchmark model with data from all three operators, representing around 100 million subscribers across the country. RESULTS:We show that mobility sources can vary significantly in their coverage of travel routes and geographic mobility patterns. Differences in projected outbreak dynamics are more pronounced at finer spatial scales, especially if the outbreak is seeded in smaller and/or geographically isolated regions. In some instances, a simple diffusion (gravity) model was better able to capture the timing and spatial spread of the outbreak compared to the sparser mobility sources. CONCLUSIONS:Our results highlight the potential biases in predicted outbreak dynamics from a metapopulation model parameterized with non-population representative data, and the limits to the generalizability of models built on these types of novel human behavioral data.
Objective/backgroundTransmission-dynamic models are commonly used to study infectious disease epidemiology. Calibration involves identifying model parameter values that align model outputs with observed data or other evidence. Inaccurate calibration and inconsistent reporting produce inference errors and limit reproducibility, compromising confidence in the validity of modeled results. No standardized framework exists for reporting on calibration of infectious disease models, and an understanding of current calibration approaches is lacking.MethodsWe developed the Purpose-Inputs-Process-Outputs (PIPO) framework for reporting calibration practices and applied it in a scoping review to assess calibration approaches and evaluate reporting comprehensiveness in transmission-dynamic models of tuberculosis, HIV and malaria published between January 1, 2018, and January 16, 2024. We searched relevant databases and websites to identify eligible publications, including peer-reviewed studies where these models were calibrated to empirical data or published estimates.ResultsWe identified 411 eligible studies encompassing 419 models, with 74% (n = 309) being compartmental models and 20% (n = 81) individual-based models (IBMs). The predominant analytical purpose was to evaluate interventions (71% of models, n = 298). Parameters were calibrated mainly because they were unknown or ambiguous (40%, n = 168), or because determining their value was relevant to the scientific question beyond being necessary to run the model (20%, n = 85). The choice of calibration method was significantly associated with model structure (p-value<0.001) and stochasticity (p-value = 0.006), with approximate Bayesian computation more frequently used with IBMs and Markov-Chain Monte Carlo with compartmental models. Regarding reporting comprehensiveness, all PIPO framework items were reported in 4% (n = 18) of models; 11-14 items in 66% (n = 277), and 10 or fewer items in 28% (n = 124). Implementation code was the least reported, available in only 20% (n = 82) of models.ConclusionsReporting on calibration is heterogeneous in recent infectious disease modeling literature. Our proposed framework for reporting of calibration approaches could support improved reproducibility and credibility of modeled analyses.
Systematic, long-term, and spatially representative monitoring of insecticide resistance in mosquito populations is urgently needed to quantify its impact on malaria transmission, and to combat failing interventions when resistance emerges. Resistance assays on wild-caught adult mosquitoes (known as adult-capture) offer an alternative to the current protocols, which recommend larval capture. Adult-capture assays can be done in a shorter time frame, in more locations, and in the absence of an insectary. However, unlike insectary-raised mosquitoes, a group of adults captured in the wild represents different ages and may have previous exposure to insecticides. Since age and prior exposure are critically important in determining the likelihood of death during the assay, taking these factors into account is important for assessing the relative utility of the assay. Currently such quantitative assessments are lacking. We developed a discrete-time deterministic model to simulate the mosquito life cycle, including insecticide exposure due to insecticide-treated bed nets. We incorporated non-lethal effects of insecticide exposure demonstrated in laboratory experiments and the impact of multiple exposure to insecticides on mosquito death rates during the assay. We then sampled from this population using both larval-captured and adult-captured mosquito collection and simulated insecticide resistance assays. To quantify possible biases in adult-capture assays, we compared the results of these assays to the true resistance allele frequency in the population. In simulated samples of 100 test mosquitoes, reflecting WHO-recommended sample sizes, we found that adult-capture samples had a 94% positive predictive value (PPV) for resistance at the WHO’s 10% resistance cutoff, and a 97% negative predictive value (NPV), compared to 98% PPV and 19% NPV for larval-captured samples. Bias in the adult-capture assays was primarily dependent on the level of insecticide resistance rather than coverage of bed nets or exposure heterogeneity. Using adult-captured mosquitoes for resistance assays may have advantages over larval-capture collection in many settings, and in our model does not appear to be significantly less accurate than larval-capture, especially when used to categorize resistance under the binary WHO criteria. These results suggest that adult-captured assays could be deployed for resistance monitoring programs at a more widespread scale.
Variation in malaria infection risk, a product of disease exposure and immunity, is poorly understood. We genotypically profiled over 13,000 blood samples from a six-year longitudinal cohort in Mali to characterize malaria infection dynamics with detail. We generated Plasmodium falciparum amplicon sequencing data from 464 participants (aged 3 months - 25 years) across the six-month 2011 transmission season and profiled a subset of 120 participants across the subsequent five annual transmission seasons. We measured infection rate as the molecular force of infection (molFOI, number of genetically distinct parasites acquired over time). We found that molFOI varied extensively among individuals (0-55 in 2011) but was independent of age and consistent within individuals over multiple seasons. Reported bednet usage was nearly universal. The HbS allele was associated with lower molFOI, and functional antibody signatures for the CSP C-term and RH5 antigens were correlated with low molFOI participants, identifying candidate immune correlates of protection. The large inter-individual variability in molFOI and consistency of intra-individual infection rate over time exhibits much greater dynamic range than malaria case incidence, and is most likely due to heterogeneous exposure to infectious mosquito bites. This and other factors contributing to variable infection risk should be considered in future clinical trials and implementation of malaria interventions.
Background:Mathematical modeling of infectious diseases is an important decision-making tool for outbreak control. However, in Africa, limited expertise reduces the use and impact of these tools on policy. Therefore, there is a need to build capacity in Africa for the use of mathematical modeling to inform policy. Here we describe our experience implementing a mathematical modeling training program for public health professionals in East Africa. Methods:We used a deliverable-driven and learning-by-doing model to introduce trainees to the mathematical modeling of infectious diseases. The training comprised two two-week in-person sessions and a practicum where trainees received intensive mentorship. Trainees evaluated the content and structure of the course at the end of each week, and this feedback informed the strategy for subsequent weeks. Findings:Out of 875 applications from 38 countries, we selected ten trainees from three countries - Rwanda (6), Kenya (2), and Uganda (2) - with guidance from an advisory committee. Nine trainees were based at government institutions and one at an academic organization. Participants gained skills in developing models to answer questions of interest and critically appraising modeling studies. At the end of the training, trainees prepared policy briefs summarizing their modeling study findings. These were presented at a dissemination event to policymakers, researchers, and program managers. All trainees indicated they would recommend the course to colleagues and rated the quality of the training with a median score of 9/10. Conclusions:Mathematical modeling training programs for public health professionals in Africa can be an effective tool for research capacity building and policy support to mitigate infectious disease burden and forecast resources. Overall, the course was successful, owing to a combination of factors, including institutional support, trainees' commitment, intensive mentorship, a diverse trainee pool, and regular evaluations.
OBJECTIVE:Mathematical models are vital tools to understand transmission dynamics and assess the impact of interventions to mitigate COVID-19. However, historically, their use in Africa has been limited. In this scoping review, we assess how mathematical models were used to study COVID-19 vaccination to potentially inform pandemic planning and response in Africa. METHODS:We searched six electronic databases: MEDLINE, Embase, Web of Science, Global Health, MathSciNet and Africa-Wide NiPAD, using keywords to identify articles focused on the use of mathematical modelling studies of COVID-19 vaccination in Africa that were published as of October 2022. We extracted the details on the country, author affiliation, characteristics of models, policy intent and heterogeneity factors. We assessed quality using 21-point scale criteria on model characteristics and content of the studies. RESULTS:The literature search yielded 462 articles, of which 32 were included based on the eligibility criteria. Nineteen (59%) studies had a first author affiliated with an African country. Of the 32 included studies, 30 (94%) were compartmental models. By country, most studies were about or included South Africa (n = 12, 37%), followed by Morocco (n = 6, 19%) and Ethiopia (n = 5, 16%). Most studies (n = 19, 59%) assessed the impact of increasing vaccination coverage on COVID-19 burden. Half (n = 16, 50%) had policy intent: prioritising or selecting interventions, pandemic planning and response, vaccine distribution and optimisation strategies and understanding transmission dynamics of COVID-19. Fourteen studies (44%) were of medium quality and eight (25%) were of high quality. CONCLUSIONS:While decision-makers could draw vital insights from the evidence generated from mathematical modelling to inform policy, we found that there was limited use of such models exploring vaccination impacts for COVID-19 in Africa. The disparity can be addressed by scaling up mathematical modelling training, increasing collaborative opportunities between modellers and policymakers, and increasing access to funding.
Importance:Post-acute sequelae of SARS-CoV-2, referred to as "long COVID", are a globally pervasive threat. While their many clinical determinants are commonly considered, their plausible social correlates are often overlooked. Objective:To compare social and clinical predictors of differences in quality of life (QoL) with long COVID. Additionally, to measure how much adjusted associations between social factors and long COVID-associated quality of life are unexplained by important clinical intermediates. Design Setting and Participants:Data from the ISARIC long COVID multi-country prospective cohort study. Subjects from Norway, the United Kingdom (UK), and Russia, aged 16 and above, with confirmed acute SARS-CoV-2 infection reporting >= 1 long COVID-associated symptoms 1+ month following infection. Exposure:The social exposures considered were educational attainment (Norway), employment status (UK and Russia), and female vs male sex (all countries). Main outcome and measures:Quality of life-adjusted days, or QALDs, with long COVID. Results:This cohort study included a total of 3891 participants. In all three countries, educational attainment, employment status, and female sex were important predictors of long COVID QALDs. Furthermore, a majority of the estimated relationships between each of these social correlates and long COVID QALDs could not be attributed to key long COVID-predicting comorbidities. In Norway, 90% (95% CI: 77%, 100%) of the adjusted association between the top two quintiles of educational attainment and long COVID QALDs was not explained by clinical intermediates. The same was true for 86% (73%, 100%) and 93% (80%,100%) of the adjusted associations between full-time employment and long COVID QALDs in the United Kingdom (UK) and Russia. Additionally, 77% (46%,100%) and 73% (52%, 94%) of the adjusted associations between female sex and long COVID QALDs in Norway and the UK were unexplained by the clinical mediators. Conclusions and Relevance:This study highlights the role of socio-economic status indicators and female sex, in line with or beyond commonly cited clinical conditions, as predictors of long COVID-associated QoL, and further reveal that other (non-clinical) mechanisms likely drive their observed relationships. Our findings point to the importance of COVID interventions which go further than an exclusive focus on comorbidity management in order to help redress inequalities in experiences with this chronic disease.
Plasmodium parasites, the causal agents of malaria, are eukaryotic organisms that obligately undergo sexual recombination within mosquitoes. In low transmission settings, parasites recombine with themselves, and the clonal lineage is propagated rather than broken up by outcrossing. We investigated whether stochastic/neutral factors drive the persistence and abundance of Plasmodium falciparum clonal lineages in Guyana, a country with relatively low malaria transmission, but the only setting in the Americas in which an important artemisinin resistance mutation (pfk13 C580Y) has been observed. We performed whole genome sequencing on 1,727 Plasmodium falciparum samples collected from infected patients across a five-year period (2016-2021). We characterized the relatedness between each pair of monoclonal infections (n = 1,409) through estimation of identity-by-descent (IBD) and also typed each sample for known or candidate drug resistance mutations. A total of 160 multi-isolate clones (mean IBD ≥ 0.90) were circulating in Guyana during the study period, comprising 13 highly related clusters (mean IBD ≥ 0.40). In the five-year study period, we observed a decrease in frequency of a mutation associated with artemisinin partner drug (piperaquine) resistance (pfcrt C350R) and limited co-occurence of pfcrt C350R with duplications of plasmepsin 2/3, an epistatic interaction associated with piperaquine resistance. We additionally observed 61 nonsynonymous substitutions that increased markedly in frequency over the study period as well as a novel pfk13 mutation (G718S). However, P. falciparum clonal dynamics in Guyana appear to be largely driven by stochastic factors, in contrast to other geographic regions, given that clones carrying drug resistance polymorphisms do not demonstrate enhanced persistence or higher abundance than clones carrying polymorphisms of comparable frequency that are unrelated to resistance. The use of multiple artemisinin combination therapies in Guyana may have contributed to the disappearance of the pfk13 C580Y mutation.
AbstractThe malaria parasites Plasmodium falciparum and Plasmodium vivax differ in key biological processes and associated clinical effects, but consequences on population-level transmission dynamics are difficult to predict. This co-endemic malaria study from Guyana details important epidemiological contrasts between the species by coupling population genomics (1396 spatiotemporally matched parasite genomes, primarily from 2020–21) with sociodemographic analysis (nationwide patient census from 2019). We describe how P. falciparum forms large, interrelated subpopulations that sporadically expand but generally exhibit restrained dispersal, whereby spatial distance and patient travel statistics predict parasite identity-by-descent (IBD). Case bias towards working-age adults is also strongly pronounced. P. vivax exhibits 46% higher average nucleotide diversity (π) and 6.5x lower average IBD. It occupies a wider geographic range, without evidence for outbreak-like expansions, only microgeographic patterns of isolation-by-distance, and weaker case bias towards adults. Possible latency-relapse effects also manifest in various analyses. For example, 11.0% of patients diagnosed with P. vivax in Greater Georgetown report no recent travel to endemic zones, and P. vivax clones recur in 11 of 46 patients incidentally sampled twice during the study. Polyclonality rate is also 2.1x higher than in P. falciparum, does not trend positively with estimated incidence, and correlates uniquely to selected demographics. We discuss possible underlying mechanisms and implications for malaria control.
Genomic epidemiology has guided research and policy for various viral pathogens and there has been a parallel effort towards using genomic epidemiology to combat diseases that are caused by eukaryotic pathogens, such as the malaria parasite. However, the central concept of viral genomic epidemiology, namely that of measurably mutating pathogens, does not apply easily to sexually recombining parasites. Here we introduce the related but different concept of measurably recombining malaria parasites to promote convergence around a unifying theoretical framework for malaria genomic epidemiology. Akin to viral phylodynamics, we anticipate that an inferential framework developed around recombination will help guide practical research and thus realize the full public health potential of genomic epidemiology for malaria parasites and other sexually recombining pathogens.
During the COVID-19 pandemic, the use of mobile phone data for monitoring human mobility patterns has become increasingly common, both to study the impact of travel restrictions on population movement and epidemiological modeling. Despite the importance of these data, the use of location information to guide public policy can raise issues of privacy and ethical use. Studies have shown that simple aggregation does not protect the privacy of an individual, and there are no universal standards for aggregation that guarantee anonymity. Newer methods, such as differential privacy, can provide statistically verifiable protection against identifiability but have been largely untested as inputs for compartment models used in infectious disease epidemiology. Our study examines the application of differential privacy as an anonymisation tool in epidemiological models, studying the impact of adding quantifiable statistical noise to mobile phone-based location data on the bias of ten common epidemiological metrics. We find that many epidemiological metrics are preserved and remain close to their non-private values when the true noise state is less than 20, in a count transition matrix, which corresponds to a privacy-less parameter ϵ = 0.05 per release. We show that differential privacy offers a robust approach to preserving individual privacy in mobility data while providing useful population-level insights for public health. Importantly, we have built a modular software pipeline to facilitate the replication and expansion of our framework.
Objectives Convenience sampling is an imperfect but important tool for seroprevalence studies. For COVID-19, local geographic variation in cases or vaccination can confound studies that rely on the geographically skewed recruitment inherent to convenience sampling. The objectives of this study were: (1) quantifying how geographically skewed recruitment influences SARS-CoV-2 seroprevalence estimates obtained via convenience sampling and (2) developing new methods that employ Global Positioning System (GPS)-derived foot traffic data to measure and minimise bias and uncertainty due to geographically skewed recruitment. Design We used data from a local convenience-sampled seroprevalence study to map the geographic distribution of study participants’ reported home locations and compared this to the geographic distribution of reported COVID-19 cases across the study catchment area. Using a numerical simulation, we quantified bias and uncertainty in SARS-CoV-2 seroprevalence estimates obtained using different geographically skewed recruitment scenarios. We employed GPS-derived foot traffic data to estimate the geographic distribution of participants for different recruitment locations and used this data to identify recruitment locations that minimise bias and uncertainty in resulting seroprevalence estimates. Results The geographic distribution of participants in convenience-sampled seroprevalence surveys can be strongly skewed towards individuals living near the study recruitment location. Uncertainty in seroprevalence estimates increased when neighbourhoods with higher disease burden or larger populations were undersampled. Failure to account for undersampling or oversampling across neighbourhoods also resulted in biased seroprevalence estimates. GPS-derived foot traffic data correlated with the geographic distribution of serosurveillance study participants. Conclusions Local geographic variation in seropositivity is an important concern in SARS-CoV-2 serosurveillance studies that rely on geographically skewed recruitment strategies. Using GPS-derived foot traffic data to select recruitment sites and recording participants’ home locations can improve study design and interpretation.
Extreme weather events including wildfires and hurricanes are becoming increasingly hazardous due to climate change, and often result in transient or permanent population displacements. Disaster-related disruptions in infrastructure, workforce, wages, and social networks can combine with population displacements to result in interruptions in health care access and prolonged impacts on morbidity and mortality. The data needed to make health systems and emergency management approaches more resilient to these hazards, and more responsive to the needs of affected populations, are sequestered in silos across private corporations and public agencies. In two case studies, we describe how our research team at CrisisReady negotiated access to privately held and novel data sources like anonymized geolocation data from cell-phones, while striking a balance between data security and public health utility. We describe how our analytic tools are embedded into disaster response workflows by co-developing our research questions and outputs with responders and policy-makers. ReadyMapper, an interactive data visualization tool to track population mobility, infrastructure damage, and health system capacity, in near real-time, was deployed during wildfires in California and during the Hurricane Ida response in Louisiana. The Data-Methods-Translational framework we have developed is scalable and relies on sharing science and co-creating products with policy makers and response agencies to ensure real-world applicability. These attributes make the framework particularly useful for formulating evidence-based approaches to protect human health through climate change adaptation.
We thank Dr Tibayrenc for his thoughtful response to our article on the population genomics of malaria parasites, and his enthusiasm for a general framework to understand its idiosyncratic structure [ 1. Tibayrenc M. Towards a general, worldwide, Plasmodium population genomics framework. Trends Parasitol. 2023; 39: 229-230 Abstract Full Text Full Text PDF PubMed Scopus (1) Google Scholar , 2. Camponovo F. et al. Measurably recombining malaria parasites. Trends Parasitol. 2023; 39: 17-25 Abstract Full Text Full Text PDF PubMed Scopus (3) Google Scholar ]. Measurably recombining malaria parasitesCamponovo et al.Trends in ParasitologyNovember 23, 2022In BriefGenomic epidemiology has guided research and policy for various viral pathogens and there has been a parallel effort towards using genomic epidemiology to combat diseases that are caused by eukaryotic pathogens, such as the malaria parasite. However, the central concept of viral genomic epidemiology, namely that of measurably mutating pathogens, does not apply easily to sexually recombining parasites. Here we introduce the related but different concept of measurably recombining malaria parasites to promote convergence around a unifying theoretical framework for malaria genomic epidemiology. Full-Text PDF Open AccessTowards a general, worldwide, Plasmodium population genomics frameworkMichel TibayrencTrends in ParasitologyJanuary 25, 2023In BriefI read with great interest the article by Camponovo et al. [1]. The authors are right in insisting on the necessity of an extensive approach to Plasmodium population genomic diversity. Such a broad picture is sorely needed. This important article inspired me to make the following remarks. Full-Text PDF
Tuberculosis (TB) remains a leading infectious cause of death worldwide. Reducing TB infections and TB-related deaths rests ultimately on stopping forward transmission from infectious to susceptible individuals. Critical to this effort is understanding how human host mobility shapes the transmission and dispersal of new or existing strains of Mycobacterium tuberculosis (Mtb). Important questions remain unanswered. What kinds of mobility, over what temporal and spatial scales, facilitate TB transmission? How do human mobility patterns influence the dispersal of novel Mtb strains, including emergent drug-resistant strains? This review summarizes the current state of knowledge on mobility and TB epidemic dynamics, using examples from three topic areas, including inference of genetic and spatial clustering of infections, delineating source-sink dynamics, and mapping the dispersal of novel TB strains, to examine scientific questions and methodological issues within this topic. We also review new data sources for measuring human mobility, including mobile phone-associated movement data, and discuss important limitations on their use in TB epidemiology.
The human malaria parasite Plasmodium falciparum is globally widespread, but its prevalence varies significantly between and even within countries. Most population genetic studies in P. falciparum focus on regions of high transmission where parasite populations are large and genetically diverse, such as sub-Saharan Africa. Understanding population dynamics in low transmission settings, however, is of particular importance as these are often where drug resistance first evolves. Here, we use the Pacific Coast of Colombia and Ecuador as a model for understanding the population structure and evolution of Plasmodium parasites in small populations harboring less genetic diversity. The combination of low transmission and a high proportion of monoclonal infections means there are few outcrossing events and clonal lineages persist for long periods of time. Yet despite this, the population is evolutionarily labile and has successfully adapted to changes in drug regime. Using newly sequenced whole genomes, we measure relatedness between 166 parasites, calculated as identity by descent (IBD), and find 17 distinct but highly related clonal lineages, six of which have persisted in the region for at least a decade. This inbred population structure is captured in more detail with IBD than with other common population structure analyses like PCA, ADMIXTURE, and distance-based trees. We additionally use patterns of intra-chromosomal IBD and an analysis of haplotypic variation to explore past selection events in the region. Two genes associated with chloroquine resistance, crt and aat1, show evidence of hard selective sweeps, while selection appears soft and/or incomplete at three other key resistance loci (dhps, mdr1, and dhfr). Overall, this work highlights the strength of IBD analyses for studying parasite population structure and resistance evolution in regions of low transmission, and emphasizes that drug resistance can evolve and spread in small populations, as will occur in any region nearing malaria elimination.
Official COVID-19 mortality statistics are strongly influenced by local diagnostic capacity, strength of the healthcare and vital registration systems, and death certification criteria and capacity, often resulting in significant undercounting of COVID-19 attributable deaths. Excess mortality, which is defined as the increase in observed death counts compared to a baseline expectation, provides an alternate measure of the mortality shock—both direct and indirect—of the COVID-19 pandemic. Here, we use data from civil death registers from a convenience sample of 90 (of 162) municipalities across the state of Gujarat, India, to estimate the impact of the COVID-19 pandemic on all-cause mortality. Using a model fit to weekly data from January 2019 to February 2020, we estimated excess mortality over the course of the pandemic from March 2020 to April 2021. During this period, the official government data reported 10,098 deaths attributable to COVID-19 for the entire state of Gujarat. We estimated 21,300 [95% CI: 20, 700, 22, 000] excess deaths across these 90 municipalities in this period, representing a 44% [95% CI: 43%, 45%] increase over the expected baseline. The sharpest increase in deaths in our sample was observed in late April 2021, with an estimated 678% [95% CI: 649%, 707%] increase in mortality from expected counts. The 40 to 65 age group experienced the highest increase in mortality relative to the other age groups. We found substantial increases in mortality for males and females. Our excess mortality estimate for these 90 municipalities, representing approximately at least 8% of the population, based on the 2011 census, exceeds the official COVID-19 death count for the entire state of Gujarat, even before the delta wave of the pandemic in India peaked in May 2021. Prior studies have concluded that true pandemic-related mortality in India greatly exceeds official counts. This study, using data directly from the first point of official death registration data recording, provides incontrovertible evidence of the high excess mortality in Gujarat from March 2020 to April 2021.
The dengue virus affects millions of people every year worldwide, causing large epidemic outbreaks that disrupt people's lives and severely strain healthcare systems. In the absence of a reliable vaccine against dengue or an effective treatment to manage the illness in humans, most efforts to combat dengue infections have focused on preventing its vectors, mainly the Aedes aegypti mosquito, from flourishing across the world. These mosquito-control strategies need reliable disease activity surveillance systems to be deployed. Despite significant efforts to estimate dengue incidence using a variety of data sources and methods, little work has been done to understand the relative contribution of the different data sources to improved prediction. Additionally, scholarship on the topic had initially focused on prediction systems at the national- and state-levels, and much remains to be done at the finer spatial resolutions at which health policy interventions often occur. We develop a methodological framework to assess and compare dengue incidence estimates at the city level, and evaluate the performance of a collection of models on 20 different cities in Brazil. The data sources we use towards this end are weekly incidence counts from prior years (seasonal autoregressive terms), weekly-aggregated weather variables, and real-time internet search data. We find that both random forest-based models and LASSO regression-based models effectively leverage these multiple data sources to produce accurate predictions, and that while the performance between them is comparable on average, the former method produces fewer extreme outliers, and can thus be considered more robust. For real-time predictions that assume long delays (6-8 weeks) in the availability of epidemiological data, we find that real-time internet search data are the strongest predictors of dengue incidence, whereas for predictions that assume short delays (1-3 weeks), in which the error rate is halved (as measured by relative RMSE), short-term and seasonal autocorrelation are the dominant predictors. Despite the difficulties inherent to city-level prediction, our framework achieves meaningful and actionable estimates across cities with different demographic, geographic and epidemic characteristics.