Individual-based models (IBMs) provide a mechanistic framework in which population-level outcomes emerge from interactions between individuals. We conducted a systematic review on IBMs for respiratory pathogens published in 2020-2024. We identified 855 eligible studies. Publications peaked in 2021, with a geographical distribution positively correlated with national GDP, leaving regions understudied. Most studies focused on SARS-CoV-2 and assessed public health interventions. Research priorities evolved over time, shifting from social distancing to vaccination. Age was included in 72.4% of studies; other sociodemographic factors (e.g., race/ethnicity) were rarely considered. This review maps the IBM landscape, offering a framework to guide future modeling efforts.
The COVID-19 pandemic highlighted the importance of human behavior in mitigating the spread of disease. Nonetheless, human behavior is often overlooked in models of disease spread, particularly by underutilizing real-world data. We address this by estimating probabilities that individuals engage in behaviors that influence SARS-CoV-2 transmission risk during the COVID-19 pandemic, between September 2020 and June 2022. These behaviors include wearing a mask, using public transportation, spending time with others, avoiding contact with others, and going to work. Our estimates account for the age and sex of individuals and are generated for every county in the United States. We utilized multiple open-source datasets and United States Census data to produce these estimates. Multiple datasets were used for validation, showing our estimates demonstrated comparable accuracy and robustness. Our estimates aid in understanding human behavior dynamics during the COVID-19 pandemic and could be used to inform monthly or longer-term behavior in simulations of COVID-19. Moreover, the methods presented can be applied to other behaviors and features for future simulations of infectious disease.
Accurate forecasting of infectious diseases drives modern public health interventions that reduce morbidity and mortality. However, accurate forecasting in real-time remains a challenge for the modeling community. Ensembling has emerged as a critical tool for accurate forecasting by leveraging multiple component (individual) models into a single weighted average. Traditional ensembling strategies have relied on bespoke component models that weight the contributions of individual models according to extensive historical data for specific diseases. This is impractical for an emerging disease, since there would be very little - if any - data. We propose an ensembling strategy, called epiFFORMA, that determines component weights for an ensemble model without historical data and is therefore disease-agnostic. The epiFFORMA model builds upon the FFORMA model from the M4 forecasting competition to harness epidemiological dynamics through synthetic data. We demonstrate that epiFFORMA performs better than a naive, equal-weighting ensembling strategy when forecasting outbreaks of COVID-19, diphtheria, influenza-like illness, dengue, measles, mumps, polio, rubella, smallpox, and chikungunya. We further show that epiFFORMA, on average, performs better than the individual component models in the ensemble.
Improving infectious disease models plays a crucial role in mitigating global health impacts. Population-level models have been essential for understanding and anticipating impacts of infectious diseases, but their accuracy has often been limited due to changing environments, rapid pathogen evolution, and insufficient access to detailed transmission data. In particular, due to the lack of granular transmission data, many epidemiological models assume homogeneous disease transmission rather than acknowledge differences between groups within the host population. To address this gap, we conducted a simulation study to assess the potential for inferring group-structured viral transmission dynamics from phylogenetic data and to demonstrate how age structure can be incorporated in case projections. Using a synthetic dataset of 800 age-structured viral phylogenies and their associated case time series, we estimated transmission rates within and between age groups. Our results show that the estimates derived from phylogenies are more accurate than those from case counts alone, although both approaches enabled recovery of age-specific transmission differences. These results suggest that viral phylogenies and case-count time series each provide valuable data that can improve our understanding of population-level disease dynamics, ultimately contributing to more accurate and effective models for informing public policy and reducing disease burden. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Research presented in this article was supported by the Laboratory Directed Research and Development program of Los Alamos National Laboratory under project number 20240066DR. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data and workflows for this study are generated from code publicly available in the repository at https://github.com/lanl/precog/tree/main/aim.
The COVID-19 pandemic highlighted significant differences in infectious disease burden among sociodemographic groups in the United States, underscoring the need for modelling approaches that can capture the complex dynamics driving these heterogeneities. Specifically, variation in case incidence, mortality and disease burden has been observed across subpopulations stratified by race, ethnicity, sex, age and geographic region. Accurately incorporating fine-grained sociodemographic attributes into infectious disease models remains challenging due to complex correlations among individual characteristics. Additionally, accurately modelling transmission while accounting for exposure differences among population strata requires a detailed understanding of transmission risk across interaction settings. We address these challenges by incorporating drivers of exposure risk and detailed sociodemographic data into EpiCast-a large-scale agent-based model of respiratory pathogen spread in the United States. Using this model, we demonstrate how differences in the rate of infections between key demographic groups emerge in households, workplaces and schools. Our findings show that embedding fine-grained population heterogeneity into infectious disease models can reveal uneven outcomes in predicted disease burden among racial groups, driven by factors such as household size and workplace exposure risk. This study demonstrates the potential of detailed models of infectious disease spread to inform policy intervention design for future pandemics.
Human cognitive responses, behavioral responses, and disease dynamics co-evolve over the course of any disease outbreak, and can result in complex feedbacks. We present a dynamic agent-based model that explicitly couples the spread of disease with the spread of fear surrounding the disease, implemented within the EpiCast simulation framework. EpiCast models transmission within a realistic synthetic population, capturing individual-level interactions. In our model, fear propagates through both in-person contact and broadcast media, prompting individuals to adopt protective behaviors that reduce disease spread. In order to better understand these coupled dynamics, we create and compare a range of compartmental models to ensure that introducing additional disease states does not prevent the emergence of multiple waves in these simpler models. Additionally, we compare a range of behavioral scenarios within EpiCast, varying the level and intensity of fear and behavior change. Our results show that the addition of asymptomatic, exposed, and pre-symptomatic disease states can impact both the rate at which an outbreak progresses and its overall trajectory in compartmental models. In EpiCast, the combination of non-local fear spread via broadcasters and strong behavioral responses by fearful individuals generally leads to multiple epidemic waves, an outcome that occurs only within a narrow parameter range when fear spreads purely through local contact. Accounting for the coupled spread of fear and disease is critical for understanding disease dynamics and designing timely, targeted responses to emerging infectious threats.
Mathematical models of infectious diseases are frequently used as a tool to support public health policy and decisions around the implementation of interventions such as school closures. However, most publications on policy-relevant modelling lack an ethical framework and do not explicitly consider the ethical implications of the work. This creates a risk that the unintended consequences of interventions are overlooked or that models are used to justify decisions that are inconsistent with public health ethics. In this article, we focus on the case study of school closures as a commonly modelled intervention against pandemic influenza, COVID-19 and other infectious disease threats. We briefly review some of the key concepts in public health ethics and describe approaches to modelling the effects of school closures. We then identify a series of ethical considerations involved in modelling school closures. These include accounting for population heterogeneity and inequalities; including a diversity of viewpoints and expertise in model design; considering the distribution of benefits and harms; and model transparency and contextualization. We conclude with some recommendations to ensure that policy-relevant modelling is consistent with some key ethics values.
Infectious disease modeling and forecasting have played a key role in helping assess and respond to epidemics and pandemics. Recent work has leveraged data on disease peak infection and peak hospital incidence to fit compartmental models for the purpose of forecasting and describing the dynamics of a disease outbreak. Incorporating these data can greatly stabilize a compartmental model fit on early observations, where slight perturbations in the data may lead to model fits that project wildly unrealistic peak infection. We introduce a new method for incorporating historic data on the value and time of peak incidence of hospitalization into the fit for a Susceptible-Infectious-Recovered (SIR) model by formulating the relationship between an SIR model's starting parameters and peak incidence as a system of two equations that can be solved computationally. This approach is assessed for practicality in terms of accuracy and speed of computation via simulation. To exhibit the modeling potential, we update the Dirichlet-Beta State Space modeling framework to use hospital incidence data, as this framework was previously formulated to incorporate only data on total infections.
The devastating global impacts of the COVID-19 pandemic are a stark reminder of the need for proactive and effective pandemic response. Disease modeling and forecasting are key in this response, as they enable forward-looking assessment and strategic planning. Via 85 interviews spanning 14 countries with disease modelers and those they support, conducted amid the COVID-19 pandemic response, we offer a qualitative overview of challenges faced, lessons learned, and readiness for future pandemics. The interviewees highlighted several key challenges and considerations in forecasting, particularly emphasizing the complications introduced by human behavior and various data-related issues (including data availability, quality, and standardization). They underscored the importance of effective communication among those who create models, those who make decisions based on these models, and the general public. Additionally, they pointed out the necessity for addressing global equity, debated the merits of centralized versus decentralized responses to crises, and stressed the need for establishing measures for sustainable preparedness. Their verdicts on future pandemic readiness were mixed, with only 43
The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.
The recent history of respiratory pathogen epidemics, including those caused by influenza and SARS-CoV-2, has highlighted the urgent need for advanced modeling approaches that can accurately capture heterogeneous disease dynamics and outcomes at the national scale, thereby enhancing the effectiveness of resource allocation and decision-making. In this paper, we describe Epicast 2.0, an agent-based model that utilizes a highly detailed, synthetic population and high-performance computing techniques to simulate respiratory pathogen transmission across the entire United States. This model replicates the contact patterns of over 320 million agents as they engage in daily activities at school, work, and within their communities. Epicast 2.0 supports vaccination and an array of non-pharmaceutical interventions that can be promoted or relaxed via highly granular, user specified policies. We illustrate the model's capabilities using a wide range of outbreak scenarios, highlighting the model's varied dynamics as well as its extensive support for policy exploration. This model provides a robust platform for conducting what if scenario analysis and providing insights into potential strategies for mitigating the impacts of infectious diseases.
BACKGROUND:Seasonal influenza infects 5-20% of people every year in the United States, resulting in hospitalizations, deaths, and adverse economic impacts. To mitigate these impacts, influenza vaccines are developed and distributed annually; however, growing evidence suggests that vaccine effectiveness (VE) wanes over the course of a flu season. Delaying influenza vaccination for older adults has attracted attention as a potential public health strategy. However, given the uncertainties in seasonal peak, vaccine effectiveness, and waning rates, postponing vaccination could also lead to increased morbidity, motivating an evaluation of a range of potential scenarios. The aim of this study was to investigate favorable age group-specific vaccination schedules that could lead to the greatest disease burden reduction. METHODS:We systematically investigated a broad range of vaccination start times for five age groups under six combinations of initial effectiveness and waning rates, based on influenza cases and vaccine uptake data from 10 influenza seasons. We defined the most favorable vaccination schedule as the one that resulted in the greatest reduction in disease burden. RESULTS:In scenarios with fast waning, all age groups benefit from delaying vaccination regardless of initial VE and peak timing. In scenarios with slower waning, results are mixed. For the ≥65 group, high initial VE and slow waning suggests that in early-peaking seasons, early vaccination most effectively reduces disease burden, while in late-peaking seasons delaying vaccination is most effective. For the ≥65 group in medium and low initial VE, and slow waning scenarios, delaying vaccination appears to prevent the greatest number of cases, regardless of whether the season peaks early or late. CONCLUSION:The most favorable vaccination schedule is sensitive to changes in initial VE, waning rate, and peak timing. Given estimates of these quantities from statistical and immunological models and observations, our methods can inform vaccination recommendations in order to most effectively reduce the annual disease burden caused by seasonal influenza. Specifically, accurate peak timing forecasts for the upcoming season have the potential to guide decisions on when to vaccinate.
Disease surveillance systems allow public health agencies to respond to emerging diseases before they become widespread. Developing such systems requires identifying optimal ways to monitor in the context of an epidemic outbreak; this problem is known as sensor selection. Contact networks represent the dynamics of interaction in a population and are used to model how a disease spreads in a population and to explore strategies of sensor selection. We evaluated five sensor selection strategies on their ability to provide an early warning of a COVID-like outbreak in synthetic contact networks encapsulated in four network scenarios. Three of these scenarios assessed different aspects of community structure. The fourth scenario employed a contact network representing the population and interactions of 6.8 million people in New York City, constructed from an agent-based simulation using census and transportation data. This scenario exemplifies how sensor selection strategies may perform in a real-world, urban context. Our findings suggest that the choice of the optimal strategy depends heavily on the community structure of the network. Strategies that select highly connected nodes or maximize network coverage are the optimal surveillance strategy for outbreak detection in many network community structures. However, a naive implementation of these strategies may fail to provide an early warning at all—including in the New York City scenario. Moreover, these methods are impractical for real-world use as they require knowledge of the underlying contact network. Instead, a selection strategy that starts with a set of random nodes and then performs a random walk through a chain of neighbors reliably provides early warnings without requiring prior knowledge of the network. We find this method, called "random chain", to be the most pragmatic for implementation in a real-world disease surveillance context.
Background Nonpharmaceutical interventions (NPIs) may be considered as part of national pandemic preparedness as a first line defense against influenza pandemics. Preemptive school closures (PSCs) are an NPI reserved for severe pandemics and are highly effective in slowing influenza spread but have unintended consequences. Methods We used results of simulated PSC impacts for a 1957-like pandemic (i.e., an influenza pandemic with a high case fatality rate) to estimate population health impacts and quantify PSC costs at the national level using three geographical scales, four closure durations, and three dismissal decision criteria (i.e., the number of cases detected to trigger closures). At the Chicago regional level, we also used results from simulated 1957-like, 1968-like, and 2009-like pandemics. Our net estimated economic impacts resulted from educational productivity costs plus loss of income associated with providing childcare during closures after netting out productivity gains from averted influenza illness based on the number of cases and deaths for each mitigation strategy. Results For the 1957-like, national-level model, estimated net PSC costs and averted cases ranged from $7.5 billion (2016 USD) averting 14.5 million cases for two-week, community-level closures to $97 billion averting 47 million cases for 12-week, county-level closures. We found that 2-week school-by-school PSCs had the lowest cost per discounted life-year gained compared to county-wide or school district–wide closures for both the national and Chicago regional-level analyses of all pandemics. The feasibility of spatiotemporally precise triggering is questionable for most locales. Theoretically, this would be an attractive early option to allow more time to assess transmissibility and severity of a novel influenza virus. However, we also found that county-wide PSCs of longer durations (8 to 12 weeks) could avert the most cases (31–47 million) and deaths (105,000–156,000); however, the net cost would be considerably greater ($88-$103 billion net of averted illness costs) for the national-level, 1957-like analysis. Conclusions We found that the net costs per death averted ($180,000-$4.2 million) for the national-level, 1957-like scenarios were generally less than the range of values recommended for regulatory impact analyses ($4.6 to 15.0 million). This suggests that the economic benefits of national-level PSC strategies could exceed the costs of these interventions during future pandemics with highly transmissible strains with high case fatality rates. In contrast, the PSC outcomes for regional models of the 1968-like and 2009-like pandemics were less likely to be cost effective; more targeted and shorter duration closures would be recommended for these pandemics.
Background Vector-borne diseases (VBDs) are important contributors to the global burden of infectious diseases due to their epidemic potential, which can result in significant population and economic impacts. Oropouche fever, caused by Oropouche virus (OROV), is an understudied zoonotic VBD febrile illness reported in Central and South America. The epidemic potential and areas of likely OROV spread remain unexplored, limiting capacities to improve epidemiological surveillance. Methods To better understand the capacity for spread of OROV, we developed spatial epidemiology models using human outbreaks as OROV transmission-locality data, coupled with high-resolution satellite-derived vegetation phenology. Data were integrated using hypervolume modeling to infer likely areas of OROV transmission and emergence across the Americas. Results Models based on one-support vector machine hypervolumes consistently predicted risk areas for OROV transmission across the tropics of Latin America despite the inclusion of different parameters such as different study areas and environmental predictors. Models estimate that up to 5 million people are at risk of exposure to OROV. Nevertheless, the limited epidemiological data available generates uncertainty in projections. For example, some outbreaks have occurred under climatic conditions outside those where most transmission events occur. The distribution models also revealed that landscape variation, expressed as vegetation loss, is linked to OROV outbreaks. Conclusions Hotspots of OROV transmission risk were detected along the tropics of South America. Vegetation loss might be a driver of Oropouche fever emergence. Modeling based on hypervolumes in spatial epidemiology might be considered an exploratory tool for analyzing data-limited emerging infectious diseases for which little understanding exists on their sylvatic cycles. OROV transmission risk maps can be used to improve surveillance, investigate OROV ecology and epidemiology, and inform early detection.
Infectious disease outbreaks pose a major threat to public health, economic stability, and human security. In order to mitigate disease impacts, we need to understand drivers that contribute to its spread. Latinx, Black, and American Indian racial and ethnic groups experienced disproportionate health outcomes during the COVID-19 pandemic. Current modeling approaches often fail to capture the high level of heterogeneity present in the population needed to measure the drivers of these disparities.