To better respond to threats, decision-makers are increasingly interested in predictions they can understand and trust. Collaborative modeling can help increase the relevance, transparency, and robustness of predictions. This approach can be facilitated with hubs, or centralized data repositories to collect, analyze, and communicate model output. This paper introduces the hubverse, a suite of standards and software tools to streamline the creation and operation of collaborative modeling hubs. Hubverse file structure and model output standards enable the use of common tools to validate, aggregate, visualize, evaluate, and communicate model output. Currently, the hubverse is used by nearly two dozen collaborative and local modeling hubs around the globe to support infectious disease modeling efforts, including hubs hosted and/or used by the United States Centers for Disease Control and Prevention, European Centre for Disease Prevention and Control, Australia-Aotearoa Consortium for Epidemic Forecasting and Analytics, and California Department of Public Health.
Background Up-to-date real-time disease surveillance data can provide critical public health insights, however reporting delays can create downward bias in the latest data. Nowcasting methods designed to correct for this bias remain underused in public health practice due to their complexity, lack of tailored documentation, or technical barriers. Methodological advances in nowcasting are also hampered by the absence of standardised benchmarks for evaluating new methods. Methods To address these needs, we developed a family of nowcasting methods and an accompanying R package, baselinenowcast. We validated our method against the baseline method that was used in the German COVID-19 Nowcast Hub and on which our approach was based. Using this data, we conducted an analysis to compare different specifications of our method which were designed to address common issues in epidemiology such as weekday patterns in reporting and the ability to share estimates across different strata. We used our approach on norovirus surveillance data from the United Kingdom Health Security Agency (UKHSA) and compared the performance of three of our method specifications against three methods evaluated in a previous study. Results Our baseline method improved estimates compared to unadjusted data across all case studies. We found that the optimal choice of baseline method specification depends on context but that our default method specification performed well in a range of settings. Applied to UKHSA norovirus data, our method helped us understand the performance of the model currently used in public health practice. Conclusions Our method and software can be used both as a straightforward nowcasting method and provides a benchmark for nowcasting model development.
The ability to estimate and predict pathogen variant dynamics can inform public health responses, including planning for increased transmission or severity, shifts in population immunity, or changes to vaccine or therapeutic effectiveness. The COVID-19 pandemic demonstrated the importance of monitoring SARS-CoV-2 variant evolution through viral genome sequencing, enabling predictive models to estimate variant frequencies in the recent past, present, and short-term future. Collaborative forecasting Hubs provided a valuable way to centralize predictive modeling of epidemiological indicators such as cases, hospitalizations, and deaths during the pandemic; however, none existed for variant dynamics. Here, we discuss the creation of the United States SARS-CoV-2 Variant Nowcast Hub, designed to solicit estimates of the relative abundance of a specified set of SARS-CoV-2 variants at the U.S. state level. We discuss the design decisions and challenges in building the Hub and its scoring procedures. Using submissions from the Hub's first respiratory virus season (nowcast dates October 9th, 2024 to June 4th, 2025), we evaluate five individual models and a baseline model. We found that the baseline model, which pools sequences across the U.S., performs well overall, with most individual models performing similarly or slightly worse. Locations with lower sequencing volumes exhibited greater variability in model performance. Models submitted for a single location outperformed those submitted for all locations, potentially due to greater timeliness and magnitude of local data. Much remains to be investigated regarding relative model performance across different phases of variant emergence, and we conclude by proposing future directions within and beyond this Hub.
ABSTRACT Infectious disease modelling has become an increasingly prominent tool in public health decision-making, with its use accelerating markedly during the COVID-19 pandemic. This growth calls for an understanding of how modelling evidence is received, interpreted, and used – or not used – by decision-makers. There is a recognised need for evaluations of modelling-to-policy systems to understand how to best integrate modelling evidence into decision-making. So far, no comprehensive evaluation framework exists that maps modelling-to-policy pathways. This study addresses that gap by developing a theory of change for modelling-to-policy systems that could serve as a foundation for future evaluations. A qualitative study design was employed, comprising semi-structured interviews with 35 modellers, knowledge brokers, and decision-makers across five continents and diverse institutional settings, spanning high-income and low- and middle-income countries, as well as national and international modelling-to-policy contexts. Thematic analysis was combined with a backward-mapping-informed approach to develop a multi-level evaluative framework. The protocol for this study has been published on March 20, 2025, at OSF ( https://doi.org/10.17605/OSF.IO/J9QXV ). Participants diverged in their conceptualisations of successful modelling evidence use – ranging from instrumental use to accurate understanding and consideration of modelling outputs – yet converged on shared risks: decisions informed by inadequately specified models or by evidence that is misinterpreted due to communication failures. The resulting three-level framework identifies factors directly influencing modelling evidence use across three domains (policy relevance, model quality, and communication and interaction), traces these to enabling conditions, and maps them to systemic enablers – including local and embedded modelling capacity, data infrastructure, knowledge brokering capacity, formal knowledge translation structures, established networks, and funding. The proposed framework represents a first theory of change for modelling-to-policy systems. While it requires further testing and application across diverse decision-making contexts, it offers a structured basis for evaluating existing systems and informing the design of new ones.
The COVID-19 crisis required scientists worldwide to contribute to complex, hectic, and unfamiliar governmental decision-processes. In this Perspective, we reflect on this intense interaction between science and health policy for pandemic response, drawing from the experience of infectious disease modellers across the world. We highlight the diversity of actors and interests in government, aiming to demystify the elusive ‘policy makers’. We present a general taxonomy to help research scientists more effectively support evidence-based policy. We stress the importance of building and maintaining relationships appropriate to the diverse pool of government actors. Not recognising this diversity may lead to miscommunication and reduce the positive impacts of scientific evidence for public policy in crisis and non-crisis situations.
Background Up-to-date real-time disease surveillance data can provide critical public health insights, however reporting delays can create downward bias in the latest data. Nowcasting methods designed to correct for this bias remain underused in public health practice due to their complexity, lack of tailored documentation, or technical barriers. Methodological advances in nowcasting are also hampered by the absence of standardised benchmarks for evaluating new methods. Methods To address these needs, we developed a family of nowcasting methods and an accompanying R package, baselinenowcast . We validated our method against the baseline method that was used in the German COVID-19 Nowcast Hub and on which our approach was based. Using this data, we conducted an analysis to compare different specifications of our method which were designed to address common issues in epidemiology such as weekday patterns in reporting and the ability to share estimates across different strata. We used our approach on norovirus surveillance data from the United Kingdom Health Security Agency (UKHSA) and compared the performance of three of our method specifications against three methods evaluated in a previous study. Results Our baseline method improved estimates compared to unadjusted data across all case studies. We found that the optimal choice of baseline method specification depends on context but that our default method specification performed well in a range of settings. Applied to UKHSA norovirus data, our method helped us understand the performance of the model currently used in public health practice. Conclusions Our method and software can be used both as a straightforward nowcasting method and provides a benchmark for nowcasting model development.
Aygün et al (2026, https://doi.org/10.1038/s41586-026-10658-6) claim that their AI-driven Empirical Research Assistance (ERA) system produces COVID-19 hospitalisation forecasts which outperform the state-of-the-art CDC ensemble by a considerable margin for the 2024/25 season. We demonstrate that the observed performance gain is attributable to information leakage in the retrospective forecasting setup, which resulted because data revisions were not taken into account. As similar mechanisms are at play in many other forecasting fields, our cautionary tale applies not just to epidemic forecasting, but is relevant to the entire emerging field of AI-assisted predictive modelling.
To better respond to a range of threats, decision-makers in diverse fields are increasingly interested in predictions they can understand and trust. Collaborative modeling can help increase the relevance, transparency, and robustness of predictions. This approach can be facilitated with hubs, or centralized data portals to collect, analyze, and communicate model output. This paper introduces the hubverse, an open-source suite of standards and software tools to streamline the creation and operation of collaborative modeling hubs. Hubverse file structure and model output standards enable the use of common tools to validate, aggregate, visualize, evaluate, and communicate model output. Currently, the hubverse is used by nearly two dozen collaborative and local modeling hubs around the globe to support infectious disease modeling efforts, including hubs hosted and/or used by the United States Centers for Disease Control and Prevention, the European Centre for Disease Prevention and Control, the Australia-Aotearoa Consortium for Epidemic Forecasting and Analytics, and the California Department of Public Health.
We reflect on the sustainability of modelling infectious disease outbreaks from the perspective of modelling as a field of practice. We formed a community of practice among UK infectious disease modellers who had contributed to the UK COVID-19 response. We previously used a participatory workshop approach to highlight issues in the infrastructure and incentives for outbreak modelling, and synthesized our experience into a set of 12 specific recommendations. Here, we track changes in the field of infectious disease modelling 1 year later, collecting the quantitative and qualitative views of change among 14 participants. We found participants continued to highlight a lack of ongoing, sufficient or appropriate action to develop outbreak modelling capacity in the UK, while positively noting collaborations among public health facing institutions. We emphasize the under-prioritization of funding for outbreak modelling outside of emergency response periods, and the continuation of unsustainable working practices. Correcting this is crucial to supporting evidence-based public health policy for outbreak preparedness and response.
The ongoing H5N1 panzootic in mammals has amplified zoonotic pathways to facilitate human infection. Characterising key epidemiological parameters for H5N1 is critical should it become widespread. To identify and estimate critical epidemiological parameters for H5N1 from past and current outbreaks, and to compare their characteristics with human influenza subtypes and the 2003 Netherlands H7N7 outbreak. We searched PubMed, Embase, and Cochrane Library for systematic reviews reporting parameter estimates from primary data or meta-analyses. To address gaps, we searched PubMed and Google Scholar for studies of any design providing relevant estimates. We estimated the basic reproduction number for the recent outbreak in the United States (US) and the 2003 Netherlands H7N7 outbreak. In addition, we estimated the serial interval for H5N1 using data from previous household clusters in Indonesia. We also applied a branching process model to simulate transmission chain size and duration to assess if simulated transmission patterns align with observed dynamics. From 46 articles, we identified H5N1’s epidemiological profile as having lower transmissibility (R0 < 0.2) but higher severity compared to other human subtypes. Evidence suggests H5N1 has a longer incubation (∼ 4 days vs. ∼ 2 days) and serial intervals (∼ 6 days vs. ∼ 3 days) than human subtypes, impacting transmission dynamics. The epidemiology of the US H5 outbreak is similar to the 2003 Netherlands H7N7 outbreak. Key gaps remain regarding latent and infectious periods. We characterised critical epidemiological parameters for H5N1 infection. The current US outbreak shows lower pathogenicity, but similar transmissibility compared to prior outbreaks. Longer incubation and serial intervals may enhance contact tracing feasibility. These estimates offer a baseline for monitoring changes in H5N1 epidemiology. Not applicable.
Baseline models are essential reference points for evaluating forecasting methods, yet their selection often receives insufficient attention. We present a systematic framework for baseline model selection in epidemiological forecasting, establish- ing criteria for suitable baselines and demonstrating the consequences of different choices. Analysing data from COVID-19 and influenza forecast hubs, we evalu- ated ten baseline model frameworks. Our results reveal that baseline selection profoundly impacts forecast evaluation: for influenza, the proportion of models outperforming the baseline ranged from 10% to 93% depending on the baseline chosen. No single baseline satisfied all evaluation criteria. The choice of base- line also affected model rankings, with some baselines producing substantially different orderings of forecast model performance. We found that well-calibrated baselines do not necessarily align with good forecast performance, highlighting a fundamental tension in baseline selection. These findings highlight the need for careful baseline selection in forecast evaluation, particularly in collaborative efforts where fair comparison across multiple models is essential. We provide prac- tical recommendations for baseline selection and suggest strategies for improving evaluation fairness when ideal baselines cannot be identified. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by funds from Wellcome (210758/Z/18/Z) and the National Institute for Health and Care Research (NIHR) Health Protection Research Unit (HPRU) in Health Analytics and Modelling. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data is publicly available and summarised at https://doi.org/ 10.5281/zenodo.16407890. Application code is available at https://github.com/ ManuelStapper/Baselines Application.
This study examines the use and translation of epidemiological modelling by policy and decision makers in response to the COVID-19 outbreak. Prior to COVID-19, there was little readiness for global health systems, and many science-policy networks were assembled ad-hoc. Moreover, in the field of epidemiological modelling, one with significant sudden influence, there is still no international guidance or standard of practice on how modelled evidence should guide policy during major health crises. Here we use a multi-country case study on the use of epidemiological modelling in emergency COVID-19 response, to examine the effective integration of crisis science and policy in different countries. We investigated COVID-19 modelling-policy systems and practices in 13 countries, spanning all six UN geographic regions. Data collection took the form of expert interviews with a range of national policy/ decision makers, scientific advisors, and modellers. We examined the current use of epidemiological modelling, introduced a classification framework for outbreak modelling and policy on which best practice can be structured, and provided preliminary recommendations for future practice. Full analysis and interpretation of the breadth of interview responses is presented, providing evidence for the current and future use of modelling in disease outbreaks. We found that interviewees in countries with a similar size and type of modelling infrastructure, and similar level of government interaction with modelling reported similar experiences and recommendations on using modelling in outbreak response. From this, we introduced a helpful grouping of country experience upon which a tailored future best practice could be structured. We concluded the article by outlining context-specific activities that modellers and policy actors could consider implementing in their own countries. This article serves as a first evidence base for the current use of modelling in a recent major health crisis and provides a robust framework for developing epidemiological modelling-to-policy best practice.
Infectious disease modelling plays a critical role in guiding decisions during outbreaks. However, ongoing debates over the utility of these models highlight the need for a deeper understanding of their exact role in decision-making. In this scoping review we sought to fill this gap, focusing on challenges and facilitators of translating modelling insights into actionable policies. We searched the Ovid database to identify modelling studies that included an assessment of utility in informing policy and decision-making from January 2019 onwards. We further identified studies based on expert judgement. Results were analysed descriptively. The study was registered on the Open Science Framework platform. Out of 4007 screened and 12 additionally suggested studies, a total of 33 studies were selected for our review. None of the included articles provided objective assessments of utility but rather reflected subjectively on modelling efforts and highlighted individual key aspects for utility. 27 of the included articles considered the COVID-19 pandemic and 25 of the articles were from high-income countries. Most modelling efforts aimed to forecast outbreaks and evaluate mitigation strategies. Participatory stakeholder engagement and collaboration between academia, policy, and non-governmental organizations were identified as key facilitators of the modelling-for-decisions pathway. However, barriers such as data inconsistencies and quality, uncoordinated decision-making, limited funding and misinterpretation of uncertainties hindered effective use of modelling in decision-making. While our review identifies crucial facilitators and barriers for the modelling-for-decisions pathway, the lack of rigorous assessments of the utility of modelling for decisions highlights the need to systematically evaluate the impact of infectious disease modelling on decisions in future.
Background: Spatial models of infectious disease transmission often rely on administrative boundaries and simple measure of proximity, such as the neighbourhood order. While these methods effectively capture key aspects of disease spread, they may not fully account for nuances in transmission risk. This study introduces a new approach that defines transmission risk at the individual level and aggregates it to the district level. This approach aims to improve both forecast accuracy and the spatial resolution of risk mapping. Method: We propose the Fine–Grid Spatial Interaction Matrix (FGSIM), which models transmission risk between districts based on distances between individuals. Five distance measures are transformed into contact intensity using a power–law function, forming a weight matrix applicable in the endemic–epidemic framework. We evaluate FGSIM using influenza data from Germany (2001–2020) and compare its performance to established methods and a simplified FGSIM variant. Forecasts for one to eight weeks ahead are assessed across four study regions using the weighted interval score (WIS) and ranked probability score (RPS). Results: FGSIM outperforms its simplified variants and models without spatial dependence in most cases when evaluated in-sample. For one–week–ahead forecasts, a centroid–based model performs best in three of four regions (two when evaluated on the logarithmic scale). For longer–term forecasts (four or eight weeks ahead), one FGSIM model consistently outperforms most others. Risk maps at 100m resolution demonstrate the ability of FGSIM to identify high–risk areas not aligned with administrative boundaries. Conclusion: FGSIM provides a flexible and computationally feasible approach to incorporate individual–level spatial structure into district–level infectious disease models. While established methods perform well in short–term forecasts, FGSIM demonstrates competitive performance for longer forecasting horizons. In addition, it enables fine–scale risk mapping, making it a valuable tool for public health planning beyond administrative districts. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study was supported by Wellcome (210758/Z/18/Z) ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
OBJECTIVES:There are few standards for what information about an infectious disease outbreak should be reported to the public and when. To address this problem, we undertook a consensus process to develop recommendations for what epidemiological information public health authorities should report to the public during an outbreak. STUDY DESIGN:We conducted a Delphi study following the steps outlined in the ACcurate COnsensus Reporting Document (ACCORD) for health-related activities or research. METHODS:We assembled a steering committee of nine experts representing federal and state public health, academia, and international partners to develop a candidate list of reporting items. We then invited 45 experts, 35 of whom agreed to participate in a Delphi panel. Of those, 25 participated in voting in the first round, 25 in the second round, and 25 in the third round, demonstrating consistent engagement in the consensus-building process. The final stage of the Delphi process consisted of a hybrid consensus meeting to finalize the voting items. RESULTS:The Delphi process yielded nine core reporting items representing a minimum standard for public outbreak reporting: numbers of new confirmed cases, new hospital admissions, new deaths, cumulative confirmed cases, cumulative hospital admissions, and cumulative deaths, each reported weekly and at Administrative Level 1 (typically state or province), and stratified by sex, age group, and race/ethnicity. CONCLUSIONS:This minimum reporting standard creates a strong framework for uniform sharing of outbreak information and promotes consistency of data between jurisdictions, enabling effective response by promoting access to information about an unfolding epidemic.