Retractions serve as an indicator of failures in research integrity, yet most analyses focus on absolute counts rather than risk per paper. We use one of the largest open bibliographic databases to develop incidence metrics normalized by population: retractions per publication and per active author annually. Applying an epidemiological framework that models counts with exposure, we find evidence of exponential growth in retraction incidence, with approximately a 5-year doubling time at both the paper and author levels. These patterns vary significantly across fields, publishers, and countries. While scientific output is becoming more democratized globally, retractions are concentrated in fewer countries, creating a "concentration" paradox that calls for targeted monitoring. Despite exponential growth, the absolute incidence remains low (0.12
Multilayer network science has emerged as a central framework for analysing interconnected and interdependent complex systems. Its relevance has grown substantially with the increasing availability of rich, heterogeneous data, which makes it possible to uncover and exploit the inherently multilayered organisation of many real-world networks. In this review, we summarise recent developments in the field. On the theoretical and methodological front, we outline core concepts and survey advances in community detection, dynamical processes, temporal networks, higher-order interactions, and machine-learning-based approaches. On the application side, we discuss progress across diverse domains, including interdependent infrastructures, spreading dynamics, computational social science, economic and financial systems, ecological and climate networks, science-of-science studies, network medicine, and network neuroscience. We conclude with a forward-looking perspective, emphasizing the need for standardised datasets and software, deeper integration of temporal and higher-order structures, and a transition toward genuinely predictive models of complex systems.
Abstract Increasing human mobility and population connectivity have intensified the risks of global pathogen spread, while concurrent shifts in human demographic patterns, ecological factors, and climatic conditions have altered the global landscape of this risk. Genomic surveillance can serve as a critical tool for early detection of emerging pathogen threats; however, challenges remain in deciding where to monitor, in understanding trade-offs among surveillance modalities, and in translating detections into actionable estimates of importation and local transmission for public health decision-making. Here we develop a computational framework to evaluate strategies for respiratory pathogen detection that integrates an established clinical surveillance modality, intensive care unit (ICU) sampling, with an emerging environmental modality, aircraft wastewater (AWW) sampling. Detections are translated into risk via a multi-scale, stochastic global transmission model that combines international flight data with a detailed agent-based local transmission model. The resulting model-based estimates contrast the time to pathogen detection via AWW at airports with that in the community via realistic healthcare testing pathways. Using real-world data from England and Wales (EW), we find that employing AWW in EW airports can improve first detection times by 12.5-37.7 days for a range of epidemiological parameters under realistic healthcare testing scenarios and random aircraft sampling between 25 and 50%. In particular, for a SARS-CoV-2-like pathogen, we expect AWW to outperform ICU in first detection timing by 22.0-25.6 days, with ∼21.9-42.6 times fewer cases at their respective time of detection. While false detection remains a risk, we show that follow-up confirmatory testing can improve detection confidence substantially. Together our results demonstrate the potential utility of AWW surveillance and how it can reduce detection times and improve global health security.
Scoring rules are critical for evaluating the predictive performance of epidemic models by quantifying how well their projections and forecasts align with observed data. In this study, we introduce the energy score as a performance metric for stochastic trajectory-based epidemic models. As a multivariate extension of the continuous ranked probability score (CRPS), the energy score provides a single, unified measure for time-series predictions. It evaluates both calibration and sharpness by considering the distances between individual trajectories and observed data, as well as the inter-trajectory variability. We provide an overview of how the energy score can be applied to assess both scenario projections and forecasts in this format, with a particular focus on a detailed analysis of the Scenario Modeling Hub results for the 2023-2024 influenza season. By comparing the energy score to the widely used weighted interval score (WIS), we demonstrate its utility as a tool for evaluating epidemic models, especially in scenarios requiring integration of predictions across multiple target outcomes into a single, interpretable metric.
Infectious disease spread is a multiscale process composed of within-host (biological) and between-host (social) drivers and disentangling them from each other is a central challenge in epidemiology. Here, we introduce VIBES, a multiscale modeling framework that explicitly integrates viral dynamics based on patient-level data with population-level transmission on a data-driven network of social contacts. Using SARS-CoV-2 as a case study, we analyze three emergent epidemic properties, namely the generation time, serial interval, and presymptomatic transmission. First, we established a purely biological baseline, thus independent of the reproduction number (R), from the within-host model, estimating a generation time of 6.3 d for symptomatic individuals and 43.1% presymptomatic transmission. Then, using the full model incorporating social contacts, we found a shorter generation time (5.4 d at R = 3.0) and an increase in presymptomatic transmission (52.8% at R = 3.0), disentangling the impact of social drivers from a purely biological baseline. We further show that as pathogen transmissibility increases (R from 1.3 to 6), competition among infectious individuals shortens the generation time and serial interval by up to 21% and 13%, respectively. Conversely, a social intervention, like isolation, increases the proportion of presymptomatic transmission by about 30%. Our framework also estimates metrics that are challenging to obtain empirically, such as the generation time for asymptomatic individuals (5.6 d; 95%CI: 5.1 to 6.0 at R = 1.3). Our findings establish multiscale modeling as a powerful tool for mechanistically quantifying how pathogen biology and human social behavior shape epidemic dynamics as well as for assessing public health interventions.
In recent years, the use of multi-model ensemble projections in infectious disease modeling has become an established methodological approach to account for and integrate across uncertainties and structural differences present in individual models. However, the creation of long-term ensemble projections through these coordinated efforts is resource-intensive, demanding the input of multiple research teams and substantial computational power. This typically limits the ability to refine projections, update the selection of plausible epidemic trajectories, or expand the number of scenarios that can be assessed, even as new empirical data become available. To address this challenge, we define an adaptive ensemble approach that, analogously to a multi-model particle filtering method, dynamically selects individual model trajectories based on observed data throughout the epidemic projection period. We demonstrate the effectiveness of this methodology using the U.S. Flu Scenario Modeling Hub (SMH) projections for influenza hospitalizations in the United States during the 2023-2024 and 2024-2025 winter seasons. Our findings show that the adaptive ensemble yields improved predictive accuracy with respect to the original SMH ensemble projections across several scoring rules and geographical resolutions. Furthermore, the adaptive ensemble approach offers two additional applications: i) the dynamic assignment of posterior probabilities to epidemic scenarios, identifying the most plausible scenario, and representing how reality is captured by a combination of scenarios, and ii) the potential use for short-term forecasting. The adaptive ensemble approach is able to identify the most likely scenarios for the 2023-2024 and 2024-2025 U.S. influenza seasons, even in the early stages of the epidemic. It outperforms, retrospectively, a baseline model in short-term forecasting of influenza hospitalizations in the United States during the two seasons across various horizons and scoring rules, showing potential to contribute to real-time collaborative forecasting challenges such as CDC's FluSight. The proposed approach offers an efficient or low-resource strategy to increase the impact of multi-model epidemic projections by providing real-time support to modeling teams, public health authorities, and decision-makers.
In 2024, the European Centre for Disease Prevention and Control (ECDC) launched RespiCompass, a scenario modelling hub for infectious respiratory diseases. The hub uses an ensemble approach, whereby projections from models developed by separate groups are combined to obtain mid-to-long term scenario-based projections. RespiCompass provides outputs designed to help public health experts and authorities assess respiratory virus risks and evaluate the potential impact of interventions, including vaccination. In building RespiCompass, ECDC has convened modelling teams, disease experts, and policymakers to work together in identifying policy-relevant questions, designing scenario rounds, and interpreting multi-model outputs. Drawing on lessons from other modelling hubs, RespiCompass successfully navigated the complex challenges of its first year, fostering cross-sector collaboration and delivering timely insights for the 2024/25 respiratory virus season. This article describes RespiCompass and distils key operational, technical, and collaborative lessons to inform future initiatives; the quantitative results from the first round are reported separately. 1
Individual-based models (IBMs) provide a mechanistic framework in which population-level outcomes emerge from interactions between individuals. We conducted a systematic review on IBMs for respiratory pathogens published in 2020-2024. We identified 855 eligible studies. Publications peaked in 2021, with a geographical distribution positively correlated with national GDP, leaving regions understudied. Most studies focused on SARS-CoV-2 and assessed public health interventions. Research priorities evolved over time, shifting from social distancing to vaccination. Age was included in 72.4% of studies; other sociodemographic factors (e.g., race/ethnicity) were rarely considered. This review maps the IBM landscape, offering a framework to guide future modeling efforts.
Accurate representations of the World Air Transportation Network (WAN) are fundamental inputs to models of global mobility, epidemic risk, and infrastructure planning. However, high-resolution, real-time data on the WAN are largely commercial and proprietary, therefore often inaccessible to the research community. Here we introduce a generative model of the WAN that treats air travel as a stochastic process within a maximum-entropy framework. The model uses airport-level passenger flows to probabilistically generate connections while preserving traffic volumes across geographic regions. The resulting reconstructed networks reproduce key structural properties of the WAN and enable simulations of dynamic spreading that closely match those obtained using the real network. Our approach provides a scalable, interpretable, and computationally efficient framework for forecasting and policy design in global mobility systems.
Abstract Background The 2026 FIFA World Cup may bring over one million visitors to North America from around the globe to participate in mass gathering events. The nature of the event and recent news have raised concerns for some that the tournament could lead to infectious disease outbreaks or fuel existing epidemics. Objective To systematically assess the infectious disease threat posed to the United States by the tournament. Design A multi-institutional team evaluated pathogen-specific risk across three dimensions: importation, outbreak potential, and impact to identify a priority pathogen list. A systematic screening protocol ensured common criteria and that pathogen information was collected when necessary to inform inclusion. Results Increased risk from the World Cup is near zero for 63 of 77 evaluated pathogens. Pathogens were predominantly excluded as threats due to low excess importation risk and low outbreak potential if introduced. The remaining priority pathogens fall into five categories: (a) mosquito borne pathogens with the potential for sustained transmission in some host cities, (b) seasonal respiratory viruses, (c) chronic infections with high prevalence outside the United States, (d) pathogens present in the United States with likely increased transmission at World Cup activities, and (e) high-consequence infectious threats. Limitations Data availability is variable across diseases. Impact calculations may not reflect actual costs to host cities. Disease incidence in World Cup travelers may differ from national incidence rates. Conclusion While infectious disease outbreaks at the 2026 FIFA World Cup are possible, in an already highly connected world where large gatherings are frequent, the elevated risk from the tournament is not as extreme as it first may seem. Primary Funding Source US Centers for Disease Control and Prevention
In 2023 the European Centre for Disease Prevention and Control (ECDC) launched RespiCast, the first European Respiratory Diseases Forecasting Hub, to provide probabilistic forecasts for influenza-like illness (ILI) and acute respiratory infection (ARI) incidence across 26 European countries. During the 2023/24 and 2024/25 winter seasons, RespiCast collected one- to four-week-ahead forecasts from multiple models contributed by different international teams and combined them into an ensemble. Our analysis shows that, when evaluated using the weighted interval score (WIS) and the absolute error (AE), the ensemble consistently outperformed the baseline model (defined as a persistence model that projects the last observed value forward) as well as individual models across most countries and forecasting rounds for both ILI and ARI incidence in the two seasons. Analysis of ensemble coverage (defined as the proportion of times observed values fall within the specified prediction intervals) indicated that forecast prediction intervals were reliable, although a general overconfidence trend (i.e., prediction intervals that are too narrow) was observed, particularly in specific countries. The relative performance of the ensemble declined in certain weeks, likely due to reduced participation from modelling teams, epidemic dynamics, higher data noise, and reporting delays. Forecast scores varied across countries, with some exhibiting consistently higher relative errors than others. Overall, the findings highlight the strengths of ensemble approaches in improving the accuracy and reliability of epidemiological forecasts while identifying areas for improvement, such as managing overconfidence and addressing variability in performance across countries and over time.
Human mobility, climate change and demographic trends increase the risk of pathogen spillover and expansion. Data that can inform our responses to outbreaks have increased in availability and volume, but access to highly confidential outbreak data and commercially sensitive contextual information remains difficult. Despite ongoing efforts to adopt global health data infrastructures and sharing protocols, there remain regulatory, logistical, human and computational barriers to data sharing. Federated approaches-in which data remain stored locally but analyses are performed across datasets from different sources-offer a potential way to address these challenges. While federated approaches have been used in some clinical and biomedical contexts, their adoption in infectious disease surveillance and modeling has been limited. Here, we discuss global approaches to infectious disease modeling and analysis, with a focus on federated methods. We outline how these can be used to address key epidemiological questions during outbreaks by enabling the secure use of multimodal data and integration with existing surveillance and modeling efforts. We summarize current methods for combining distributed and locally stored data and identify limitations, opportunities and organizational structures needed to achieve equitable global public health impacts.
Six years after its emergence, SARS-CoV-2 continues to have a substantial burden, however, the impact of vaccination and the optimal timing of its rollout remain uncertain. To explore these uncertainties, the US Scenario Modeling Hub convened its 19th round of ensemble projections for COVID-19 hospitalizations and deaths in the United States. Eight teams provided outcomes for each US state and nationally from April 2025 to April 2026 under five scenarios regarding vaccine recommendations and timing. We assessed recommendations with two eligibility scenarios (high-risk individuals only and all-eligible) and two timing scenarios (classic start: mid-August, earlier start: late June). These were crossed to create four scenarios and were compared against a counterfactual scenario with no vaccination. We found that compared to no vaccination, our ensemble projections estimated 90,000 (95% PI 53,000-126,000) hospitalizations averted in the high-risk and classic timing scenario across the US. Expanding coverage averted an additional 26,000 (95% PI 14,000-39,000) hospitalizations, which when coupled with earlier vaccination timing further reduced national hospitalizations by 15,000 (95% PI -3,000-33,000). These findings estimate significant benefits from a broad all-eligible vaccination recommendation, and suggest an additional benefit is likely to be gained from an earlier vaccination campaign.
Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making. ### Competing Interest Statement J.S. and Columbia University disclose partial ownership of SK Analytics. N.G.R. discloses paid consulting for Google Inc. J.Lemaitre discloses paid consulting for Pfizer Inc. The remaining authors declare no competing interests. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The forecast data for each model are publicly accessible from the FluSight Forecast Hub GitHub repository(https://github.com/cdcepi/FluSight-forecast-hub). The target hospitalization data are also available as weekly counts for each jurisdiction from U.S. Department of Health & Human Services. Weekly Hospital Respiratory Data (HRD) Metrics by Jurisdiction, National Healthcare Safety Network (NHSN), https://data.cdc.gov/Public-Health-Surveillance/Weekly-Hospital-Respiratory-Data-HRD-Metrics-by-Ju/ua7e-t2fy/about_data All data analyzed were publicly available prior to the initiation of this study. The influenza hospital admission data are aggregated weekly counts reported at the jurisdiction level and contain no individually identifiable information. The forecast data are model outputs submitted by participating teams to a public repository. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The forecast data for each model are publicly accessible from the FluSight Forecast Hub GitHub repository(https://github.com/cdcepi/FluSight-forecast-hub; https://doi.org/10.5281/zenodo.22101290). The target data are also available as weekly counts for each jurisdiction from HHS CDC's Center for Forecasting and Outbreak Analytics, CDC-RFA-FT-23-0069 CSTE/CDC, NU38OT000297, NU38PW000005, NU38OT000297 Centers for Disease Control and Prevention, U01IP001121, 75D30123C15907, U01IP001122, 6NU50CK000555-03-01 Natural Sciences and Engineering Research Council of Canada, ALLRP 581756-23 National Institute of General Medical Sciences, R35GM119582, R24GM153920, R35GM156799 National Institute of Allergy and Infectious Diseases, R01AI163023 National Institutes of Health / National Institute of General Medical Sciences, R35GM156799, DBI-2412389 National Institutes of Health / National Institute of General Medical Sciences, R01GM111510 National Institutes of Health, 5R01AI102939
Aircraft wastewater surveillance has been proposed as a new approach to monitor the global spread of pathogens. Here we develop a computational framework providing actionable information for the design and estimation of the effectiveness of global aircraft-based wastewater surveillance networks (WWSNs). We study respiratory diseases of varying transmission potential and find that networks of 10-20 strategically placed wastewater sentinel sites can provide timely situational awareness and function effectively as an early warning system. The model identifies potential blind spots and suggests optimization strategies to increase WWSN effectiveness while minimizing resource use. Our findings indicate that increasing the number of sentinel sites beyond a critical threshold does not proportionately improve WWSN capabilities, emphasizing the importance of resource optimization. We show, through retrospective analyses, that WWSNs can notably shorten detection time for emerging pathogens. The approach presented offers a realistic analytic framework for the analysis of WWSNs at airports.
West Nile virus (WNV) is the most common cause of mosquito-borne disease in the continental USA, with an average of 1200 severe, neuroinvasive cases reported annually from 2005 to 2021 (range 386–2873). Despite this burden, efforts to forecast WNV disease to inform public health measures to reduce disease incidence have had limited success. Here, we analyze forecasts submitted to the 2022 WNV Forecasting Challenge, a follow-up to the 2020 WNV Forecasting Challenge. Forecasting teams submitted probabilistic forecasts of annual West Nile virus neuroinvasive disease (WNND) cases for each county in the continental USA for the 2022 WNV season. We assessed the skill of team-specific forecasts, baseline forecasts, and an ensemble created from team-specific forecasts. We then characterized the impact of model characteristics and county-specific contextual factors (e.g., population) on forecast skill. Ensemble forecasts for 2022 anticipated a season at or below median long-term WNND incidence for nearly all (> 99
We present Epydemix, an open-source Python package for the development and calibration of stochastic compartmental epidemic models. The framework supports flexible model structures that incorporate demographic information, age-stratified contact matrices, and dynamic public health interventions. A key feature of Epydemix is its integration of Approximate Bayesian Computation (ABC) techniques to perform parameter inference and model calibration through comparison between observed and simulated data. The package offers a range of ABC methods such as simple rejection sampling, simulation-budget-constrained rejection, and Sequential Monte Carlo (ABC-SMC). Epydemix is modular, and supports ABC-based calibration both for models defined within the package and for those developed externally. To demonstrate the computational framework capabilities, we discuss usage examples that include (i) simulating an intervention-driven model with time-varying parameters, and (ii) benchmarking calibration performance using synthetic epidemic data. We further illustrate the use of the package in a retrospective case study that includes scenario projections under alternative intervention assumptions. By lowering the barrier for the implementation of computational and inference approaches, Epydemix makes epidemic modeling more accessible to a wider range of users, from academic researchers to public health professionals.
Human-AI coevolution, defined as a process in which humans and AI algorithms continuously influence each other, increasingly characterises our society, but is understudied in artificial intelligence and complexity science literature. Recommender systems and assistants play a prominent role in human-AI coevolution, as they permeate many facets of daily life and influence human choices through online platforms. The interaction between users and AI results in a potentially endless feedback loop, wherein users' choices generate data to train AI models, which, in turn, shape subsequent user preferences. This human-AI feedback loop has peculiar characteristics compared to traditional human-machine interaction and gives rise to complex and often “unintended” systemic outcomes. This paper introduces human-AI coevolution as the cornerstone for a new field of study at the intersection between AI and complexity science focused on the theoretical, empirical, and mathematical investigation of the human-AI feedback loop. In doing so, we: (i) outline the pros and cons of existing methodologies and highlight shortcomings and potential ways for capturing feedback loop mechanisms; (ii) propose a reflection at the intersection between complexity science, AI and society; (iii) provide real-world examples for different human-AI ecosystems; and (iv) illustrate challenges to the creation of such a field of study, conceptualising them at increasing levels of abstraction, i.e., scientific, legal and socio-political.
Importance:COVID-19 remains a disease with high burden in the US, prompting continued debate about optimal targets for annual vaccination. Objective:To project COVID-19 burden in the US for April 2024 to April 2025 under 6 scenarios of immune escape (20% and 50% per year) and levels of vaccine recommendation (no recommendation, vaccination for individuals at high risk only, vaccination for all eligible groups) and to assess the potential benefit of vaccine recommendations in reducing disease burden. Design, Setting, and Participants:For this decision analytical model, the US Scenario Modeling Hub, a collaborative modeling effort, convened 9 teams to provide scenario projections of US COVID-19 hospitalizations and deaths for April 2024 to April 2025, under 6 scenarios combining levels of immune escape and possible vaccine recommendations. Exposure:Annually reformulated vaccines were assumed to be 75% effective against hospitalization for variants circulating on June 15, 2024, and available on September 1, 2024. Age- and state-specific coverage was assumed to be as reported in September 2023 to April 2024. Main Outcomes and Measures:Ensemble estimates were made for weekly COVID-19 hospitalizations and deaths. Projections are presented for relative and absolute prevented hospitalizations and deaths averted due to vaccination over the April 2024 to April 2025 period. Results:For the US population (332 million, with an estimated 58 million aged ≥65 years), COVID-19 was expected to cause 814 000 (95% projection interval [PI], 400 000-1.2 million) hospitalizations and 54 000 (95% PI, 17 000-98 000) deaths for April 2024 to April 2025, comparable in magnitude to the prior year. Vaccination of high-risk groups only was projected to reduce hospitalizations (compared to no vaccination recommendation) by 76 000 (95% CI, 34 000-118 000) and deaths by 7000 (95% CI, 3000-11 000) across both immune escape scenarios. Compared with vaccinating high-risk groups only, a universal vaccine recommendation was projected to provide direct and indirect benefits, further preventing 11 000 hospitalizations and 1000 deaths in those aged 65 years and older. Conclusions and Relevance:In this decision analytical modeling study of COVID-19 burden in the US in 2024 to 2025, ensemble projections suggested that although vaccinating high-risk groups had substantial benefits in reducing disease burden, maintaining the vaccine recommendation for all individuals had the potential to save thousands more lives. Despite divergence of projections from observed disease trends in 2024 to 2025-possibly driven by variant emergence patterns and immune escape-averted COVID-19 burden due to vaccination was robust across immune escape scenarios, emphasizing the substantial benefit of broader vaccine availability for all individuals.
Human-AI coevolution, defined as a process in which humans and AI algorithms continuously influence each other, increasingly characterises our society, but is understudied in artificial intelligence and complexity science literature. Recommender systems and assistants play a prominent role in human-AI coevolution, as they permeate many facets of daily life and influence human choices through online platforms. The interaction between users and AI results in a potentially endless feedback loop, wherein users' choices generate data to train AI models, which, in turn, shape subsequent user preferences. This human-AI feedback loop has peculiar characteristics compared to traditional human-machine interaction and gives rise to complex and often "unintended" systemic outcomes. This paper introduces human-AI coevolution as the cornerstone for a new field of study at the intersection between AI and complexity science focused on the theoretical, empirical, and mathematical investigation of the human-AI feedback loop. In doing so, we: (i) outline the pros and cons of existing methodologies and highlight shortcomings and potential ways for capturing feedback loop mechanisms; (ii) propose a reflection at the intersection between complexity science, AI and society; (iii) provide real-world examples for different human-AI ecosystems; and (iv) illustrate challenges to the creation of such a field of study, conceptualising them at increasing levels of abstraction, i.e., scientific, legal and socio-political.
Bruno Goncalves合作论文数Indiana University29