Phylodynamic analysis has been instrumental in elucidating epidemiological and evolutionary dynamics of pathogens. Bayesian phylodynamics integrates out phylogenetic uncertainty, which is typically substantial in phylodynamic datasets due to limited genetic diversity. Phylodynamic inference does not, however, scale with modern datasets, partly due to difficulties in traversing tree space. Here, we characterize tree space and landscape in phylodynamic inference and assess its impacts on analysis difficulty and key biological estimates. By running extensive Bayesian analyses of 15 classic large phylodynamic datasets and carefully analyzing the posterior samples, we find that the posterior tree landscape is diffuse yet rugged, leading to widespread tree sampling problems that usually stem from sequences in a small part of the tree. We develop clade-specific diagnostics to show that a few sequences-including putative recombinants and recurrent mutants-frequently drive the ruggedness and sampling problems, although existing data-quality tests show limited power to detect them. The sampling problems can significantly impact phylodynamic inferences or distort major biological conclusions; the impact is usually stronger on "local" estimates (e.g., introduction history) associated with particular clades than on "global" parameters (e.g., demographic trajectory) governed by general tree shape. We evaluate existing Markov chain Monte Carlo diagnostics and diagnostics developed here, and offer strategies for optimizing phylodynamic analysis settings and mitigating sampling problem impacts. Our findings highlight the need and directions to develop efficient traversal over rugged tree landscapes, ultimately advancing scalable and reliable phylodynamics.
Bayesian inference has predominantly relied on the Markov chain Monte Carlo (MCMC) algorithm for many years. However, MCMC is computationally laborious, especially for complex phylogenetic models of time trees. This bottleneck has led to the search for alternatives, such as variational Bayes, which can scale better to large data sets. In this paper, we introduce torchtree, a framework written in Python that allows developers to easily implement rich phylogenetic models and algorithms using a fixed tree topology. One can either use automatic differentiation or leverage torchtree's plug-in system to compute gradients analytically for model components for which automatic differentiation is slow. We demonstrate that the torchtree variational inference framework performs similarly to BEAST in terms of speed, and delivers promising approximation results, though accuracy varies across scenarios. Furthermore, we explore the use of the forward Kullback-Leibler (KL) divergence as an optimizing criterion for variational inference, which can handle discontinuous and nondifferentiable models. Our experiments show that inference using the forward KL divergence is frequently faster per iteration compared with the evidence lower bound (ELBO) criterion, although the ELBO-based inference may converge faster in some cases. Overall, torchtree provides a flexible and efficient framework for phylogenetic model development and inference using PyTorch.
Since late 2020, the emergence of variants of concern (VOCs) of SARS-CoV-2 has been of concern to public health, researchers and policymakers. Mutations in the SARS-CoV-2 genome-for which clear evidence is available indicating a significant impact on transmissibility, severity and/or immunity-illustrate the importance of genomic surveillance and monitoring the evolution and geographic spread of novel lineages. Lineage B.1.619 was first detected in Switzerland in January 2021, in international travellers returning from Cameroon. This lineage was subsequently also detected in Rwanda, Belgium, Cameroon, France, and many other countries and is characterised by spike protein amino acid mutations N440K and E484K in the receptor binding domain, which are associated with immune escape and higher infectiousness. In this study, we perform a phylogeographic analysis to track the geographic origin and subsequent dispersal of SARS-CoV-2 lineage B.1.619. We employ a recently developed travel history-aware phylogeographic model, enabling us to incorporate genomic sequences with associated travel information. We estimate that B.1.619 most likely originated in Cameroon, in November 2020. We estimate the influence of the number of air-traffic passengers on the dispersal of B.1.619 but find no significant effect, illustrative of the complex dispersal patterns of SARS-CoV-2 lineages. Finally, we examine the metadata associated with infected Belgian patients and report a wide range of symptoms and medical interventions.
We define symmetric and asymmetric branching trees, a class of processes particularly suited for modeling genealogies of inhomogeneous populations where individuals may reproduce throughout life. In this framework, a broad class of Crump-Mode-Jagers processes can be constructed as (a)symmetric Sevast'yanov processes, which count the branches of the tree. Analogous definitions yield reduced (a)symmetric Sevast'yanov processes, which restrict attention to branches that lead to extant progeny. We characterize their laws through generating functions. The genealogy obtained by pruning away branches without extant progeny at a fixed time is shown to satisfy a branching property, which provides distributional characterizations of the genealogy.
BACKGROUND: Randomized clinical trials (RCTs) define evidence-based medicine, but quantifying their generalizability to real-world patients remains challenging. We propose a multidimensional approach to compare individuals in RCT and electronic health record (EHR) cohorts by quantifying their representativeness and estimating real-world effects based on individualized treatment effects (ITE) observed in RCTs. METHODS: We identified 65 pre-randomization characteristics of an RCT of heart failure with preserved ejection fraction (HFpEF), the Treatment of Preserved Cardiac Function Heart Failure with an Aldosterone Antagonist Trial (TOPCAT), and extracted those features from patients with HFpEF from the EHR within the Yale New Haven Health System. We then assessed the real-world generalizability of TOPCAT by developing a multidimensional machine learning-based phenotypic distance metric between TOPCAT stratified by region including the United States (US) and Eastern Europe (EE) and EHR cohorts. Finally, from the ITE identified in TOPCAT participants, we assessed spironolactone benefit within the EHR cohorts. RESULTS: There were 3,445 patients in TOPCAT and 8,121 patients with HFpEF across 4 hospitals. Across covariates, the EHR patient populations were more similar to each other than the TOPCAT-US participants (median SMD 0.065, IQR 0.011-0.144 vs median SMD 0.186, IQR 0.040-0.479). At the multi-variate level using the phenotypic distance metric, our multidimensional similarity score found a higher generalizability of the TOPCAT-US participants to the EHR cohorts than the TOPCAT-EE participants. By phenotypic distance, a 47% of TOPCAT-US participants were closer to each other than any individual EHR patient. Using a TOPCAT-US-derived model of ITE from spironolactone, all patients were predicted to derive benefit from spironolactone treatment in the EHR cohort, while a TOPCAT-EE-derived model predicted 13% of patients to derive benefit. CONCLUSIONS: This novel multidimensional approach evaluates the real-world representativeness of RCT participants against corresponding patients in the EHR, enabling the evaluation of an RCT's implication for real-world patients.
Here we present the open-source and cross-platform BEAST X software that combines molecular phylogenetic reconstruction with complex trait evolution, divergence-time dating and coalescent demographics in an efficient statistical inference engine. BEAST X significantly advances the flexibility and scalability of evolutionary models supported. Novel clock and substitution models leverage a large variety of evolutionary processes; discrete, continuous and mixed traits with missingness and measurement errors; and fast, gradient-informed integration techniques that rapidly traverse high-dimensional parameter spaces.
Infectious disease threats to individual and public health are numerous, varied and frequently unexpected. Artificial intelligence (AI) and related technologies, which are already supporting human decision making in economics, medicine and social science, have the potential to transform the scope and power of infectious disease epidemiology. Here we consider the application to infectious disease modelling of AI systems that combine machine learning, computational statistics, information retrieval and data science. We first outline how recent advances in AI can accelerate breakthroughs in answering key epidemiological questions and we discuss specific AI methods that can be applied to routinely collected infectious disease surveillance data. Second, we elaborate on the social context of AI for infectious disease epidemiology, including issues such as explainability, safety, accountability and ethics. Finally, we summarize some limitations of AI applications in this field and provide recommendations for how infectious disease epidemiology can harness most effectively current and future developments in AI.
Scientific studies in many areas of biology routinely employ evolutionary analyses based on inference of phylogenetic trees from molecular sequence data. Evolutionary processes that act at the molecular level are highly variable, and properly accounting for heterogeneity is crucial for more accurate phylogenetic inference. Nucleotide substitution rates and patterns are known to vary among sites in multiple sequence alignments, and such variation can be modeled by partitioning alignments into categories corresponding to different substitution models. Determining a priori appropriate partitions can be difficult, however, and better model fit can be achieved through flexible Bayesian infinite mixture models that simultaneously infer the number of partitions, the partition that each site belongs to, and the evolutionary parameters corresponding to each partition. Here, we consider several different types of infinite mixture models, including classic Dirichlet process mixtures, as well as novel approaches for modeling across-site evolutionary variation: hierarchical models for data with a natural group structure, and infinite hidden Markov models that account for spatial patterns in alignments. In analyses of several viral data sets, we find that different types of models perform best in different scenarios, but infinite hidden Markov models emerge as particularly promising for larger data sets and complex evolutionary patterns characterized by multiple genes and overlapping reading frames. To enable these models to scale to large data sets, we adapt efficient Markov chain Monte Carlo algorithms and exploit opportunities for parallel computing. We implement this infinite mixture modeling framework in BEAST X, a widely-used software package for phylogenetic inference.
Propensity score adjustment addresses confounding by balancing covariates in subject treatment groups through matching, stratification, or weighting. Diagnostics test the success of adjustment. For example, if the standardized mean difference (SMD) for a relevant covariate exceeds a threshold like 0.1, the covariate is considered imbalanced and the study may be biased. Unfortunately, for studies with small or moderate numbers of subjects, the probability of identifying a study as biased because of chance imbalance can be grossly larger than a given nominal level like 0.05, yet that chance imbalance may not cause significant bias. In this paper, we illustrate that chance imbalance is operative in real-world settings even for moderate sample sizes of 2000. We identify a previously unrecognized challenge that as meta-analyses increase the precision of an effect estimate, the diagnostics must also undergo meta-analysis for a corresponding increase in precision. We propose an alternative diagnostic that checks whether the SMD statistically significantly exceeds the threshold. Through simulation and real-world data, we find that this diagnostic achieves a better trade-off of type 1 error rate and power than standard nominal threshold tests and not testing for sample sizes from 250 to 4000 and for 20 to 100 000 covariates. We confirm that in network studies, meta-analysis of effect estimates must be accompanied by meta-analysis of the diagnostics or else systematic confounding may overwhelm the estimated effect. Our procedure supports the review of large numbers of covariates, enabling more rigorous diagnostics.
Multiproposal Markov chain Monte Carlo (MCMC) algorithms choose from multiple proposals to generate their next chain step in order to sample from challenging target distributions more efficiently. However, on classical machines, these algorithms require 𝒪 ( P ) target evaluations for each Markov chain step when choosing from P proposals. Recent work demonstrates the possibility of quadratic quantum speedups for one such multiproposal MCMC algorithm. After generating P proposals, this quantum parallel MCMC (QPMCMC) algorithm requires only 𝒪 ( P ) target evaluations at each step, outperforming its classical counterpart. However, generating P proposals using classical computers still requires 𝒪 ( P ) time complexity, resulting in the overall complexity of QPMCMC remaining 𝒪 ( P ) . Here, we present a new, faster quantum multiproposal MCMC strategy, QPMCMC2. With a specially designed Tjelmeland distribution that generates proposals close to the input state, QPMCMC2 requires only 𝒪 ( 1 ) target evaluations and 𝒪 ( log P ) qubits when computing over a large number of proposals P . Unlike its slower predecessor, the QPMCMC2 Markov kernel (1) maintains detailed balance exactly and (2) is fully explicit for a large class of graphical models. We demonstrate this flexibility by applying QPMCMC2 to novel Ising-type models built on bacterial evolutionary networks and obtain significant speedups for Bayesian ancestral trait reconstruction for 248 observed salmonella bacteria.
Bayesian phylogenetics typically estimates a posterior distribution, or aspects thereof, using Markov chain Monte Carlo methods. These methods integrate over tree space by applying local rearrangements to move a tree through its space as a random walk. Previous work explored the possibility of replacing this random walk with a systematic search, but was quickly overwhelmed by the large number of probable trees in the posterior distribution. In this paper we develop methods to sidestep this problem using a recently introduced structure called the subsplit directed acyclic graph (sDAG). This structure can represent many trees at once, and local rearrangements of trees translate to methods of enlarging the sDAG. Here we propose two methods of introducing, ranking, and selecting local rearrangements on sDAGs to produce a collection of trees with high posterior density. One of these methods successfully recovers the set of high posterior density trees across a range of data sets. However, we find that a simpler strategy of aggregating trees into an sDAG in fact is computationally faster and returns a higher fraction of probable trees.
The rapid rate of virus evolution, while useful for outbreak investigations, poses a challenge for accurately estimating long-term viral evolutionary divergence and leaves us with little genomic traces at deep evolutionary timescales, complicating the reconstruction of deep virus evolutionary history. Recent advancements in protein structure prediction and computational biology have opened up new avenues and enabled us to peer back further in time and with greater clarity than ever before. Here, we review recent approaches to reconstructing the deep evolutionary history of viruses. In particular, we focus on how Bayesian models that account for evolutionary rates that are time-dependent may provide better estimates of the timescale of virus evolution. We then outline approaches to structural phylogenetics and their application to reconstructing the evolutionary history of viruses. Despite current limitations, including structural prediction uncertainty, conformational variation, and limited benchmarking, structural phylogenetics appears promising, particularly where sequence-level homology is eroded. The availability of and ease with which virus structures can now be predicted is likely to drive additional statistical and software developments in this area. Ultimately, answering fundamental questions of virus origins and early diversification, long-term host associations, virus classification, and the timescale of viral diseases will likely require unifying sequence and structural information into a temporally aware evolutionary inference framework.
Background Fluoroquinolones (FQs) are commonly used to treat urinary tract infections (UTIs), but some studies have suggested they may increase the risk of aortic aneurysm or dissection (AA/AD). However, no large-scale international study has thoroughly assessed this risk. Methods A retrospective cohort study was conducted using a large, distributed network analysis across 14 databases from 5 countries (United States, South Korea, Japan, Taiwan, and Australia). The study included 13,588,837 patients aged 35 or older who initiated systemic fluoroquinolones (FQs) or comparable antibiotics (trimethoprim with or without sulfamethoxazole [TMP] or cephalosporins [CPHs]) for UTI treatment in the outpatient setting between JAN 01, 2010 and DEC 31, 2019. Patients were included if at the index date they had at least 365 days of prior observation and were not hospitalised for any reason on or within 7 days prior to the index date. The primary outcome was AA/ AD occurrence within 60 days of exposure, with secondary outcomes examining AA and AD separately. Cox proportional hazards models with 1:1 propensity score (PS) matching were used to estimate the risk, with results calibrated using negative control outcomes. Analyses were subjected to pre-defined study diagnostics, and only those passing all diagnostics were reported. Hazard ratios (HRs) were pooled using Bayesian random-effects meta-analysis. Findings Among analyses that passed diagnostics there were 1,954,798 and 1,195,962 propensity-matched pairs for the FQ versus TMP and FQ versus CPH comparisons respectively. For the 60-day follow-up there was no difference in risk of AA/AD between FQ and TMP (absolute rate difference [ARD], 0.21 per 1000 person-year; calibrated HR, 0.91 [95% CI 0.73-1.10]). There was no significant difference in risk for FQ versus CPH (ARD, 0.11 per 1000 person-year; calibrated HR, 1.01 [95% CI 0.82-1.25]). Interpretation This large-scale study used a rigorous design with objective diagnostics to address bias and confounding. There was no increased risk of AA/AD associated with FQ compared to TMP or CPH in patients treated for UTI in the outpatient setting. As we only examined FQ used to treat UTIs in the outpatient setting, the results may not be generalisable to other indications with different severity. Funding Yonsei University College of Medicine, Government-wide R&D Fund project for infectious disease research (GFID), Republic of Korea, National Health and Medical Research Council (NHMRC) Australian Government. Department of Veterans Affairs (VA) Informatics and Computing Infrastructure (VINCI), Department of Veterans Affairs, the United States Government. Copyright (c) 2025 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
The widespread adoption of real-world data has given rise to numerous healthcare-distributed research networks, but multi-site analyses still face administrative burdens and data privacy challenges. In response, we developed a Collaborative One-shot Lossless Algorithm for Generalized Linear Mixed Models (COLA-GLMM), the first-ever algorithm that achieves both lossless and one-shot properties. COLA-GLMM ensures accuracy against the gold standard of pooled data while requiring only summary statistics and completes within a single communication round, eliminating the usual back-and-forth overhead. We further introduced an enhanced version that employs homomorphic encryption to reduce the risks of summary statistics misuse at the coordinating center. The simulation studies showed near-exact agreement with the gold standard in parameter estimation, with relative differences of 7.8 × 10-6%-3.0% under various cell suppression settings. We also validated COLA‑GLMM on eight international decentralized databases to identify risk factors for COVID‑19 mortality. Together, these results show that COLA‑GLMM enables accurate, low‑burden, and privacy-preserving multi‑site research.
The 2022 global mpox epidemic was caused by transmission of MPXV clade IIb, lineage B.1 through sexual contact networks, with New York City (NYC) experiencing the first and largest outbreak in the United States. By performing phylogeographic analysis of MPXV genomes sampled from 757 individuals in NYC between April 2022 and April 2023, and 3,287 MPXV genomes sampled around the world, we identify over 200 introductions of MPXV into NYC with at least 84 leading to onward transmission. These infections primarily occurred among men who have sex with men, transgender women and nonbinary individuals. Through a comparative analysis with HIV in NYC, we find that both MPXV and HIV genomic cluster sizes are best fit by scale-free distributions, and that people in MPXV clusters are more likely to have previously received an HIV diagnosis and be a member of a recently growing HIV transmission cluster. We model MPXV transmission through sexual contact networks and show that highly connected individuals would be disproportionately infected at the start of an epidemic, which would likely result in the exhaustion of the most densely connected parts of the network, and, therefore, explain the rapid expansion and decline of the NYC outbreak. By coupling the genomic epidemiology of MPXV and HIV with epidemic modeling, we demonstrate that the transmission dynamics of MPXV in NYC can be understood by general principles of sexually transmitted pathogens.
While advances in molecular epidemiology and computational modeling have enhanced our capacity to track pathogen evolution, the accurate reconstruction of spatiotemporal transmission dynamics remains essential for developing epidemic preparedness frameworks and implementing outbreak response measures. Structured coalescent models offer a phylogeographic framework by restricting lineage coalescence events to geographically proximate host populations. Although the Bayesian structured coalescent approximation (BASTA) provides a tractable approach, contemporary phylogeographic analyses involving dozens of geographic localities and hundreds to thousands of viral genomes substantially exceed the computational capacity of existing implementations. The BASTA likelihood scales cubically with deme count and quadratically with sequence count due to matrix exponentiation and pairwise coalescent probability calculations. Here, we introduce a comprehensive algorithmic restructuring of the structured coalescent likelihood that eliminates redundancies, optimizes memory access, and exposes parallelization opportunities. Our approach reorganizes computations along three dimensions: (i) independent calculation of deme-transition probability matrices across time intervals; (ii) simultaneous evaluation of partial likelihood vectors within temporal slices; and (iii) concurrent aggregation of coalescent probabilities. Algorithmic restructuring cuts average coalescent likelihood computation by 7-8 fold, and parallelization further boosts performance to 10-26 fold, enabling joint phylogeographic analyses of dengue virus across 10 South American countries and H5N1 avian influenza across 20 Eurasian regions to finish in a fraction of prior time. This computational efficiency also enables comparison between backward-in-time structured coalescent approximations and forward-in-time phylogeographic methods, revealing that the former provides appropriately conservative posterior estimates, particularly at intermediate phylogenetic depths. We integrate our implementation into the popular BEAST X and BEAGLE software packages, with an accompanying interface in BEAUti X to easily set up the analyses, providing researchers with an accessible and scalable tool for real-time phylogeographic surveillance of rapidly evolving pathogens.