
Abstract Rabies virus (RABV) is a lethal zoonotic pathogen that remains endemic in some parts of the world despite long-standing vaccination efforts. We present a comprehensive genomic and evolutionary analysis of the sylvatic rabies epidemic in Switzerland (1967–1997) and propose a molecular epidemiological framework describing RABV dynamics from emergence to elimination, with relevance for countries pursuing wildlife rabies elimination. We generated almost whole-genome sequences from 88 RABV-positive brain samples collected across 17 Swiss cantons and 13 host species over three decades. Using Illumina sequencing, 76 genomes were recovered at high coverage, enabling robust Single Nucleotide Polymorphism (SNP) and molecular clock analyses. Phylogenetic reconstruction based on Swiss whole genomes and glycoprotein sequences including additional European strains identified three temporally structured RABV lineages—red (1967–1976), blue (1976–1988), and green (1984–1996)—each consistent with independent introductions from neighbouring countries. Evolutionary analyses indicated a slow substitution rate and predominant purifying selection, with limited adaptive amino-acid change. However, four mutations emerged during the epidemiological peak, suggesting possible modest fitness effects. We found no evidence of major genomic changes associated with cross-species spillovers. Four synonymous variants were underrepresented in carnivore hosts; however, larger datasets may provide stronger evidence for host-associated genetic differences. Genetic divergence showed a strong temporal signal and moderate correlation with geographic distance, consistent with natural and vaccination-induced barriers limiting RABV dispersal, in agreement with the known ecology and control history of the epidemic. These findings provide a genomic framework for understanding sylvatic rabies emergence and elimination and may inform strategies to achieve and sustain the World Health Organization (WHO) goal of zero human deaths from dog-mediated rabies.
Human adenovirus type C (HAdV-C) causes upper respiratory infections in children and may lead to severe pneumonia. During the implementation and subsequent relaxation of non-pharmaceutical interventions, HAdV-C emerged as a transiently dominant circulating strain in Beijing. However, the fine-scale genomic architecture and the evolutionary trajectories governing its recombination remains insufficiently characterized. Between March 2023 and August 2024, respiratory samples from Beijing patients were collected and screened by quantitative PCR. In conjunction with high-throughput whole-genome sequencing, five complete HAdV-C genomes were characterized, predominantly identified as genotypes C1 and C108, maintaining over 98% intra-typic sequence identity. Phylogenetic analysis based on whole genomic sequences revealed at least five evolutionary branches within C108. Phylogenetic reconstruction revealed a complex diversification within C108, delineating at least five distinct evolutionary clades. Specifically, the four C108 strains partitioned into two divergent sub-lineages, exhibiting close phylogenetic affinities with sequences from China and the United States. Furthermore, recombination analysis identified six discrete recombination patterns. Selection pressure analysis further demonstrated heterogenous evolutionary constraints across the genome; notably, immune-relevant early genes such as E1a_26KD exhibited elevated dN/dS ratios, harbouring multiple positive selection sites. These adaptive mutations were distributed across 23 of 33 annotated genes (69.7%), suggesting extensive diversifying selection. These findings elucidate that HAdV-C evolution is synergistically driven by frequent recombination events and potent selective pressures. This study provides critical evidence for the spatiotemporal dynamics and genomic surveillance of the emerging HAdV-C variants.
Killer yeasts secrete protein toxins that inhibit other yeast strains, a trait often encoded by M satellites of L-A dsRNA viruses. These systems serve as important molecular models, yet their adaptive significance in natural settings remains unclear. This study surveyed 60 strains from a natural Saccharomyces paradoxus population in the UK to characterize dsRNA viral genomic diversity. Our survey revealed 27% of strains showed killer activity mediated by dsRNA viruses. Additionally, five of 18 nonkiller strains also contained dsRNA viruses. Deep sequencing of pooled dsRNAs from 17 strains identified six distinct M satellite types, including three novel lineages, with single-nucleotide and structural polymorphisms creating multiple variant sequences within types. Predicted preprotoxin proteins generally contained post-translational modification sites necessary for toxin maturation, although with variable numbers among different types, potentially affecting processing and expression. Eight L-A virus sequences were also assembled. These were all closely related to each other and clustered phylogenetically with viruses from other European S. paradoxus strains. Extended 5' sequence analysis revealed novel structural features in both L-As and their M satellites, including a pair of inverted repeats ending in a pair of inverted conserved motifs (GA5-6 and corresponding U5-6C in Ms and GAAUA and corresponding UAUUC in L-As), which is, in turn, flanked by a pair of direct repeats. The pair of conserved motifs also exists in all described S. cerevisiae M satellites. The observed diversity of dsRNA satellites within a single yeast population is a challenge to explain, with many evolutionary forces potentially contributing.
Understanding the sources of genetic diversity in Influenza A Virus (IAV) infections is crucial for understanding the mechanisms of viral evolution and immune escape. Whereas prior studies have characterized the effects of population bottlenecks during host-to-host transmission and intrahost tissue-to-tissue dissemination, the role of intracellular replication processes on IAV genetic diversity remains largely unexplored. In this study, we used stochastic mathematical modelling to simulate the replication of genetically distinct IAV strains within individual cells and tissues. Our results reveal significant bottleneck effects within a single infection cycle of individual cells. Intracellular bottleneck effects are driven by stochastic molecular processes and lead to the expansion or elimination of neutral variants creating large-scale differences between the initial and final frequencies of genetic variants in individual cells. By expanding our findings to a population-level model, we show that IAV intracellular replication reduces the effective population size, thereby diminishing the impact of selection and increasing the role of genetic drift. Our findings highlight the impact of intracellular replication processes on IAV genetic diversity.
Parvoviridae (small, nonenveloped ssDNA viruses) currently includes 281 species in two vertebrate- and four invertebrate-infecting subfamilies. While parvovirus-derived sequences are frequently identified in viromes, their taxonomy and host affiliation can be challenging due to high host and genetic diversity. We investigated the faecal parvovirome of 46 bats (7 insectivorous and 1 frugivorous species) from Eswatini and identified 28 novel viral species in 29 individuals (63.0%). The majority of these (22/28, 78.6%) belonged to nine genera (including two that are previously undescribed) within the invertebrate-infecting subfamily Densovirinae. A novel virus in the genus Brevipenbrevirus (arthropod-infecting subfamily Penbrevirinae) was found in 19.6% of the animals, including several frugivorous Epomophorus wahlbergi bats. A novel bat protoparvovirus (vertebrate-infecting subfamily Parvovirnae) was found both in the faeces and blood of one Afronycteris nanus bat. A highly divergent virus (Swazi bat-associated megaparvovirus 1, SwaBA-MePV-1) was found in the faeces, but not in the blood, of two insectivorous bats (Mops pumilus and Scotophilus viridis). Compared to other parvoviruses, SwaBA-MePV-1 presented two additional coding cassettes, significantly increasing its genome size. Homology modelling showed capsid protein C-terminal elongation, a previously undescribed strategy of parvoviral particle size expansion. Exploring public repositories identified 10 related uncharacterized viruses with similar genome organization and complete endogenous viral elements (EVEs) in eight beetle species, suggesting a coleopteran host affiliation. The complete genome of another highly divergent virus (SwaBA microparvovirus 1), without any detectable exogenous or endogenous relatives, was found in the faeces, but not in the blood, of one insectivorous Mops condylurus bat. Importantly, when comparing our sequences to references in Genbank, we observed that taxonomic mislabelling in sequence repositories can seriously misguide automatic taxonomy assignments (~75% of sequences initially identified as parvoviruses were discarded as false positives). These errors are amplified as new mislabelled sequences become dominant, highlighting the importance of prioritizing taxonomy validation and correct annotations in repositories. This study demonstrates that faecal samples from insectivorous chiropterans are rich in (novel) parvoviruses from various hosts. The discovery of highly divergent lineages (outside current sub-families) and EVEs helps clarify parvovirus evolutionary history and emphasizes how much of the parvoviral world remains unexplored.
Respiratory syncytial virus (RSV) remains a leading cause of hospitalization despite the recent implementation of infant immunoprophylaxis programs. The inter- and intrahost levels of genetic variation across the genome have not been extensively investigated. Hence, the aim of this study was to characterize the whole genome of RSV strains circulating within two Italian areas during the 2024-5 RSV season, corresponding to the first season after implementation of the national nirsevimab prophylaxis program. Whole genome sequencing (WGS) was performed on 88 respiratory samples that tested positive for RSV. Clinical records showed that none of the sequenced patients had received nirsevimab prophylaxis. The WGS strains obtained in the present study were correlated with the RSV strains circulating in Europe using phylogenetic inference. A newly designed bioinformatic pipeline was developed to evaluate the extent of genetic variation within hosts and reveal the distribution of minority mutations among RSV lineages. RSV-A (n = 54) showed significantly higher levels of nucleotide diversity than RSV-B (n = 34). Among the RSV-A strains, lineage A.D.3 was the most prevalent (35/54, 64%), followed by A.D.1 (18/54, 33%). The A.D.5 lineage was identified in only one strain, while all 34 RSV-B strains were found to be lineage B.D.E. Analysis of variants across the major lineages (A.D.1, A.D.3, and B.D.E) revealed heterogeneous mutation patterns. The G, L, and F genes showed the highest variability. Minority viral variants (0.02-0.05 frequency) followed the same gene-specific dynamics as the broader quasispecies population. A > G transitions consistent with ADAR activity were detected in all samples, but their frequency was similar to the other nucleotide substitutions across the genomes. The analysis of minor variants also targeted mutations known to reduce nirsevimab efficacy, and none were detected. Overall, RSV-A showed substantially greater intragenotypic and intrahost diversity than RSV-B. The elevated diversity in RSV-A resulted from a wider distribution of mutations across the viral genome, including both high-frequency substitutions that shaped the consensus sequence and low-frequency variants contributing to the underlying quasispecies complexity. Although no nirsevimab-treated patients were included in this study, the observed intrahost diversity highlights the importance of continued genomic surveillance to detect potential escape variants as prophylactic programs expand.
Viral genes sometimes use certain codons more than others due to their nucleotide content, translational efficiency, and selection pressure from the host immune system. The rabies virus (RABV) is a negative strand RNA virus which can infect a broad range of mammalian hosts, with many of its clades circulating predominantly in specific host species. Previous work on codon usage in RABV has focused only on broader viral clades. We use publicly available RABV nucleoprotein gene sequences to investigate how dinucleotide content and codon usage biases differ between host-associated clades, and what drives these differences. We found that codon usage varies most between bat- and carnivore-associated RABV clades, and more subtly between host-species-specific minor clades within these groups. C1, GA3, and GT3 content were found to have a strong influence over codon usage patterns, and cytosine-guanine dinucleotide (CpG) content was higher in carnivore-associated RABV clades than in bat-associated clades. This, along with lower numbers of zinc-finger antiviral protein binding motifs than would be expected based on (di)nucleotide composition in bat-associated RABV sequences, suggests that bat-associated RABV clades may be under higher selection pressure from the host's zinc-finger antiviral protein than carnivore-associated clades are, warranting further investigation of the mechanism underpinning this change.
It is known that several endogenous retroviruses, remnants of ancient retroviral integrations, retain envelope (env) genes encoding fusogenic proteins in primates. Whilst most env genes are degraded, a few have been co-opted by hosts, yet the entire evolutionary dynamics of env sequences remain poorly understood. To explore this, we screened and compared env open reading frames (ORFs) from 247 primate genomes. In total, 8683 nearly intact env-ORFs encoding 400 or more amino acids were identified, and their copy numbers ranged from 3 to 429 across primate species. Sequence similarity clustering revealed a clear evolutionary signature that distinguishes the small set of long-term co-opted env-derived genes, which are retained at low copy numbers across lineages, from the broader pool of rapidly turning-over env-ORFs. Applying this framework, we identified a previously unrecognized, single-copy env ortholog conserved across Tarsiiformes, named env-Tar1, and experimentally demonstrated its cell fusion activity in vitro, providing the first functional evidence of an env-derived gene co-opted in a tarsier lineage. Further, we found that certain co-opted env-derived genes may have lost their functions due to nonsense or indel mutations within specific primate lineages. Considering that many env genes tend to be maintained at low copy numbers and are reported to be under purifying selection, such dynamic evolutionary turnover of env-derived genes may be driven by host-virus arms races, as viruses and endogenous retroviruses often share cell-surface receptors.
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), responsible for the Coronavirus Disease of 2019 (COVID-19) pandemic, can productively infect a variety of animal hosts. The evidence suggests that both wild and farmed animal populations supporting continuous transmission present us with novel concerning genotypes of SARS-CoV-2. This is especially worrisome within large and dense populations of farmed animals, such as minks, for which spillbacks of 'mink genotypes' to the human population have been recorded on numerous occasions. In this study, we present the results of continuous clinical, virological, serological, and genomic surveillance during the period of 11 months of the SARS-CoV-2 outbreak at the largest Latvian mink farm with ˃300 000 animals. Using the One Health approach, the COVID-19 status of minks and farm workers was constantly monitored during the surveillance period. The presence of the SARS-CoV-2 genome was confirmed in 299 minks and 32 farm workers during the outbreak. The phylogenetic analysis of 188 mink and 14 human SARS-CoV-2 isolates linked to the farm provided insight into the evolution of the virus in situ and contributed to epidemiological investigation. The results revealed that SARS-CoV-2 lineage B.1.177.60 was initially introduced to the mink farm by an infected farm worker between 17 February and 23 March 2021, and subsequently spread among the minks. Surveillance in the affected farm showed fluctuating virus circulation. Although the average seroprevalence in samples taken from live minks was 76.92%, a fluctuating course of infection was observed from April 2021 to March 2022. Despite the implementation of strict preventive and control measures at the farm, several additional SARS-CoV-2 strain introductions were identified during 11 months. The initial introduction of a common viral strain circulating among the people at the time soon resulted in the co-circulation of multiple sister sublineages that have evolved some concerning spike protein amino acid substitutions. Subsequent other lineage introductions into the farm were not able to spread among minks. Several independent cases of farm workers infected with genotypes restricted to this mink farm were documented throughout the study timeframe. However, the 'Latvian mink genotypes' were not detected in the general human population beyond epidemiologic association with the given mink farm.
Arboviruses evolve under unique ecological constraints imposed by their dual replication cycles in vertebrate and arthropod hosts. This dual-host requirement results in markedly low substitution rates, which complicate molecular clock calibration, particularly when temporal sampling spans only narrow time windows. For RNA arboviruses, wide sampling windows are especially rare due to the intrinsic instability of RNA. Here, we demonstrate that ethanol-preserved museum specimens can help overcome these temporal limitations. We successfully recovered the coding-complete genome of a bandavirus, a negative-sense segmented RNA virus that clusters with the highly pathogenic human severe fever with thrombocytopenia syndrome virus. The virus was detected in a Common pipistrelle (Pipistrellus pipistrellus) bat collected in northern Germany in 1919, making it one of the oldest sequenced mammalian RNA viruses, only comparable to historic measles and influenza A viruses from 1912 and 1918. Screening 1086 contemporary bat samples revealed strains of the same virus species in nine Common pipistrelle bats (2010-2018) and one Serotine bat (Eptesicus serotinus, 1999) from Germany and the Netherlands. Coding-complete genomes indicate frequent genome segment reassortment and widespread circulation of reassortants of this understudied virus species (Bandavirus zwieselense). We detected an exceptionally low substitution rate (< 6.88 × 10-5 substitutions/site/year) between the RdRp coding sequence of the ancient genome and its nearest modern counterpart. Additionally, functional assays demonstrated that the virus's non-structural (NSs) protein effectively inhibits interferon induction in human HEK-293T cells. Our findings highlight the feasibility and scientific value of extracting and analysing ancient viral RNA from ethanol-preserved museum specimens to substantially enhance our understanding of RNA arbovirus evolution.
Using an in silico data mining approach, we previously expanded the known diversity within the family Bornaviridae by identifying novel bornavirids in publicly available sequence datasets from bony and cartilaginous fish. Building on this discovery, we systematically screened numerous additional sequence datasets from fish, amphibians, reptiles, and birds in the Sequence Read Archive and a new RNA-seq dataset to further explore bornaviral diversity and evolutionary history. Our analysis identified nearly complete and partial bornaviral genomes across multiple host taxa, that represented either new viral variants, previously known viruses detected in so far unreported hosts/sources, or novel viruses. Several novel orthobornaviral genomes were identified in snake and lizard datasets. Sequences of the species Orthobornavirus alphapsittaciforme and Orthobornavirus serini were identified in avian datasets. We also found sequences of both known and possibly new species in the genus Cultervirus, which were detected in bony fish and, for the first time, in reptile and amphibian datasets. The identification of culterviruses in reptiles and amphibians suggests that this group of viruses, which was previously only known to be associated with fish, has a much wider host range. A genetically divergent bornaviral genome was discovered in the gold-striped salamander (Chioglossa lusitanica), possessing the genomic architecture of orthobornaviruses but likely representing a new genus within the family Bornaviridae. Interestingly, multiple distinct viral genomes identified in passerine birds were classified as carboviruses based on sequence similarities, but they exhibited distinct genomic features and lacked the genes that code for glycoprotein (G) and, in some cases, the accessory protein (X) and/or matrix protein (M). Subsequent screening of samples from passerine birds further confirmed these findings and showed that these carboviral sequences likely represent genuine viruses rather than endogenous bornaviral elements. Interestingly, these carboviruses appear to co-occur with enveloped viruses in potential co-infection or helper virus scenarios. These findings contribute substantially to our understanding of Bornaviridae diversity, their evolutionary adaptation as well as modes of transcriptional regulation.
Lassa fever is a viral haemorrhagic fever that poses a persistent public health threat in several West African countries, particularly Nigeria. The scarcity of Lassa virus (LASV) sequences isolated from small mammal reservoirs limits our knowledge and understanding of LASV genomic diversity and transmission dynamics. To address this knowledge gap, we sampled 1189 small mammals, including mice, rats, and shrews, from two LASV-endemic states in southern Nigeria (Ondo and Ebonyi States) and tested them for the presence of LASV RNA using reverse transcription-quantitative polymerase chain reaction. Selected quantitative polymerase chain reaction-positive samples were subjected to whole genome sequencing and small mammal speciation through next-generation sequencing outputs. We recorded an overall polymerase chain reaction positivity rate of 61.6%, with rat species demonstrating the highest LASV prevalence. We also conducted a serosurvey of 269 small rodents using indirect Enzyme-Linked Immunosorbent Assay (ELISA) and obtained an overall anti-LASV seroprevalence of 45%. Using the Nextera XT metagenomic sequencing protocol, we produced 55 LASV partial (n = 28) and full-length genomes (n = 27) from small mammals sampled, all of which clustered within sublineage 2g. LASV sequences generated from this study suggest that LASV variation is mostly driven by location, as isolates from this study tend to cluster more closely with other isolates collected from within the same region, rather than by collection date or host. However, samples collected from Ebonyi State were more closely related to isolates collected in Ondo State than to isolates from Edo, despite a larger physical distance. Overall, the data from this study suggest free movement of the virus across states in Nigeria, among humans and various non-human taxa. The finding of LASV in additional small mammal hosts suggests that the virus reservoir is vast and may include many small mammals not well-characterized.
Matryoshka RNA virus 1 (MaRNAV-1) is a bi-segmented and single-stranded RNA virus associated with Plasmodium vivax, a cause of human malaria. Little has been uncovered about the epidemiology and ecology of this virus since its discovery in 2019. To address this, we used a combination of primary and publicly available metatranscriptomic data to map the geographic distribution and host associations of MaRNAV-1. We detected this virus throughout Southeast Asia, in parts of South America, and, for the first time, in Oceania. Despite its broad distribution, MaRNAV-1 was found exclusively in metatranscriptomes containing P. vivax, suggesting that there is a specific virus-host relationship that has shaped the evolutionary history of this virus. We were unable to estimate the emergence date of the MaRNAV-1 lineage; however, phylogeographic mapping analysis suggested that MaRNAV-1 is widely dispersed throughout Southeast Asia. Our findings have both evolutionary and public health implications and can serve as the basis for future investigations in these fields.
When the temperate phage [Formula: see text] infects its host, the bacterium Escherichia coli, it can either replicate to form new progeny (lytic growth) or integrate its genome into the host chromosome (lysogenization). Crucially, the probability of lysis or lysogeny is tightly regulated and varies with phage and host genotypes and the environment. In particular, the number of viruses coinfecting the same host cell is known to have a strong effect on the outcome of the infection, leading to phenotypic plasticity: the probability of lysogenization is higher when the cellular multiplicity of infection (MOI) increases. However, the selective forces driving the evolution of plasticity in phage [Formula: see text] remain unclear. Here, we analyze the evolution of the plasticity of lysogenization and show that a MOI-varying strategy is adaptive when the abundance of susceptible cells fluctuates periodically. We study how the speed of these fluctuations and various within-host decision rules affect the evolution of viral plasticity. Our results suggest that the complex genetic regulation of lysogeny in temperate phages can be shaped by natural selection allowing viruses to use the MOI as indirect information to persist in environments with fluctuating host densities.
Small ruminants (sheep and goats) are one of the few mammals in which an exogenous retrovirus (XRV) and closely related endogenous retroviral (ERV) elements coexist within the same host genome. The betaretroviruses Jaagsiekte sheep retrovirus (JSRV) and Enzootic Nasal Tumour Virus (ENTV) cause pulmonary and nasal adenocarcinomas, respectively, and share extensive sequence similarity with their endogenous counterparts. Consequently, molecular surveillance must rely on assays that can unequivocally distinguish true exogenous infection from ERV-derived templates; failure to do so compromises diagnosis, phylogenetic inference, and epidemiological conclusions. We retrieved all complete JSRV, ENTV-1/2, and related ERV genomes deposited in public repositories and performed a comprehensive alignment. Only a limited number of genomic segments were capable of distinguishing exogenous from endogenous sequences. We refer to these as discriminating regions (DRs). Phylogenies built using DRs revealed that several entries annotated as XRV are, in fact, ERV-derived or chimeric artefacts generated by short-amplicon reconstruction. A systematic literature review of over 100 articles identified 286 distinct primers and probes used for the XRV amplification. In-silico mapping of each oligonucleotide onto the full alignment showed that only 28% reliably differentiate XRV from ERV. We experimentally validated the predictive power of this approach for 17 primer/probe sets, confirming that non-discriminating assays produce false-positive signals from endogenous templates. The misannotation of ERV sequences as exogenous viruses has resulted in the population of databases with dubious entries, fostering erroneous hypotheses such as vector-borne transmission of JSRV and ENTV. To address this issue, we propose a concise set of criteria for assay design, validation, and database annotation, emphasizing DR targeting, specificity testing against endogenous templates, and transparent reporting. Although this framework was developed for small ruminants, it is readily applicable to any host-virus system in which exogenous viruses coexist with endogenous viral elements. This will strengthen viral surveillance, phylogenetics, and the One Health initiatives.
Investigating potential zoonotic viruses in animal reservoirs is crucial to anticipate viral emergence. Seals can represent large populations of coastal mammals with unknown consequences on the microbiological quality of their surrounding environment. To assess this, we conducted a metaviromics analysis of feces collected from two species of seals in the North-Western Atlantic (Saint-Pierre et Miquelon archipelago). We focused on the Caliciviridae family, which regroups several genera with viruses infecting humans and other mammals, including marine mammals, but none identified in seals (Phocidae). Among the assembled sequences identified as Caliciviridae, there were four known genera (norovirus, sapovirus, vesivirus, and salovirus) and unknown, distantly related viruses. Complete or nearly-complete genomes could be assembled for each genus. Norovirus and sapovirus sequences from seals were diverse and likely represent several new genogroups or genotypes. Seal vesivirus formed a monophyletic group, representing a potential new species related to the canine vesivirus. Salovirus, which are fish viruses, were likely diet-derived, like the distant sequences which exhibited the hallmarks of caliciviruses and were more closely related to fish and reptile viruses. In conclusion, seals are a reservoir for a large diversity of Caliciviridae, some related to norovirus or sapovirus genotypes known to infect humans, and their impact on the quality of coastal water or shellfish should be further assessed. This study expands the knowledge on Caliciviridae genetic diversity and circulation in marine mammals.
RNA viruses form genetically diverse populations structured as mutant spectra, or quasispecies, whose internal organization influences their evolutionary and adaptive dynamics. While genetic diversity has been extensively characterized, the structural organization of viral populations in sequence space remains less explored. Here, we compare genotype network architectures in two RNA viruses with markedly different evolutionary contexts: bacteriophage Qβ evolving in controlled laboratory conditions and SARS-CoV-2 evolving within infected human hosts. Using deep sequencing data, we reconstruct the genotype network of mutationally coupled variants within viral populations and analyze their topological properties. Despite large differences in genome size, mutation rate, and ecological setting, both viruses exhibit a common organization: a highly abundant central haplotype surrounded by layers of variants of diminishing abundance as Hamming distance to the central haplotype increases. All reconstructed networks share qualitative and quantitative topological features, displaying a hierarchical structure. The robust organization of both populations under multiple conditions suggests that RNA viruses may share a common genotype network architecture governed by fundamental properties of sequence space and the generic mechanisms of replication and mutation. Genotype networks provide a unifying framework to describe viral population structure beyond conventional diversity measures and, by revealing how local constraints shape mutational search, offers insights into the predictability of viral evolution.
The order Picornavirales is a group of highly diverse RNA viruses that includes many pathogens of significance to human and veterinary health, agriculture, and the wider environment. However, the wide range of viruses assigned to the order, together with their genomic variability, and the recent description of numerous 'picorna-like' viruses derived from metagenomic analyses of environmental samples, challenge the established taxonomic classification of members of the order and the criteria for their classification. Here, we combine the existing gold standard, hallmark RNA-directed RNA-polymerase (RdRP) gene sequence-based analysis with helicase sequence-based phylogeny, RdRP structural prediction through the use of ColabFold and Fold Tree, and analysis of coding-complete genomes using GRAViTy-V2, to genetically classify 525 picornaviral genomes and recently described 'picorna-like' viruses. All analyses were conducted with a bespoke, fully automated pipeline for retrieval of genome sequences, domain prediction and extraction, phylogenetic analysis, and output conditioning, which is available as open-source software. Our results reveal broad support for established families as well as for 6 novel families, and 32 new genera. In instances where inconsistencies were found between classification methods, we demonstrate how examination of the pipeline's output may be used to reconcile differences with respect to the genomic features quantified by the analysis. Automated multimodal taxonomic analysis may save significant resources over manual methods and better define demarcation criteria for families and genera.
Determining the sequence of the transmitted founder virus, the virus that establishes infection in a new host, is critical for understanding early viral dynamics and evolution. Methods to estimate transmitted founder virus sequences using ancestral sequence reconstruction require either sequences collected early in infection or longitudinal samples that can capture the evolutionary history of the viral sequences, which can be challenging to collect. In human immunodeficiency virus infection, viral genomes are integrated into host cells which can persist, creating a proviral archive of the evolutionary history of the virus. We can potentially utilize these proviral sequences, which can be collected later during infection and while the individual is on therapy, to perform ancestral sequence reconstruction to estimate founder virus sequences. We analysed a previously described data set of 12 participants from Zambia who had human immunodeficiency virus sequences collected within months of infection and proviral sequences collected before and after suppressive therapy initiation. We investigated the accuracy of root placement and founder virus sequence reconstruction in these individuals from their proviral sequences using a variety of phylogenetic methods. We had limited success in reconstructing founder virus sequences across all ancestral sequence reconstruction and rooting methods. However, we observed lower error in founder virus sequence reconstruction with participants that had proviral sequences similar to their founder sequence. Our results highlight a need for new methods to be developed in order to effectively reconstruct founder virus sequences from proviral sequences.
Previous studies based on clinical data from hepatitis C virus and hepatitis E virus infections revealed a deterministic evolution of quasispecies structure, irrespective of haplotype identities, towards a flat-like landscape, characterized by the absence of dominance and high evenness, combined with high haplotype synonymy. Here, two idealized limiting quasispecies states, A and Z, are defined, and it is shown that the A-to-Z evolution describes a parabolic trajectory between these states. The initial phase is dominated by increasing genetic diversity, whereas the subsequent phase is driven primarily by increasing evenness in the haplotype distribution. This evolutionary progression confers a broad domain within the genetic space, resulting in increased fitness and resilience, accompanied by a diminished response to antiviral therapies and multiple escape routes. Finally, a normalized quasispecies maturity score is proposed to position a given quasispecies along this structural evolutionary trajectory. This conceptual framework helps to account for the challenges in treating advanced chronic infections in which therapeutic failure frequently occurs in the absence of resistance-associated mutations.