Crimean-Congo hemorrhagic fever virus (CCHFV) is a segmented RNA virus that can cause severe hemorrhagic fever and is primarily transmitted to humans and other mammals through tick bites. Since segment reassortment can increase genetic diversity and disease severity, we sampled infected ticks and humans in Sivas Province, Türkiye, a region with high CCHFV prevalence. Analysis of 113 human and 23 Hyalomma tick CCHFV samples revealed three main phylogenetic clades in Sivas. The most common clade (SIVAS-2) was involved in multiple reassortants. Phylogenetic analysis indicated multiple independent reassortment events involving all three CCHFV genome segments. The most common reassortant (1.1.2) was detected only in the Zara district of Sivas in both humans and ticks. We found that most infections were geographically limited to local spread and that infections with more than one variant and reassortment were exceptional. We found no significant clinical differences between human SIVAS-1, SIVAS-2 and the 1.1.2 reassortant infections. Thus, within the limits of the available sample size, we found no evidence that the two parental forms differed in disease severity or that the 1.1.2 reassortant was associated with increased severity. Continued CCHFV surveillance could uncover future forms of the virus with altered characteristics.
Motivation:Understanding how virus sequences are shaped by selection can inform vaccine design and transmission inference. Modeling within-host evolution to interrogate these questions requires a detailed mechanistic framework that accurately captures sequence diversification. The CD8+ cytotoxic T-lymphocyte (CTL) response plays an important role in immune-mediated selection and can leave strong signatures in virus sequences; however, existing sequence-based within-host virus modeling frameworks do not explicitly include a human leukocyte antigen (HLA)-aware CTL response. Results:We extended our previously published within-host sequence evolution simulator, wavess, to include an explicit CTL response, and share a method for identifying HLA-specific CTL epitopes given a founder virus sequence. We also updated the model to permit a variable recombination rate, which allows for modeling non-adjacent genes, segmented genomes, and recombination hotspots. These extensions to wavess allow for more accurate simulation of viruses and virus genes, particularly in regions of the genome where the immune response is dominated by CTLs (rather than antibodies). It also provides the foundation for investigations of how these newly-added biological mechanisms influence within-host evolution. Availability and implementation:The core of wavess is written in Python 3, with helper functions written in R. It is available at https://github.com/MolEvolEpid/wavess.
Crimean-Congo haemorrhagic fever virus (CCHFV) is an important human tick-borne pathogen, able to cause severe haemorrhagic fever. CCHFV is endemic in Tajikistan, which records between 5-38 cases of CCHF a year from southern regions. Molecular surveillance of CCHFV is crucial to implement effective prevention and control strategies, understand viral evolution, study transmission dynamics, and develop effective diagnostics, therapeutics, and vaccines. While the presence of Asia-1 and Asia-2 genotypes has been previously reported, only two historical samples from Tajikistan have been fully sequenced. In this study we developed and applied a genotype IV-specific tiling PCR enrichment approach recovering 52 CCHFV genome segment sequences from clinical and Hyalomma tick samples collected between 2017-2023. Most sequences belonged to the Asia-2 genotype, but one virus exhibited an Asia-1 S segment combined with Asia-2 M and L segments, representing the first evidence of such viral reassortment event in Tajikistan.
Motivation:The detection of APOBEC3F- and APOBEC3G-induced mutations in virus sequences is useful for identifying hypermutated sequences. These sequences are not representative of viral evolution and can therefore alter the results of downstream sequence analyses if included. We previously published the software Hypermut, which detects hypermutation events in sequences relative to a reference. Two versions of this method are available as a webtool. Neither of these methods consider multistate characters or gaps in the sequence alignment. Results:Here, we present an updated, user-friendly web and command-line version of Hypermut with functionality to handle multistate characters and gaps in the sequence alignment. This tool allows for straightforward integration of hypermutation detection into sequence analysis pipelines. As with the previous tool, while the main purpose is to identify G to A hypermutation events, any mutational pattern and context can be specified. Availability and implementation:Hypermut 3 is written in Python 3. It is available as a command-line tool at https://github.com/MolEvolEpid/hypermut3 and as a webtool at https://www.hiv.lanl.gov/content/sequence/HYPERMUT/hypermutv3.html.
The human immunodeficiency virus (HIV) can persist in a latent form as integrated DNA (provirus) in resting CD4+ T cells unaffected by antiretroviral therapy. Despite being a major obstacle for eradication efforts, it remains unclear which infected cells survive, persist, and ultimately enter the long-lived reservoir. Here, we determine the genetic divergence and integration times of simian immunodeficiency virus (SIV) envelope sequences collected from infected macaques. We show that the proviral divergence and the phylogenetically estimated integration times display a biphasic decline over time. Investigating the dynamics of the mutational distributions, we show that SIV genomes in short-lived cells are, on average, more diverged, while long-lived cells contain less diverged virus. The change in the mutational distributions over time explains the observed biphasic decline in the divergence of the proviruses. This suggests that long-lived cells harbor viruses deposited earlier in infection, while short-lived cells predominantly harbor more recent viruses.
Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007-2008 Swedish HIV-1 outbreak and the broader 1998-2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.
Simulating within-host virus sequence evolution allows for the investigation of factors such as the role of recombination in virus diversification and the impact of selective pressures on virus evolution. Here, we provide a new software to simulate virus within-host evolution called wavess (within-host adaptive virus evolution sequence simulator), a discrete-time individual-based model and a corresponding user-friendly R package. The underlying model simulates recombination, a latent infected cell reservoir, and three forms of selection: conserved sites fitness and replicative fitness in comparison to a reference sequence, and immune fitness including cross-reactivity imposed by a co-evolving immune response. In the R package, we also provide functions to generate model inputs from empirical data, as well as functions to analyze the simulation outputs. At user-defined time points, the software returns various counts related to the virus population(s) and a set of sampled virus sequences. We applied this model to investigate the selection pressures on HIV-1 env sequences longitudinally collected from 11 individuals. The best-fitting immune cost differed across individuals, mirroring the real-world expectation of heterogeneous immune responses among human hosts. Furthermore, the phylogenies reconstructed from these simulated sequences were similar to the phylogenies reconstructed from the real sequences for all summary statistics tested. To our knowledge, compared to other similar models, wavess has been more rigorously validated against real within-host virus sequences, and is the first to be implemented as an R package. The wavess R package can be downloaded from https://github.com/MolEvolEpid/wavess.
Reassortment is an evolutionary process common in viruses with segmented genomes. These viruses can swap whole genomic segments during cellular co-infection, giving rise to novel progeny formed from the mixture of parental segments. Since large-scale genome rearrangements have the potential to generate new phenotypes, reassortment is important to both evolutionary biology and public health research. However, statistical inference of the pattern of reassortment events from phylogenetic data is exceptionally difficult, potentially involving inference of general graphs in which individual segment trees are embedded. In this paper, we argue that, in general, the number and pattern of reassortment events are not identifiable from segment trees alone, even with theoretically ideal data. We call this fact the fundamental problem of reassortment, which we illustrate using the concept of the “first-infection tree,” a potentially counterfactual genealogy that would have been observed in the segment trees had no reassortment occurred. Further, we illustrate four additional problems that can arise logically in the inference of reassortment events and show, using simulated data, that these problems are not rare and can potentially distort our observation of reassortment even in small data sets. Finally, we discuss how existing methods can be augmented or adapted to account for not only the fundamental problem of reassortment, but also the four additional situations that can complicate the inference of reassortment.
Background. Hepatitis C virus (HCV) has high genetic diversity and is classified into 8 genotypes and >90 subtypes, with some endemic to specific world regions. This could compromise direct-acting antiviral efficacy and global HCV elimination. Methods. We characterized HCV subtypes "rare" in the United Kingdom (non-1a/1b/2b/3a/4d) by means of whole-genome sequencing via a national surveillance program. Genetic analyses to determine the genotype of samples with unresolved genotypes were undertaken by comparison with International Committee on Taxonomy of Viruses HCV reference sequences. Results. Two HCV variants were characterized as being closely related to the recently identified genotype (GT) 8, with >85% pairwise genetic distance similarity to GT8 sequences and within the typical intersubtype genetic distance range. The individuals infected by the variants were UK residents originally from Pakistan and India. In contrast, a third variant was only confidently identified to be more similar to GT6 compared with other genotypes across 6% of the genome and was isolated from a UK resident originally from Guyana. All 3 were cured with pangenotypic direct-acting antivirals (sofosbuvir-velpatasvir or glecaprevir-pibrentasvir) despite the presence of resistance polymorphisms in NS3 (80K/168E), NS5A (28V/30S/62L/92S/93S) and NS5B (159F). Conclusions. This study expands our knowledge of HCV diversity by identifying 2 new GT8 subtypes and potentially a new genotype.
BackgroundSweden reached the UNAIDS 90-90-90 target in 2015. It is important to reassess the HIV epidemiological situation due to ever-changing migration patterns, the roll-out of PrEP and the impact of the COVID-19 pandemic.AimWe aimed to assess the progress towards the UNAIDS 95-95-95 targets in Sweden by estimating the proportion of undiagnosed people with HIV (PWHIV) and HIV incidence trends.MethodsWe used routine laboratory data to inform a biomarker model of time since infection. When available, we used previous negative test dates, arrival dates for PWHIV from abroad and transmission modes to inform our incidence model. We also used data collected from the Swedish InfCareHIV register on antiretroviral therapy (ART).ResultsThe yearly incidence of HIV in Sweden decreased after 2014. In part, this was because the fraction of undiagnosed PWHIV had decreased almost twofold since 2006. After 2015, three of four PWHIV in Sweden were diagnosed within 1.9 and 3.2 years after infection among men who have sex with men and in heterosexual groups, respectively. While 80% of new PWHIV in Sweden acquired HIV before immigration, they make up 50% of the current PWHIV in Sweden. By 2022, 96% of all PWHIV in Sweden had been diagnosed, and 99% of them were on ART, with 98% virally suppressed.ConclusionsBy 2022, about half of all PWHIV in Sweden acquired HIV abroad. Using our new biomarker model, we assess that Sweden has reached the UNAIDS goal at 96-99-98.
When the time of an HIV transmission event is unknown, methods to identify it from virus genetic data can reveal the circumstances that enable transmission. We developed a single-parameter Markov model to infer transmission time from an HIV phylogeny constructed of multiple virus sequences from people in a transmission pair. Our method finds the statistical support for transmission occurring in different possible time slices. We compared our time-slice model results to previously-described methods: a tree-based logical transmission interval, a simple parsimony-like rules-based method, and a more complex coalescent model. Across simulations with multiple transmitted lineages, different transmission times relative to the source's infection, and different sampling times relative to transmission, we found that overall our time-slice model provided accurate and narrower estimates of the time of transmission. We also identified situations when transmission time or direction was difficult to estimate by any method, particularly when transmission occurred long after the source was infected and when sampling occurred long after transmission. Applying our model to real HIV transmission pairs showed some agreement with facts known from the case investigations. We also found, however, that uncertainty on the inferred transmission time was driven more by uncertainty from time-calibration of the phylogeny than from the model inference itself. Encouragingly, comparable performance of the Markov time-slice model and the coalescent model-which make use of different information within a tree-suggests that a new method remains to be described that will make full use of the topology and node times for improved transmission time inference.
With a single circulating vector-borne virus, the basic reproduction number incorporates contributions from tick-to-tick (co-feeding), tick-to-host and host-to-tick transmission routes. With two different circulating vector-borne viral strains, resident and invasive, and under the assumption that co-feeding is the only transmission route in a tick population, the invasion reproduction number depends on whether the model system of ordinary differential equations possesses the property of neutrality. We show that a simple model, with two populations of ticks infected with one strain, resident or invasive, and one population of co-infected ticks, does not have Alizon's neutrality property. We present model alternatives that are capable of representing the invasion potential of a novel strain by including populations of ticks dually infected with the same strain. The invasion reproduction number is analysed with the next-generation method and via numerical simulations.
Background and aimHepatitis C virus (HCV) infection is a major global public health concern, being a leading cause of chronic liver diseases such as chronic hepatitis, cirrhosis, and hepatocellular carcinoma. The virus is classified into 8 genotypes and 93 subtypes, each displaying distinct geographic distributions. Genotype 4 is the most predominant in the Middle East and Eastern Mediterranean and is associated with high rates of hepatitis C infection worldwide. This study used next-generation sequencing to fully characterize the HCV genome and identify a novel subtype within genotype 4 isolated from a 64-year-old Saudi man diagnosed with hepatitis C.MethodsWe analyzed the complete genome of the 141-HCV isolate using whole-genome sequencing.ResultsOur phylogenetic reconstructions, based on the entire genome of HCV-4 strains, revealed that the 141-HCV isolate formed a distinct group within the genotype 4 classification, providing valuable new insights into the variability of HCV.ConclusionThis discovery of a previously unclassified HCV subtype within genotype 4 sheds light on the ongoing evolution and diversity of the virus. Such knowledge has significant implications for diagnostic and therapeutic approaches, as different subtypes may exhibit varying drug sensitivities and resistance profiles.
The decay kinetics of HIV-1-infected cells are critical to understand virus persistence. We evaluated the frequency of simian immunodeficiency virus (SIV)-infected cells for 4 years of antiretroviral therapy (ART). The intact proviral DNA assay (IPDA) and an assay for hypermutated proviruses revealed short- and long-term infected cell dynamics in macaques starting ART ∼1 year after infection. Intact SIV genomes in circulating CD4+T cells showed triphasic decay with an initial phase slower than the decay of the plasma virus, a second phase faster than the second phase decay of intact HIV-1, and a stable third phase reached after 1.6–2.9 years. Hypermutated proviruses showed bi- or mono-phasic decay, reflecting different selective pressures. Viruses replicating at ART initiation had mutations conferring antibody escape. With time on ART, viruses with fewer mutations became more prominent, reflecting decay of variants replicating at ART initiation. Collectively, these findings confirm ART efficacy and indicate that cells enter the reservoir throughout untreated infection.
HIV can persist in a latent form as integrated DNA (provirus) in resting CD4 + T cells of infected individuals and as such is unaffected by antiretroviral therapy (ART). Despite being a major obstacle for eradication efforts, the genetic variation and timing of formation of this latent reservoir remains poorly understood. Previous studies on when virus is deposited in the latent reservoir have come to contradictory conclusions. To reexamine the genetic variation of HIV in CD4 + T cells during ART, we determined the divergence in envelope sequences collected from 10 SIV infected rhesus macaques. We found that the macaques displayed a biphasic decline of the viral divergence over time, where the first phase lasted for an average of 11.6 weeks (range 4-28 weeks). Motivated by recent observations that the HIV-infected CD4 + T cell population is composed of short- and long-lived subsets, we developed a model to study the divergence dynamics. We found that SIV in short-lived cells was on average more diverged, while long-lived cells harbored less diverged virus. This suggests that the long-lived cells harbor virus deposited starting earlier in infection and continuing throughout infection, while short-lived cells predominantly harbor more recent virus. As these cell populations decayed, the overall proviral divergence decline matched that observed in the empirical data. This model explains previous seemingly contradictory results on the timing of virus deposition into the latent reservoir, and should provide guidance for future eradication efforts. Significance statement HIV can persist in a latent reservoir unaffected by antiretroviral drugs. The genetic variation of this latent virus population is a major obstacle for eradication efforts, but also a clue to when HIV variants are deposited in the reservoirs. Unfortunately, previous studies assessing when the virus was deposited in latent reservoirs have come to contradictory conclusions. Here, we propose SIV proviral DNA exists in both short- and long-lived CD4 + T cells, and that these two cell subsets harbor different genetically diverged virus populations. Our model explains the contradictory findings and shows that when CD4 + T cells decay under effective drug treatment, which prevents virus replication, the resulting virus divergence decreases and recapitulates observed data. This knowledge should help in improving future eradication efforts.
Reassortment is an evolutionary process common in viruses with segmented genomes. These viruses can swap whole genomic segments during cellular co-infection, giving rise to new viral variants. Large-scale genome rearrangements, such as reassortment, have the potential to quickly generate new phenotypes, making the understanding of viral reassortment important to both evolutionary biology and public health research. In this paper, we argue that reassortment cannot be reliably inferred from incongruities between segment phylogenies using the established remove-and-rejoin or coalescent approaches. We instead show that reassortment must be considered in the context of a broader population process that includes the dynamics of the infected hosts. Using illustrative examples and simulation we identify four types of evolutionary events that are difficult or impossible to reconstruct with incongruence-based methods. Further, we show that these specific situations are very common and will likely occur even in small samples. Finally, we argue that existing methods can be augmented or modified to account for all the problematic situations that we identify in this paper. Robust assessment of the role of reassortment in viral evolution is difficult, and we hope to provide conceptual clarity on some important methodological issues that can arise in the development of the next generation of tools for studying reassortment.
ABSTRACT The evolution of HIV-1 in a host is shaped by many evolutionary forces, including recombination of virus genomes and the potential isolation of viruses into different tissues with compartmentalized evolution. Recombination and compartmentalization have opposite effects on viral diversification, with the former causing global mixing and the latter countering it through spatial segregation of the virus population. Therefore, recombination and compartmentalization together give rise to complex evolutionary dynamics that convolutes their individual effects in standard, bifurcating phylogenetic trees. Although there are various theoretical methods available to infer the presence of recombination or compartmentalization individually, there is little knowledge of their combined effect. To study their interaction and whether that could explain the many disparate results that have been described in the HIV-1 literature, we developed an age-structured forward-time evolutionary model that includes compartments, migration, and recombination. By tracking the evolutionary history of individual virus variants in an Ancestral Recombination Graph (ARG) and resolving the ARG into standard bifurcating trees, we reexamined 771 anatomical tissue pairs infected with HIV-1. Remarkably, we found that recombination can make the resulting bifurcating tree appear more compartmentalized than the virus population actually is. However, we found migration between all 771 tissue pairs, typically at 2.8-4.9×10 −3 taxa -1 day -1 . Thus, while different point mutations may arise in different parts of the body, migration eventually brings these variants together and recombination merges them into a relatively homogeneous cloud of a universally evolving quasispecies. Modelling this process, we explain the many different results previous research found among distinct anatomical tissues. We also show that the popular Slatkin-Maddison test comes to different results about compartmentalization at the same migration rate depending on the sample size. SIGNIFICANCE Whether HIV-1 sampled in different anatomical tissues are compartmentalized or not has been a long-standing quest because it relates to both clinical and biological insight. Previous studies of this question have reported disparate results from the same as well as different tissue comparisons. Here, we consolidate and explain these contrasting results with a new evolutionary model that includes compartments, migration, and recombination. We show that no anatomical tissues are fully compartmentalized, and that the migration of HIV-1 between the tissues allows recombination to homogenize the diversity that may arise in separate tissues of an infected person.
The study aimed to characterize the genotype and subgenotypes of HBV circulating in Saudi Arabia, the presence of clinically relevant mutations possibly associated with resistance to antivirals or immune escape phenomena, and the possible impact of mutations in the structural characteristics of HBV polymerase. Plasma samples from 12 Saudi Arabian HBV-infected patients were analyzed using an in-house PCR method and direct sequencing. Saudi patients were infected with mainly subgenotype D1. A number of mutations in the RT gene (correlated to antiviral resistance) and within and outside the major hydrophilic region of the S gene (claimed to influence immunogenicity and be related to immune escape) were observed in almost all patients. Furthermore, the presence of mutations in the S region caused a change in the tertiary structure of the protein compared with the consensus region. Clinical manifestations of HBV infection may change dramatically as a result of viral and host factors: the study of mutations and protein-associated cofactors might define possible aspects relevant for the natural and therapeutic history of HBV infection.
Within-host Human immunodeficiency virus (HIV) evolution involves several features that may disrupt standard phylogenetic reconstruction. One important feature is reactivation of latently integrated provirus, which has the potential to disrupt the temporal signal, leading to variation in the branch lengths and apparent evolutionary rates in a tree. Yet, real within-host HIV phylogenies tend to show clear, ladder-like trees structured by the time of sampling. Another important feature is recombination, which violates the fundamental assumption that evolutionary history can be represented by a single bifurcating tree. Thus, recombination complicates the within-host HIV dynamic by mixing genomes and creating evolutionary loop structures that cannot be represented in a bifurcating tree. In this paper, we develop a coalescent-based simulator of within-host HIV evolution that includes latency, recombination, and effective population size dynamics that allows us to study the relationship between the true, complex genealogy of within-host HIV evolution, encoded as an ancestral recombination graph (ARG), and the observed phylogenetic tree. To compare our ARG results to the familiar phylogeny format, we calculate the expected bifurcating tree after decomposing the ARG into all unique site trees, their combined distance matrix, and the overall corresponding bifurcating tree. While latency and recombination separately disrupt the phylogenetic signal, remarkably, we find that recombination recovers the temporal signal of within-host HIV evolution caused by latency by mixing fragments of old, latent genomes into the contemporary population. In effect, recombination averages over extant heterogeneity, whether it stems from mixed time signals or population bottlenecks. Furthermore, we establish that the signals of latency and recombination can be observed in phylogenetic trees despite being an incorrect representation of the true evolutionary history. Using an approximate Bayesian computation method, we develop a set of statistical probes to tune our simulation model to nine longitudinally sampled within-host HIV phylogenies. Because ARGs are exceedingly difficult to infer from real HIV data, our simulation system allows investigating effects of latency, recombination, and population size bottlenecks by matching decomposed ARGs to real data as observed in standard phylogenies.