Low-abundance members of microbial communities are difficult to study in their native habitats, including Escherichia coli, a minor but common inhabitant of the gastrointestinal tract, and key opportunistic pathogen of the urinary tract. While multi-omic analyses have detailed interactions between uropathogenic Escherichia coli (UPEC) and the bladder mediating urinary tract infection (UTI), little is known about UPEC in its pre-infection reservoir, the gastrointestinal tract, partly due to its low relative abundance (<1%). To sensitively explore the genomes and transcriptomes of diverse gut E. coli, we develop E. coli PanSelect, which uses probes designed to specifically capture E. coli's broad pangenome. We demonstrate its ability to enrich diverse E. coli by orders of magnitude, in a mock community and in human stool from a study investigating recurrent UTI (rUTI). Comparisons of transcriptomes between gut E. coli of women with and without history of rUTI suggest rUTI gut E. coli are responding to increased oxygen and nitrate, suggestive of mucosal inflammation, which may have implications for recurrent disease. E. coli PanSelect is well suited for investigations of in vivo E. coli biology in other low-abundance environments, and the framework described here has broad applicability to other diverse, low-abundance organisms.
Effective infectious disease surveillance in high-risk regions is critical for clinical care and pandemic preemption; however, few clinical diagnostics are available for the wide range of potential human pathogens. Here, we conduct unbiased metagenomic sequencing of 593 samples from febrile Nigerian patients collected in three settings: i) population-level surveillance of individuals presenting with symptoms consistent with Lassa Fever (LF); ii) real-time investigations of outbreaks with suspected infectious etiologies; and iii) undiagnosed clinically challenging cases. We identify 13 distinct viruses, including the second and third documented cases of human blood-associated dicistrovirus, and a highly divergent, unclassified dicistrovirus that we name human blood-associated dicistrovirus 2. We show that pegivirus C is a common co-infection in individuals with LF and is associated with lower Lassa viral loads and favorable outcomes. We help uncover the causes of three outbreaks as yellow fever virus, monkeypox virus, and a noninfectious cause, the latter ultimately determined to be pesticide poisoning. We demonstrate that a local, Nigerian-driven metagenomics response to complex public health scenarios generates accurate, real-time differential diagnoses, yielding insights that inform policy.
Background Carbapenem-resistant Enterobacterales (CRE) are an urgent global health threat. Inferring the dynamics of local CRE dissemination is currently limited by our inability to confidently trace the spread of resistance determinants to unrelated bacterial hosts. Whole-genome sequence comparison is useful for identifying CRE clonal transmission and outbreaks, but high-frequency horizontal gene transfer (HGT) of carbapenem resistance genes and subsequent genome rearrangement complicate tracing the local persistence and mobilization of these genes across organisms. Methods To overcome this limitation, we developed a new approach to identify recent HGT of large, near-identical plasmid segments across species boundaries, which also allowed us to overcome technical challenges with genome assembly. We applied this to complete and near-complete genome assemblies to examine the local spread of CRE in a systematic, prospective collection of all CRE, as well as time- and species-matched carbapenem-susceptible Enterobacterales , isolated from patients from four US hospitals over nearly 5 years. Results Our CRE collection comprised a diverse range of species, lineages, and carbapenem resistance mechanisms, many of which were encoded on a variety of promiscuous plasmid types. We found and quantified rearrangement, persistence, and repeated transfer of plasmid segments, including those harboring carbapenemases, between organisms over multiple years. Some plasmid segments were found to be strongly associated with specific locales, thus representing geographic signatures that make it possible to trace recent and localized HGT events. Functional analysis of these signatures revealed genes commonly found in plasmids of nosocomial pathogens, such as functions required for plasmid retention and spread, as well survival against a variety of antibiotic and antiseptics common to the hospital environment. Conclusions Collectively, the framework we developed provides a clearer, high-resolution picture of the epidemiology of antibiotic resistance importation, spread, and persistence in patients and healthcare networks.
Unusually large outbreaks of mumps across the United States in 2016 and 2017 raised questions about the extent of mumps circulation and the relationship between these and prior outbreaks. We paired epidemiological data from public health investigations with analysis of mumps virus whole genome sequences from 201 infected individuals, focusing on Massachusetts university communities. Our analysis suggests continuous, undetected circulation of mumps locally and nationally, including multiple independent introductions into Massachusetts and into individual communities. Despite the presence of these multiple mumps virus lineages, the genomic data show that one lineage has dominated in the US since at least 2006. Widespread transmission was surprising given high vaccination rates, but we found no genetic evidence that variants arising during this outbreak contributed to vaccine escape. Viral genomic data allowed us to reconstruct mumps transmission links not evident from epidemiological data or standard single-gene surveillance efforts and also revealed connections between apparently unrelated mumps outbreaks.
Metagenomic sequencing has the potential to transform microbial detection and characterization, but new tools are needed to improve its sensitivity. Here we present CATCH, a computational method to enhance nucleic acid capture for enrichment of diverse microbial taxa. CATCH designs optimal probe sets, with a specified number of oligonucleotides, that achieve full coverage of, and scale well with, known sequence diversity. We focus on applying CATCH to capture viral genomes in complex metagenomic samples. We design, synthesize, and validate multiple probe sets, including one that targets the whole genomes of the 356 viral species known to infect humans. Capture with these probe sets enriches unique viral content on average 18-fold, allowing us to assemble genomes that could not be recovered without enrichment, and accurately preserves within-sample diversity. We also use these probe sets to recover genomes from the 2018 Lassa fever outbreak in Nigeria and to improve detection of uncharacterized viral infections in human and mosquito samples. The results demonstrate that CATCH enables more sensitive and cost-effective metagenomic sequencing.
During 2018, an unusual increase in Lassa fever cases occurred in Nigeria, raising concern among national and international public health agencies. We analyzed 220 Lassa virus genomes from infected patients, including 129 from the 2017-2018 transmission season, to understand the viral populations underpinning the increase. A total of 14 initial genomes from 2018 samples were generated at Redeemer's University in Nigeria, and the findings were shared with the Nigerian Center for Disease Control in real time. We found that the increase in cases was not attributable to a particular Lassa virus strain or sustained by human-to-human transmission. Instead, the data were consistent with ongoing cross-species transmission from local rodent populations. Phylogenetic analysis also revealed extensive viral diversity that was structured according to geography, with major rivers appearing to act as barriers to migration of the rodent reservoir.
In one patient over time, we found that concentration of Ebola virus RNA in semen during recovery is remarkably higher than blood at peak illness. Virus in semen is replication-competent with no change in viral genome over time. Presence of sense RNA suggests replication in cells present in semen.
The 2013–2016 West African epidemic caused by the Ebola virus was of unprecedented magnitude, duration and impact. Here we reconstruct the dispersal, proliferation and decline of Ebola virus throughout the region by analysing 1,610 Ebola virus genomes, which represent over 5% of the known cases. We test the association of geography, climate and demography with viral movement among administrative regions, inferring a classic ‘gravity’ model, with intense dispersal between larger and closer populations. Despite attenuation of international dispersal after border closures, cross-border transmission had already sown the seeds for an international epidemic, rendering these measures ineffective at curbing the epidemic. We address why the epidemic did not spread into neighbouring countries, showing that these countries were susceptible to substantial outbreaks but at lower risk of introductions. Finally, we reveal that this large epidemic was a heterogeneous and spatially dissociated collection of transmission clusters of varying size, duration and connectivity. These insights will help to inform interventions in future epidemics.
Zika virus (ZIKV) is causing an unprecedented epidemic linked to severe congenital syndromes 1,2 . In July 2016, mosquito-borne ZIKV transmission was first reported in the continental United States and since then, hundreds of locally-acquired infections have been reported in Florida 3 . To gain insights into the timing, source, and likely route(s) of introduction of ZIKV into the continental United States, we tracked the virus from its first detection in Miami, Florida by direct sequencing of ZIKV genomes from infected patients and Aedes aegypti mosquitoes. We show that at least four distinct ZIKV introductions contributed to the outbreak in Florida and that local transmission likely started in the spring of 2016 - several months before its initial detection. By analyzing surveillance and genetic data, we discovered that ZIKV moved among transmission zones in Miami. Our analyses show that most introductions are phylogenetically linked to the Caribbean, a finding corroborated by the high incidence rates and traffic volumes from the region into the Miami area. By comparing mosquito abundance and travel flows, we describe the areas of southern Florida that are especially vulnerable to ZIKV introductions. Our study provides a deeper understanding of how ZIKV initiates and sustains transmission in new regions.
One hundred and ten Zika virus genomes from ten countries and territories involved in the Zika virus epidemic reveal rapid expansion of the epidemic within Brazil and multiple introductions to other regions. Three papers in this issue present a wealth of new Zika virus (ZIKV) genome sequences and further insights into the genetic epidemiology of ZIKV. Nathan Grubaugh et al. provide 39 new ZIKV genome sequences from infected patients and Aedes aegypti mosquitoes in Florida. Phylogenetic analysis suggests that the virus has been introduced on multiple separate occasions, probably linked to travel from the Caribbean. They find a low probability of long-term persistence of ZIKV transmission chains within Florida, suggesting that the potential for future ZIKV outbreaks there will depend on transmission dynamics in the Americas. Nuno Faria et al. and Hayden Metsky et al. reconstruct the spread of ZIKV in Brazil and the Americas. Faria et al. provide 54 new ZIKV genomes, several sequenced in real time in a mobile genomics laboratory. They trace the spatial origins and spread of ZIKV in Brazil and the Americas and date the timing of the international spread of ZIKV from Brazil. They find that northeast Brazil had a crucial role in the establishment of the epidemic and the spread of the virus within Brazil and the Americas. Metsky et al. generate 110 ZIKV genomes from clinical and mosquito samples from ten regions. They also see rapid expansion of the epidemic within Brazil and multiple introductions to other geographic areas. In agreement with Faria et al., they find that ZIKV circulated unobserved for many months before transmission was detected. Metsky et al. additionally describe ZIKV evolution and discuss how the accumulation of mutations might affect the performance of diagnostic tests in the future. Although the recent Zika virus (ZIKV) epidemic in the Americas and its link to birth defects have attracted a great deal of attention1,2, much remains unknown about ZIKV disease epidemiology and ZIKV evolution, in part owing to a lack of genomic data. Here we address this gap in knowledge by using multiple sequencing approaches to generate 110 ZIKV genomes from clinical and mosquito samples from 10 countries and territories, greatly expanding the observed viral genetic diversity from this outbreak. We analysed the timing and patterns of introductions into distinct geographic regions; our phylogenetic evidence suggests rapid expansion of the outbreak in Brazil and multiple introductions of outbreak strains into Puerto Rico, Honduras, Colombia, other Caribbean islands, and the continental United States. We find that ZIKV circulated undetected in multiple regions for many months before the first locally transmitted cases were confirmed, highlighting the importance of surveillance of viral infections. We identify mutations with possible functional implications for ZIKV biology and pathogenesis, as well as those that might be relevant to the effectiveness of diagnostic tests.
Here we outline a next-generation RNA sequencing protocol that enables de novo assemblies and intra-host variant calls of viral genomes collected from clinical and biological sources. The method is unbiased and universal; it uses random primers for cDNA synthesis and requires no prior knowledge of the viral sequence content. Before library construction, selective RNase H-based digestion is used to deplete unwanted RNA - including poly(rA) carrier and ribosomal RNA - from the viral RNA sample. Selective depletion improves both the data quality and the number of unique reads in viral RNA sequencing libraries. Moreover, a transposase-based 'tagmentation' step is used in the protocol as it reduces overall library construction time. The protocol has enabled rapid deep sequencing of over 600 Lassa and Ebola virus samples-including collections from both blood and tissue isolates-and is broadly applicable to other microbial genomics studies.
The 2013-2015 Ebola virus disease (EVD) epidemic is caused by the Makona variant of Ebola virus (EBOV). Early in the epidemic, genome sequencing provided insights into virus evolution and transmission and offered important information for outbreak response. Here, we analyze sequences from 232 patients sampled over 7 months in Sierra Leone, along with 86 previously released genomes from earlier in the epidemic. We confirm sustained human-to-human transmission within Sierra Leone and find no evidence for import or export of EBOV across national borders after its initial introduction. Using high-depth replicate sequencing, we observe both host-to-host transmission and recurrent emergence of intrahost genetic variants. We trace the increasing impact of purifying selection in suppressing the accumulation of nonsynonymous mutations over time. Finally, we note changes in the mucin-like domain of EBOV glycoprotein that merit further investigation. These findings clarify the movement of EBOV within the region and describe viral evolution during prolonged human-to-human transmission.
ABSTRACT The introduction of West Nile virus (WNV) into North America in 1999 is a classic example of viral emergence in a new environment, with its subsequent dispersion across the continent having a major impact on local bird populations. Despite the importance of this epizootic, the pattern, dynamics, and determinants of WNV spread in its natural hosts remain uncertain. In particular, it is unclear whether the virus encountered major barriers to transmission, or spread in an unconstrained manner, and if specific viral lineages were favored over others indicative of intrinsic differences in fitness. To address these key questions in WNV evolution and ecology, we sequenced the complete genomes of approximately 300 avian isolates sampled across the United States between 2001 and 2012. Phylogenetic analysis revealed a relatively star-like tree structure, indicative of explosive viral spread in the United States, although with some replacement of viral genotypes through time. These data are striking in that viral sequences exhibit relatively limited clustering according to geographic region, particularly for those viruses sampled from birds, and no strong phylogenetic association with well-sampled avian species. The genome sequence data analyzed here also contain relatively little evidence for adaptive evolution, particularly of structural proteins, suggesting that most viral lineages are of similar fitness and that WNV is well adapted to the ecology of mosquito vectors and diverse avian hosts in the United States. In sum, the molecular evolution of WNV in North America depicts a largely unfettered expansion within a permissive host and geographic population with little evidence of major adaptive barriers. IMPORTANCE How viruses spread in new host and geographic environments is central to understanding the emergence and evolution of novel infectious diseases and for predicting their likely impact. The emergence of the vector-borne West Nile virus (WNV) in North America in 1999 represents a classic example of this process. Using approximately 300 new viral genomes sampled from wild birds, we show that WNV experienced an explosive spread with little geographical or host constraints within birds and relatively low levels of adaptive evolution. From its introduction into the state of New York, WNV spread across the United States, reaching California and Florida within 4 years, a migration that is clearly reflected in our genomic sequence data, and with a general absence of distinct geographical clusters of bird viruses. However, some geographically distinct viral lineages were found to circulate in mosquitoes, likely reflecting their limited long-distance movement compared to avian species.
The 2013-2015 West African epidemic of Ebola virus disease (EVD) reminds us of how little is known about biosafety level 4 viruses. Like Ebola virus, Lassa virus (LASV) can cause hemorrhagic fever with high case fatality rates. We generated a genomic catalog of almost 200 LASV sequences from clinical and rodent reservoir samples. We show that whereas the 2013-2015 EVD epidemic is fueled by human-to-human transmissions, LASV infections mainly result from reservoir-to-human infections. We elucidated the spread of LASV across West Africa and show that this migration was accompanied by changes in LASV genome abundance, fatality rates, codon adaptation, and translational efficiency. By investigating intrahost evolution, we found that mutations accumulate in epitopes of viral surface proteins, suggesting selection for immune escape. This catalog will serve as a foundation for the development of vaccines and diagnostics. VIDEO ABSTRACT.
In its largest outbreak, Ebola virus disease is spreading through Guinea, Liberia, Sierra Leone, and Nigeria. We sequenced 99 Ebola virus genomes from 78 patients in Sierra Leone to ~2000× coverage. We observed a rapid accumulation of interhost and intrahost genetic variation, allowing us to characterize patterns of viral transmission over the initial weeks of the epidemic. This West African variant likely diverged from central African lineages around 2004, crossed from Guinea to Sierra Leone in May 2014, and has exhibited sustained human-to-human transmission subsequently, with no evidence of additional zoonotic sources. Because many of the mutations alter protein sequences and other biologically meaningful targets, they should be monitored for impact on diagnostics, vaccines, and therapies critical to outbreak response.
ABSTRACT Human respiratory syncytial virus (RSV) is the leading cause of lower respiratory tract disease in infants and young children and an important respiratory pathogen in the elderly and immunocompromised. While population-wide molecular epidemiology studies have shown multiple cocirculating RSV genotypes and revealed antigenic and genetic change over successive seasons, little is known about the extent of viral diversity over the course of an individual infection, the origins of novel variants, or the effect of immune pressure on viral diversity and potential immune-escape mutations. To investigate viral population diversity in the presence and absence of selective immune pressures, we studied whole-genome deep sequencing of RSV in upper airway samples from an infant with severe combined immune deficiency syndrome and persistent RSV infection. The infection continued over several months before and after bone marrow transplant (BMT) from his RSV-immune father. RSV diversity was characterized in 26 samples obtained over 78 days. Diversity increased after engraftment, as defined by T-cell presence, and populations reflected variation mostly within the G protein, the major surface antigen. Minority populations with known palivizumab resistance mutations emerged after its administration. The viral population appeared to diversify in response to selective pressures, showing a statistically significant growth in diversity in the presence of pressure from immunity. Defining escape mutations and their dynamics will be useful in the design and application of novel therapeutics and vaccines. These data can contribute to future studies of the relationship between within-host and population-wide RSV phylodynamics. IMPORTANCE Human RSV is an important cause of respiratory disease in infants, the elderly, and the immunocompromised. RSV circulating in a community appears to change season by season, but the amount of diversity generated during an individual infection and the impact of immunity on this viral diversity has been unclear. To address this question, we described within-host RSV diversity by whole-genome deep sequencing in a unique clinical case of an RSV-infected infant with severe combined immunodeficiency and effectively no adaptive immunity who then gained adaptive immunity after undergoing bone marrow transplantation. We found that viral diversity increased in the presence of adaptive immunity and was primarily within the G protein, the major surface antigen. These data will be useful in designing RSV treatments and vaccines and to help understand the relationship between the dynamics of viral diversification within individual hosts and the viral populations circulating in a community.
RNA viruses are the causative agents for AIDS, influenza, SARS, and other serious health threats. Development of rapid and broadly applicable methods for complete viral genome sequencing is highly desirable to fully understand all aspects of these infectious agents as well as for surveillance of viral pandemic threats and emerging pathogens. However, traditional viral detection methods rely on prior sequence or antigen knowledge. In this study, we describe sequence-independent amplification for samples containing ultra-low amounts of viral RNA coupled with Illumina sequencing and de novo assembly optimized for viral genomes. With 5 million reads, we capture 96 to 100% of the viral protein coding region of HIV, respiratory syncytial and West Nile viral samples from as little as 100 copies of viral RNA. The methods presented here are scalable to large numbers of samples and capable of generating full or near full length viral genomes from clone and clinical samples with low amounts of viral RNA, without prior sequence information and in the presence of substantial host contamination.
BACKGROUND:Extensive genetic diversity in viral populations within infected hosts and the divergence of variants from existing reference genomes impede the analysis of deep viral sequencing data. A de novo population consensus assembly is valuable both as a single linear representation of the population and as a backbone on which intra-host variants can be accurately mapped. The availability of consensus assemblies and robustly mapped variants are crucial to the genetic study of viral disease progression, transmission dynamics, and viral evolution. Existing de novo assembly techniques fail to robustly assemble ultra-deep sequence data from genetically heterogeneous populations such as viruses into full-length genomes due to the presence of extensive genetic variability, contaminants, and variable sequence coverage.RESULTS:We present VICUNA, a de novo assembly algorithm suitable for generating consensus assemblies from genetically heterogeneous populations. We demonstrate its effectiveness on Dengue, Human Immunodeficiency and West Nile viral populations, representing a range of intra-host diversity. Compared to state-of-the-art assemblers designed for haploid or diploid systems, VICUNA recovers full-length consensus and captures insertion/deletion polymorphisms in diverse samples. Final assemblies maintain a high base calling accuracy. VICUNA program is publicly available at: http://www.broadinstitute.org/scientific-community/science/projects/viral-genomics/ viral-genomics-analysis-software.CONCLUSIONS:We developed VICUNA, a publicly available software tool, that enables consensus assembly of ultra-deep sequence derived from diverse viral populations. While VICUNA was developed for the analysis of viral populations, its application to other heterogeneous sequence data sets such as metagenomic or tumor cell population samples may prove beneficial in these fields of research.