Metagenomics enables direct investigation of the gene content and potential functions of gut bacteria without isolation and culture. However, metagenome-assembled genomes are often incomplete and have low contiguity due to challenges in assembling repeated genomic elements. Long-read sequencing approaches have successfully yielded circular bacterial genomes directly from metagenomes, but these approaches require high DNA input and can have high error rates. Illumina has recently launched the Illumina Complete Long Read (ICLR) assay, a new approach for generating kilobase-scale reads with low DNA input requirements and high accuracy. Here, we evaluate the performance of ICLR sequencing for gut metagenomics for the first time. We sequenced a microbial mock community and 10 human gut microbiome samples with standard, shotgun 2 × 150 paired-end sequencing, ICLR sequencing, and nanopore long-read sequencing and compared performance in read lengths, assembly contiguity, and bin quality. We find that ICLR human metagenomic assemblies have higher N50 (119.5 ± 24.8 kilobases) than short read assemblies (9.9 ± 4.5 kilobases; P = 0.002), and comparable N50 to nanopore assemblies (91.0 ± 43.8 kilobases; P = 0.32). Additionally, we find that ICLR draft microbial genomes are more complete (94.0% ± 20.6%) than nanopore draft genomes (85.9% ± 23.0%; P ≤ 0.001), and that nanopore draft genomes have truncated gene lengths (924.6 ± 114.7 base pairs) relative to ICLR genomes (954.6 ± 71.5 base pairs; P ≤ 0.001). Overall, we find that ICLR sequencing is a promising method for the accurate assembly of microbial genomes from gut metagenomes.IMPORTANCEMetagenomic sequencing allows scientists to directly measure the genome content and structure of microbes residing in complex microbial communities. Traditional short-read metagenomic sequencing methods often yield fragmented genomes, whereas advanced long-read sequencing methods improve genome assembly quality but often suffer from high error rates and are logistically limited due to high input requirements. A new method, the Illumina Complete Long Read (ICLR) assay, is capable of generating highly accurate kilobase-scale sequencing reads with minimal input material. To evaluate the utility of ICLR in metagenomic contexts, we applied short-read, long-read, and ICLR methods to simple and complex microbial communities. We found that ICLR outperforms short-read methods and yields comparable metagenomic assemblies to standard long-read approaches while requiring less input material. Overall, ICLR represents an additional option for assembling complete genomes from complex metagenomes.
Five bacterial isolates were isolated from Fragaria × ananassa in 1976 in Rydalmere, Australia, during routine biosecurity surveillance. Initially, the results of biochemical characterisation indicated that these isolates represented members of the genus Xanthomonas. To determine their species, further analysis was conducted using both phenotypic and genotypic approaches. Phenotypic analysis involved using MALDI-TOF MS and BIOLOG GEN III microplates, which confirmed that the isolates represented members of the genus Xanthomonas but did not allow them to be classified with respect to species. Genome relatedness indices and the results of extensive phylogenetic analysis confirmed that the isolates were members of the genus Xanthomonas and represented a novel species. On the basis the minimal presence of virulence-associated factors typically found in genomes of members of the genus Xanthomonas, we suggest that these isolates are non-pathogenic. This conclusion was supported by the results of a pathogenicity assay. On the basis of these findings, we propose the name Xanthomonas rydalmerensis, with DAR 34855T = ICMP 24941 as the type strain.
ABSTRACTWe describe five bacterial isolates that were isolated fromFragaria x ananassain 1976 in Rydalmere, Australia, during routine biosecurity surveillance. Initially, biochemical characterisation identified these isolates as members of theXanthomonasgenus. To determine their species, we conducted further analysis using both phenotypic and genotypic approaches. Phenotypic analysis involved using MALDI-TOF MS and BIOLOG GEN III microplates, which confirmed that the isolates belonged to theXanthomonasgenus but could not classify species. Genome relatedness indices and extensive phylogenetic analysis confirmed that the isolates belonged to theXanthomonasgenus and represented a new species. Based on the absence of virulence factors typically found inXanthomonasspp. genomes, we suggest that these isolates are non-pathogenic. This conclusion was supported by a pathogenicity assay. Based on these findings, we propose the nameXanthomonas rydalmerenesis, with DAR34855 = ICMP24941 as the type strain.
Bayesian inference for phylogenetics is a gold standard for computing distributions of phylogenies. However, Bayesian phylogenetics faces the challenging computational problem of moving throughout the high-dimensional space of trees. Fortunately, hyperbolic space offers a low dimensional representation of tree-like data. In this paper, we embed genomic sequences as points in hyperbolic space and perform hyperbolic Markov Chain Monte Carlo for Bayesian inference in this space. The posterior probability of an embedding is computed by decoding a neighbour-joining tree from the embedding locations of the sequences. We empirically demonstrate the fidelity of this method on eight data sets. We systematically investigated the effect of embedding dimension and hyperbolic curvature on the performance in these data sets. The sampled posterior distribution recovers the splits and branch lengths to a high degree over a range of curvatures and dimensions. We systematically investigated the effects of the embedding space's curvature and dimension on the Markov Chain's performance, demonstrating the suitability of hyperbolic space for phylogenetic inference.
Recombination is a key evolutionary driver in shaping novel viral populations and lineages. When unaccounted for, recombination can impact evolutionary estimations or complicate their interpretation. Therefore, identifying signals for recombination in sequencing data is a key prerequisite to further analyses. A repertoire of recombination detection methods (RDMs) have been developed over the past two decades; however, the prevalence of pandemic-scale viral sequencing data poses a computational challenge for existing methods. Here, we assessed eight RDMs: PhiPack (Profile), 3SEQ, GENECONV, recombination detection program (RDP) (OpenRDP), MaxChi (OpenRDP), Chimaera (OpenRDP), UCHIME (VSEARCH), and gmos; to determine if any are suitable for the analysis of bulk sequencing data. To test the performance and scalability of these methods, we analysed simulated viral sequencing data across a range of sequence diversities, recombination frequencies, and sample sizes. Furthermore, we provide a practical example for the analysis and validation of empirical data. We find that RDMs need to be scalable, use an analytical approach and resolution that is suitable for the intended research application, and are accurate for the properties of a given dataset (e.g. sequence diversity and estimated recombination frequency). Analysis of simulated and empirical data revealed that the assessed methods exhibited considerable trade-offs between these criteria. Overall, we provide general guidelines for the validation of recombination detection results, the benefits and shortcomings of each assessed method, and future considerations for recombination detection methods for the assessment of large-scale viral sequencing data.
We describe five bacterial isolates that were isolated from Fragaria x ananassa in 1976 in Rydalmere, Australia, during routine biosecurity surveillance. Initially, biochemical characterisation identified these isolates as members of the Xanthomonas genus. To determine their species, we conducted further analysis using both phenotypic and genotypic approaches. Phenotypic analysis involved using MALDI-TOF MS and BIOLOG GEN III microplates, which confirmed that the isolates belonged to the Xanthomonas genus but could not classify species. Genome relatedness indices and extensive phylogenetic analysis confirmed that the isolates belonged to the Xanthomonas genus and represented a new species. Based on the absence of virulence factors typically found in Xanthomonas spp. genomes, we suggest that these isolates are non-pathogenic. This conclusion was supported by a pathogenicity assay. Based on these findings, we propose the name Xanthomonas rydalmerenesis , with DAR34855 = ICMP24941 as the type strain.### Competing Interest StatementThe authors have declared no competing interest.
Prokaryotic evolution is influenced by the exchange of genetic information between species through a process referred to as recombination. The rate of recombination is a useful measure for the adaptive capacity of a prokaryotic population. We introduce Rhometa (https://github.com/sid-krish/Rhometa), a new software package to determine recombination rates from shotgun sequencing reads of metagenomes. It extends the composite likelihood approach for population recombination rate estimation and enables the analysis of modern short-read datasets. We evaluated Rhometa over a broad range of sequencing depths and complexities, using simulated and real experimental short-read data aligned to external reference genomes. Rhometa offers a comprehensive solution for determining population recombination rates from contemporary metagenomic read datasets. Rhometa extends the capabilities of conventional sequence-based composite likelihood population recombination rate estimators to include modern aligned metagenomic read datasets with diverse sequencing depths, thereby enabling the effective application of these techniques and their high accuracy rates to the field of metagenomics. Using simulated datasets, we show that our method performs well, with its accuracy improving with increasing numbers of genomes. Rhometa was validated on a real S. pneumoniae transformation experiment, where we show that it obtains plausible estimates of the rate of recombination. Finally, the program was also run on ocean surface water metagenomic datasets, through which we demonstrate that the program works on uncultured metagenomic datasets.
Bovine respiratory disease (BRD) is one of the most common diseases in intensively managed cattle, often resulting in high morbidity and mortality. Although several pathogens have been isolated and extensively studied, the complete infectome of the respiratory complex consists of a more extensive range unrecognised species. Here, we used total RNA sequencing (i.e., metatranscriptomics) of nasal and nasopharyngeal swabs collected from animals with and without BRD from two cattle feedlots in Australia. A high abundance of bovine nidovirus, influenza D, bovine rhinitis A and bovine coronavirus was found in the samples. Additionally, we obtained the complete or near-complete genome of bovine rhinitis B, enterovirus E1, bovine viral diarrhea virus (sub-genotypes 1a and 1c) and bovine respiratory syncytial virus, and partial sequences of other viruses. A new species of paramyxovirus was also identified. Overall, the most abundant RNA virus, was the bovine nidovirus. Characterisation of bacterial species from the transcriptome revealed a high abundance and diversity of Mollicutes in BRD cases and unaffected control animals. Of the non-Mollicutes species, Histophilus somni was detected, whereas there was a low abundance of Mannheimia haemolytica. This study highlights the use of untargeted sequencing approaches to study the unrecognised range of microorganisms present in healthy or diseased animals and the need to study previously uncultured viral species that may have an important role in cattle respiratory disease.
Intensive farming practices can increase exposure of animals to infectious agents against which antibiotics are used. Orally administered antibiotics are well known to cause dysbiosis. To counteract dysbiotic effects, numerous studies in the past two decades sought to understand whether probiotics are a valid tool to help re-establish a healthy gut microbial community after antibiotic treatment. Although dysbiotic effects of antibiotics are well investigated, little is known about the effects of intramuscular antibiotic treatment on the gut microbiome and a few studies attempted to study treatment effects using phylogenetic diversity analysis techniques. In this study we sought to determine the effects of two probiotic- and one intramuscularly administered antibiotic treatment on the developing gut microbiome of post-weaning piglets between their 3rd and 9th week of life. Shotgun metagenomic sequences from over 800 faecal time-series samples derived from 126 post-weaning piglets and 42 sows were analysed in a phylogenetic framework. Differences between individual hosts such as breed, litter, and age, were found to be important contributors to variation in the community composition. Host age was the dominant factor in shaping the gut microbiota of piglets after weaning. The post-weaning pig gut microbiome appeared to follow a highly structured developmental program with characteristic post-weaning changes that can distinguish hosts that were born as little as two days apart in the second month of life. Treatment effects of the antibiotic and probiotic treatments were found but were subtle and included a higher representation of Mollicutes associated with intramuscular antibiotic treatment, and an increase of Lactobacillus associated with probiotic treatment. The discovery of correlations between experimental factors and microbial community composition is more commonly addressed with OTU-based methods and rarely analysed via phylogenetic diversity measures. The latter method, although less intuitive than the former, suffers less from library size normalization biases, and it proved to be instrumental in this study for the discovery of correlations between microbiome composition and host-, and treatment factors.
Evaluating metagenomic software is key for optimizing metagenome interpretation and focus of the Initiative for the Critical Assessment of Metagenome Interpretation (CAMI). The CAMI II challenge engaged the community to assess methods on realistic and complex datasets with long- and short-read sequences, created computationally from around 1,700 new and known genomes, as well as 600 new plasmids and viruses. Here we analyze 5,002 results by 76 program versions. Substantial improvements were seen in assembly, some due to long-read data. Related strains still were challenging for assembly and genome recovery through binning, as was assembly quality for the latter. Profilers markedly matured, with taxon profilers and binners excelling at higher bacterial ranks, but underperforming for viruses and Archaea. Clinical pathogen detection results revealed a need to improve reproducibility. Runtime and memory usage analyses identified efficient programs, including top performers with other metrics. The results identify challenges and guide researchers in selecting methods for analyses.
We developed a low-cost method for the production of Illumina-compatible sequencing libraries that allows up to 14 times more libraries for high-throughput Illumina sequencing to be generated for the same cost. We call this new method Hackflex. The quality of library preparation was tested by constructing libraries from Escherichia coli MG1655 genomic DNA using either Hackflex, standard Nextera Flex (recently renamed as Illumina DNA Prep) or a variation of standard Nextera Flex in which the bead-linked transposase is diluted prior to use. In order to test the library quality for genomes with a higher and a lower G+C content, library construction methods were also tested on Pseudomonas aeruginosa PAO1 and Staphylococcus aureus ATCC 25923, respectively. We demonstrated that Hackflex can produce high-quality libraries and yields a highly uniform coverage, equivalent to the standard Nextera Flex kit. We show that strongly size-selected libraries produce sufficient yield and complexity to support de novo microbial genome assembly, and that assemblies of the large-insert libraries can be much more contiguous than standard libraries without strong size selection. We introduce a new set of sample barcodes that are distinct from standard Illumina barcodes, enabling Hackflex samples to be multiplexed with samples barcoded using standard Illumina kits. Using Hackflex, we were able to achieve a per-sample reagent cost for library prep of A$7.22 (Australian dollars) (US $5.60; UK £3.87, £1=A$1.87), which is 9.87 times lower than the standard Nextera Flex protocol at advertised retail price. An additional simple modification and further simplification of the protocol by omitting the wash step enables a further price reduction to reach an overall 14-fold cost saving. This method will allow researchers to construct more libraries within a given budget, thereby yielding more data and facilitating research programmes where sequencing large numbers of libraries is beneficial.
Hi-C is a sample preparation method that enables high-throughput sequencing to capture genome-wide spatial interactions between DNA molecules. The technique has been successfully applied to solve challenging problems such as 3D structural analysis of chromatin, scaffolding of large genome assemblies and more recently the accurate resolution of metagenome-assembled genomes (MAGs). Despite continued refinements, however, preparing a Hi-C library remains a complex laboratory protocol. To avoid costly failures and maximise the odds of successful outcomes, diligent quality management is recommended. Current wet-lab methods provide only a crude assay of Hi-C library quality, while key post-sequencing quality indicators used have-thus far-relied upon reference-based read-mapping. When a reference is accessible, this reliance introduces a concern for quality, where an incomplete or inexact reference skews the resulting quality indicators. We propose a new, reference-free approach that infers the total fraction of read-pairs that are a product of proximity ligation. This quantification of Hi-C library quality requires only a modest amount of sequencing data and is independent of other application-specific criteria. The algorithm builds upon the observation that proximity ligation events are likely to create k-mers that would not naturally occur in the sample. Our software tool (qc3C) is to our knowledge the first to implement a reference-free Hi-C QC tool, and also provides reference-based QC, enabling Hi-C to be more easily applied to non-model organisms and environmental samples. We characterise the accuracy of the new algorithm on simulated and real datasets and compare it to reference-based methods.
ABSTRACT We developed a low-cost method for the production of Illumina-compatible sequencing libraries that allows up to 14 times more libraries for high-throughput Illumina sequencing to be generated for the same cost. We call this new method Hackflex. Quality of library preparation was tested by constructing libraries from E. coli MG1655 genomic DNA using either Hackflex, standard Nextera Flex or a variation of standard Nextera Flex in which the bead-linked transposase is diluted prior to use. In order to test the library quality for genomes with a higher and a lower GC content, library construction methods were also tested on P. aeruginosa PAO1 and S. aureus ATCC25923, respectively. We demonstrated that Hackflex can produce high quality libraries and yields a highly uniform coverage, equivalent to the standard Nextera Flex kit. We show that strongly size selected libraries produce sufficient yield and complexity to support de novo microbial genome assembly, and that assemblies of the large insert libraries can be much more contiguous than standard libraries without strong size selection. We introduce a new set of sample barcodes that are distinct from standard Illumina barcodes, enabling Hackflex samples to be multiplexed with samples barcoded using standard Illumina kits. Using Hackflex, we were able to achieve a per sample reagent cost for library prep of A$7.22 (USD$5.60), which is 9.87 times lower than the Standard Nextera Flex protocol at advertised retail price. An additional simple modification and further simplification of the protocol by omitting the wash step enables a further price reduction to reach an overall 14-fold cost saving. This method will allow researchers to construct more libraries within a given budget, thereby yielding more data and facilitating research programs where sequencing large numbers of libraries is beneficial.
This is the repository of configuration details and source code necessary to reproduce the simulated sweep for the manuscript: qc3C - reference-free quality control for Hi-C sequencing data
BACKGROUND:Human milk oligosaccharides (HMO) are a diverse range of sugars secreted in breast milk that have direct and indirect effects on immunity. The profiles of HMOs produced differ between mothers.OBJECTIVE:We sought to determine the relationship between maternal HMO profiles and offspring allergic diseases up to age 18 years.METHODS:Colostrum and early lactation milk samples were collected from 285 mothers enrolled in a high-allergy-risk birth cohort, the Melbourne Atopy Cohort Study. Nineteen HMOs were measured. Profiles/patterns of maternal HMOs were determined using LCA. Details of allergic disease outcomes including sensitization, wheeze, asthma, and eczema were collected at multiple follow-ups up to age 18 years. Adjusted logistic regression analyses and generalized estimating equations were used to determine the relationship between HMO profiles and allergy.RESULTS:The levels of several HMOs were highly correlated with each other. LCA determined 7 distinct maternal milk profiles with memberships of 10% and 20%. Compared with offspring exposed to the neutral Lewis HMO profile, exposure to acidic Lewis HMOs was associated with a higher risk of allergic disease and asthma over childhood (odds ratio asthma at 18 years, 5.82; 95% CI, 1.59-21.23), whereas exposure to the acidic-predominant profile was associated with a reduced risk of food sensitization (OR at 12 years, 0.08; 95% CI, 0.01-0.67).CONCLUSIONS:In this high-allergy-risk birth cohort, some profiles of HMOs were associated with increased and some with decreased allergic disease risks over childhood. Further studies are needed to confirm these findings and realize the potential for intervention.
We introduce STrain Resolution ON assembly Graphs (STRONG), which identifies strains de novo, from multiple metagenome samples. STRONG performs coassembly, and binning into metagenome assembled genomes (MAGs), and stores the coassembly graph prior to variant simplification. This enables the subgraphs and their unitig per-sample coverages, for individual single-copy core genes (SCGs) in each MAG, to be extracted. A Bayesian algorithm, BayesPaths, determines the number of strains present, their haplotypes or sequences on the SCGs, and abundances. STRONG is validated using synthetic communities and for a real anaerobic digestor time series generates haplotypes that match those observed from long Nanopore reads.
Computational methods are key in microbiome research, and obtaining a quantitative and unbiased performance estimate is important for method developers and applied researchers. For meaningful comparisons between methods, to identify best practices and common use cases, and to reduce overhead in benchmarking, it is necessary to have standardized datasets, procedures and metrics for evaluation. In this tutorial, we describe emerging standards in computational meta-omics benchmarking derived and agreed upon by a larger community of researchers. Specifically, we outline recent efforts by the Critical Assessment of Metagenome Interpretation (CAMI) initiative, which supplies method developers and applied researchers with exhaustive quantitative data about software performance in realistic scenarios and organizes community-driven benchmarking challenges. We explain the most relevant evaluation metrics for assessing metagenome assembly, binning and profiling results, and provide step-by-step instructions on how to generate them. The instructions use simulated mouse gut metagenome data released in preparation for the second round of CAMI challenges and showcase the use of a repository of tool results for CAMI datasets. This tutorial will serve as a reference for the community and facilitate informative and reproducible benchmarking in microbiome research.
Background Early weaning and intensive farming practices predispose piglets to the development of infectious and often lethal diseases, against which antibiotics are used. Besides contributing to the build-up of antimicrobial resistance, antibiotics are known to modulate the gut microbial composition. As an alternative to antibiotic treatment, studies have previously investigated the potential of probiotics for the prevention of postweaning diarrhea. In order to describe the post-weaning gut microbiota, and to study the effects of two probiotics formulations and of intramuscular antibiotic treatment on the gut microbiota, we sampled and processed over 800 faecal time-series samples from 126 piglets and 42 sows. Results Here we report on the largest shotgun metagenomic dataset of the pig gut lumen microbiome to date, consisting of >8 Tbp of shotgun metagenomic sequencing data. The animal trial, the workflow from sample collection to sample processing, and the preparation of libraries for sequencing, are described in detail. We provide a preliminary analysis of the dataset, centered on a taxonomic profiling of the samples, and a 16S-based beta diversity analysis of the mothers and the piglets in the first 5 weeks after weaning. Conclusions This study was conducted to generate a publicly available databank of the faecal metagenome of weaner piglets aged between 3 and 9 weeks old, treated with different probiotic formulations and intramuscular antibiotic treatment. Besides investigating the effects of the probiotic and intramuscular antibiotic treatment, the dataset can be explored to assess a wide range of ecological questions with regards to antimicrobial resistance, host-associated microbial and phage communities, and their dynamics during the aging of the host.
High-throughput short-read metagenomics has enabled large-scale species-level analysis and functional characterization of microbial communities. Microbiomes often contain multiple strains of the same species, and different strains have been shown to have important differences in their functional roles. Recent advances on long-read based methods enabled accurate assembly of bacterial genomes from complex microbiomes and an as-yet-unrealized opportunity to resolve strains. Here we present Strainberry, a metagenome assembly pipeline that performs strain separation in single-sample low-complexity metagenomes and that relies uniquely on long-read data. We benchmarked Strainberry on mock communities for which it produces strain-resolved assemblies with near-complete reference coverage and 99.9% base accuracy. We also applied Strainberry on real datasets for which it improved assemblies generating 20-118% additional genomic material than conventional metagenome assemblies on individual strain genomes. We show that Strainberry is also able to refine microbial diversity in a complex microbiome, with complete separation of strain genomes. We anticipate this work to be a starting point for further methodological improvements on strain-resolved metagenome assembly in environments of higher complexities.
Using a previously described metagenomics dataset of 27 billion reads, we reconstructed over 50 000 metagenome-assembled genomes (MAGs) of organisms resident in the porcine gut, 46.5 % of which were classified as 70 % complete with a <10 % contamination rate, and 24.4 % were nearly complete genomes. Here, we describe the generation and analysis of those MAGs using time-series samples. The gut microbial communities of piglets appear to follow a highly structured developmental programme in the weeks following weaning, and this development is robust to treatments including an intramuscular antibiotic treatment and two probiotic treatments. The high resolution we obtained allowed us to identify specific taxonomic 'signatures' that characterize the gut microbial development immediately after weaning. Additionally, we characterized the carbohydrate repertoire of the organisms resident in the porcine gut. We tracked the abundance shifts of 294 carbohydrate active enzymes, and identified the species and higher-level taxonomic groups carrying each of these enzymes in their MAGs. This knowledge can contribute to the design of probiotics and prebiotic interventions as a means to modify the piglet gut microbiome.
Frederick Blattner合作论文数Scarab Genomics;Dnastar5