Acanthophis is a comprehensive pipeline for the joint analysis of both host genetic variation and variation in the composition and abundance of host-associated microbiomes (together, the "hologenome").Implemented in Snakemake (Köster & Rahmann, 2012), Acanthophis handles data from raw FASTQ read files through quality control, alignment of the reads to a plant reference, variant calling, taxonomic classification and quantification of microbes, and metagenome analysis.The workflow contains numerous practical optimisations, both to reduce disk space usage and maximise utilisation of computational resources.Acanthophis is available under the Mozilla Public Licence v2 at https://github.com/kdm9/Acanthophisas a python package installable from conda or PyPI (pip install acanthophis).
Massively parallel, second-generation short-read DNA sequencing has become an integral tool in biology for genomic studies. Offering highly accurate base-pair resolution at the most competitive price, the technology has become widespread. However, high-throughput generation of multiplexed DNA libraries can be costly and cumbersome. Here, we present a cost-conscious protocol for generating multiplexed short-read DNA libraries using a bead-linked transposome from Illumina. We prepare libraries in high-throughput with small reaction volumes that use 1/50th the amount of transposome compared to Illumina DNA Prep tagmentation protocols. By reducing transposome usage and optimising the protocol to circumvent magnetic bead-based clean-ups between steps, we reduce costs, labour time and DNA input requirements. Developing our own dual index primers further reduced costs and enables up to nine 96-well microplate combinations. This facilitates efficient usage of large-scale sequencing platforms, such as the Illumina NovaSeq 6000, which offers up to three terabases of sequencing per S4 flow cell. The protocol presented substantially reduces the cost per library by approximately 1/20th compared to conventional Illumina methods.
Many microRNAs (miRNAs) are encoded by small gene families. In a third of all conserved Arabidopsis miRNA families, members vary at two or more nucleotide positions. We have focused on the related miR159 and miR319 families, which share sequence identity at 17 of 21 nucleotides, yet affect different developmental processes through distinct targets. MiR159 regulates MYB mRNAs, while miR319 predominantly acts on TCP mRNAs. In the case of miR319, MYB targeting plays at most a minor role because miR319 expression levels and domain limit its ability to affect MYB mRNAs. In contrast, in the case of miR159, the miRNA sequence prevents effective TCP targeting. We complement these observations by identifying nucleotide positions relevant for miRNA activity with mutants recovered from a suppressor screen. Together, our findings reveal that functional specialization of miR159 and miR319 is achieved through both expression and sequence differences.
The unique ecology, pathology and undefined taxonomy of coconut foliar decay virus (CFDV), found associated with coconut foliar decay disease (CFD) in 1986, prompted analyses of old virus samples by modern methods. Rolling circle amplification and deep sequencing applied to nucleic acid extracts from virion preparations and CFD-affected palms identified twelve distinct circular DNAs, eleven of which had a size of about 1.3 kb and one of 641 nt. Mass spectrometry-based protein identification proved that a 24 kDa protein encoded by two 1.3 kb DNAs is the virus capsid protein with highest sequence similarity to that of grabloviruses (family Geminiviridae), even though CFDV particles are not geminate. The nine other 1.3 kb DNAs represent alphasatellites coding for replication initiator proteins that differ clearly from those encoded by nanovirid DNA-R. The 641 nt DNA-gamma is unique and may encode a movement protein. Three DNAs, alphasatellite CFDAR, capsid protein encoding CFDV DNA-S.1 and DNA-gamma share sequence motifs near their replication origins and were consistently present in all samples analysed. These DNAs appear to be integral components of a possibly tripartite CFDV genome, different from those of any Geminiviridae or Nanoviridae family member, implicating CFDV as representative of a new genus and family.
The development of model systems requires a detailed assessment of standing genetic variation across natural populations. The Brachypodium species complex has been promoted as a plant model for grass genomics with translation to small grain and biomass crops. To capture the genetic diversity within this species complex, thousands of Brachypodium accessions from around the globe were collected and genotyped by sequencing. Overall, 1897 samples were classified into two diploid or allopolyploid species, and then further grouped into distinct inbred genotypes. A core set of diverse B. distachyon diploid lines was selected for whole genome sequencing and high resolution phenotyping. Genome-wide association studies across simulated seasonal environments was used to identify candidate genes and pathways tied to key life history and agronomic traits under current and future climatic conditions. A total of 8, 22, and 47 QTL were identified for flowering time, early vigor, and energy traits, respectively. The results highlight the genomic structure of the Brachypodium species complex, and the diploid lines provided a resource that allows complex trait dissection within this grass model species.
The development of model systems requires a detailed assessment of standing genetic variation across natural populations. The Brachypodium species complex has been promoted as a plant model for grass genomics with translational to small grain and biomass crops. To capture the genetic diversity within this species complex, thousands of Brachypodium accessions from around the globe were collected and sequenced using genotyping by sequencing (GBS). Overall, 1,897 samples were classified into two diploid or allopolyploid species and then further grouped into distinct inbred genotypes. A core set of diverse B. distachyon diploid lines were selected for whole genome sequencing and high resolution phenotyping. Genome-wide association studies across simulated seasonal environments was used to identify candidate genes and pathways tied to key life history and agronomic traits under current and future climatic conditions. A total of 8, 22 and 47 QTLs were identified for flowering time, early vigour and energy traits, respectively. Overall, the results highlight the genomic structure of the Brachypodium species complex and allow powerful complex trait dissection within this new grass model species.
Mutants without root hairs show reduced inorganic orthophosphate (Pi) uptake and compromised growth on soils when Pi availability is restricted. What is less clear is whether root hairs that are longer than wild-type provide an additional benefit to phosphorus (P) nutrition. This was tested using transgenic Brachypodium lines with longer root hairs. The lines were transformed with the endogenous BdRSL2 and BdRSL3 genes using either a constitutive promoter or a root hair-specific promoter. Plants were grown for 32 d in soil amended with various Pi concentrations. Plant biomass and P uptake were measured and genotypes were compared on the basis of critical Pi values and P uptake per unit root length. Ectopic expression of RSL2 and RSL3 increased root hair length three-fold but decreased plant biomass. Constitutive expression of BdRSL2, but not expression of BdRSL3, consistently improved P nutrition as measured by lowering the critical Pi values and increasing Pi uptake per unit root length. Increasing root hair length through breeding or biotechnology can improve P uptake efficiency if the pleotropic effects on plant biomass are avoided. Long root hairs, alone, appear to be insufficient to improve Pi uptake and need to be combined with other traits to benefit P nutrition.
Most studies of aquatic plankton focus on either macroscopic or microbial communities, and on either eukaryotes or prokaryotes. This separation is primarily for methodological reasons, but can overlook potential interactions among groups. Here we tested whether DNA metabarcoding of unfractionated water samples with universal primers could be used to qualitatively and quantitatively study the temporal dynamics of the total plankton community in a shallow temperate lake. Significant changes in the relative proportions of normalized sequence reads of eukaryotic and prokaryotic plankton communities over a 3-month period in spring were found. Patterns followed the same trend as plankton estimates measured using traditional microscopic methods. The bloom of a conditionally rare bacterial taxon belonging to Arcicella was characterized, which rapidly came to dominate the whole lake ecosystem and would have remained unnoticed without metabarcoding. The data demonstrate the potential of universal DNA metabarcoding applied to unfractionated samples for providing a more holistic view of plankton communities.
Modern genomics techniques generate overwhelming quantities of data. Extracting population genetic variation demands computationally efficient methods to determine genetic relatedness between individuals (or "samples") in an unbiased manner, preferably de novo. Rapid estimation of genetic relatedness directly from sequencing data has the potential to overcome reference genome bias, and to verify that individuals belong to the correct genetic lineage before conclusions are drawn using mislabelled, or misidentified samples. We present the k-mer Weighted Inner Product (kWIP), an assembly-, and alignment-free estimator of genetic similarity. kWIP combines a probabilistic data structure with a novel metric, the weighted inner product (WIP), to efficiently calculate pairwise similarity between sequencing runs from their k-mer counts. It produces a distance matrix, which can then be further analysed and visualised. Our method does not require prior knowledge of the underlying genomes and applications include establishing sample identity and detecting mix-up, non-obvious genomic variation, and population structure. We show that kWIP can reconstruct the true relatedness between samples from simulated populations. By re-analysing several published datasets we show that our results are consistent with marker-based analyses. kWIP is written in C++, licensed under the GNU GPL, and is available from https://github.com/kdmurray91/kwip.
Freshwater fungi are a poorly studied paraphyletic group that include a high diversity of phyla. Most studies of aquatic fungal diversity have focussed on single habitats, thus the linkage between habitat heterogeneity and fungal diversity remains largely unexplored. We took 216 samples from 54 locations representing eight different habitats in meso-oligotrophic, temperate Lake Stechlin in northern Germany, including the pelagic and littoral water column, sediments, and biotic substrates. We pyrosequenced with an universal eukaryotic marker within the ribosomal large subunit (LSU) in order to compare fungal diversity, community structure, and species turnover among habitats. Our analysis recovered 1024 fungal OTUs (97% criterion). Diversity was highest in the sediment, biofilms, and benthic samples (293-428 OTUs), intermediate in water and reed samples (36-64 OTUs), and lowest in plankton (8 OTUs) samples. NMDS clustering clearly grouped the eight studied habitats into six clusters, indicating that total diversity was strongly influenced by turnover among habitats. Fungal communities exhibited pronounced changes at the levels of phylum and order along a gradient from littoral to pelagic habitats. The large majority of OTUs could not be classified below the order level due to the lack of aquatic fungal entries in taxonomic databases. Our study provides a first estimate of lake-wide fungal diversity and highlights the important contribution of habitat-specificity to total fungal diversity. This remarkable diversity is probably an underestimate, because most lakes undergo seasonal changes and previous studies have uncovered differences in fungal communities among lakes.
Most studies of biodiversity focus on either macroscopic or microbial communities, with little or no simultaneous study of eukaryotes and prokaryotes. We tested whether a universal metabarcoding approach could be used to study the total diversity and temporal dynamics of aquatic pico- to mesoplankton communities in a shallow temperate lake. The approach revealed significant changes in the relative abundance of eukaryotic and prokaryotic plankton communities over a period of three months. These patterns, based on sequencing reads, fit with counts using traditional methods. We also witnessed the bloom of a conditionally rare bacterial taxon belonging to Arcicella, a genus that has been largely overlooked in freshwaters. Our data demonstrate the potential of universal metabarcoding as a complement to traditional studies of plankton communities, and for long-term monitoring across a broad range of organisms.
Modern genomics techniques generate overwhelming quantities of data. Extracting population genetic variation demands computationally efficient methods to determine genetic relatedness between individuals or samples in an unbiased manner, preferably de novo . The rapid and unbiased estimation of genetic relatedness has the potential to overcome reference genome bias, to detect mix-ups early, and to verify that biological replicates belong to the same genetic lineage before conclusions are drawn using mislabelled, or misidentified samples. We present the k-mer Weighted Inner Product (kWIP), an assembly-, and alignment-free estimator of genetic similarity. kWIP combines a probabilistic data structure with a novel metric, the weighted inner product (WIP), to efficiently calculate pairwise similarity between sequencing runs from their k-mer counts. It produces a distance matrix, which can then be further analysed and visualised. Our method does not require prior knowledge of the underlying genomes and applications include detecting sample identity and mix-up, non-obvious genomic variation, and population structure. We show that kWIP can reconstruct the true relatedness between samples from simulated populations. By re-analysing several published datasets we show that our results are consistent with marker-based analyses. kWIP is written in C++, licensed under the GNU GPL, and is available from https://github.com/kdmurray91/kwip.
I n Sayou et al. (1), we explained how conserved LEAFY (LFY) homologs could recognize different DNA motifs (types I, II, and III) by determining the key residues affecting LFY DNA binding specificity. We identified a hypothetical promiscuous LFY variant (key residues His-Cys-His) through phylogenetic reconstruction, but also discovered a promiscuous variant (Gln-Cys-His) in a lineage of early diverging land plants, the hornworts. We proposed that these promiscuous forms acted as intermediates, enabling a gradual transition from type III to types I and II binding specificities. Brunkard et al. (2) identified novel paralogous LFY sequences in a single moss species and conclude that changes in LFY binding specificity evolved only through duplication. Here, we explain why we do not agree with their conclusions. First, Brunkard et al. constrain the LFY phylogeny to a single organismal topology in which liverworts, mosses, and hornworts constitute a paraphyletic grade leading to the vascular plants. This choice does not acknowledge the debate surrounding early land plant phylogeny. The topological constraint they employed is based on the phylogenetic hypothesis provided by Qiu et al. (3). However, Cox et al. reanalyzed the Qiu et al. data set and concluded that the paraphyly of bryophytes, and the support for the hornworts as the sister group to the tracheophytes, is a methodological artifact (4). Furthermore, other studies support alternative scenarios (5–10), and recent publications (4, 11) stress that the early land plant phylogeny remains profoundly uncertain, even in the postgenomic era. As emphasized in these publications and discussed in our original manuscript (1), there are four alternative topologies still in play [see figure S9 in (1)]. In this context, we think that it is inappropriate to constrain the LFY phylogeny to only one of several competing hypotheses of organismal relationships. Instead, in Sayou et al., we evaluated the incongruence between the LFY phylogeny and these four competing organismal phylogenies, establishing that our model is robust in the context of different organismal hypotheses. Second, on the basis of a sequence alignment alone, Brunkard et al. propose that the novel Polytrichum commune LFY sequences they isolated are the product of duplications at the base of mosses, rather than derived events. However, they did not perform the phylogenetic analysis necessary to establish whether these duplication events occurred before or after the changes in DNA binding specificity. In addition to duplications within mosses, their model also requires a likely duplication within hornworts and a deep duplication at the base of the land plants (mediating type III to type I specificity change). In the absence of this deep duplication, their model requires an abrupt switch between different binding specificities, which would likely be deleterious. To date, there is no support for any of these additional duplications, and Brunkard et al. do not provide a nondeleterious mechanism to account for an abrupt switch. Finally, Brunkard et al. rely on the parsimony criterion to substantiate their model. However, their parsimony reconstruction analysis [figure 2B in (2)] does not take into account that the three critical amino acid residues occupywell-separated positions and were reconstructed as such in Sayou et al. [figures 4 and S6 of (1)]; instead, Brunkard et al. appear to have reconstructed them as a single linked trait. Reanalyzing the Brunkard et al. data with the three amino acids sites individually reconstructed, but using their tree topology, we obtained an equal probability of the promiscuous intermediate (His-Cys-His) preceding key transitions in binding specificity (Fig. 1). Furthermore, the topology supplied by Brunkard et al. is the more challenging scenario for our model. In the other three competing organismal hypotheses (i.e., liverworts-plus-mosses asmonophyletic, all bryophytes asmonophyletic, or hornworts as sister to all land plants), a GlnCys-His ancestral promiscuous intermediate is always recovered. Therefore, we reject the idea that promiscuity is merely a derived state and maintain that the promiscuous model holds. We agree that gene duplication plays a role in LFY evolution because LFY duplicates may occasionally acquire different functions, likely through divergence in expression patterns (12, 13). In Sayou et al., we carefully noted all examples of known LFY duplications, even if not associated with a change in DNA binding specificity. We also agree that limited taxon sampling and genomic data make the presence of gene duplication difficult to disprove and concluded, “we cannot completely rule out the occurrence of transient ancient duplications” (1). In their Comment, Brunkard et al. treat the duplication and promiscuity scenarios as mutually exclusive. In contrast, we maintain that “it is plausible that the mechanisms we describe could also contribute to the evolution of TFs encoded bymultigene families” (1). Promiscuous intermediates may be easier to detect in a predominantly single-copy gene lineage, but promiscuous forms could themselves be duplicated and obtain novel function in derived paralogous lineages, as recently invoked in the evolution of HOX genes (14). In summary, we feel that Brunkard et al. misjudge the extent of the phylogenetic support for their model, and we suggest that their proposed scenario does not adequately explain the existence of the promiscuous form. Brunkard et al. highlight the need to densely sample LFY sequences from emerging genomic resources to RESEARCH
Land use management is a central challenge for the 21st century with unprecedented and competing demands to produce food, feed/fodder, fibre, fuel, and essential ecosystem services which sustain life. Global change requires rapid adaptation in current and emerging crops as well as in the foundation species of natural ecosystems. Revolutions in genomics and high throughput experimentation are transforming breeding so that adaptive traits in new environments can be predicted and selected more directly from germplasm collections of crops and wild species. This genomic breeding is now feasible in almost any species and has promise to help meet the need to feed and nourish over 9 billion people by 2050. Genomic techniques can accelerate our response to food security challenges of yield, quality and resilience and also address environmental security challenges. To achieve its potential there will need to be widespread and ongoing investments in the human capital to promote genomic breeding.
Periodic behavior in the climate system has important implications not only for weather prediction but also for understanding and interpreting the physical processes that drive climate variability. Here we demonstrate that the large-scale Southern Hemisphere atmospheric circulation exhibits marked periodicity on time scales of approximately 20 to 30 days. The periodicity is tied to the Southern Hemisphere baroclinic annular mode and emerges in hemispheric-scale averages of the eddy fluxes of heat, the eddy kinetic energy, and precipitation. Observational and theoretical analyses suggest that the oscillation results from feedbacks between the extratropical baroclinicity, the wave fluxes of heat, and radiative damping. The oscillation plays a potentially profound role in driving large-scale climate variability throughout much of the mid-latitude Southern Hemisphere.
Despite evolutionary conserved mechanisms to silence transposable element activity, there are drastic differences in the abundance of transposable elements even among closely related plant species. We conducted a de novo assembly for the 375 Mb genome of the perennial model plant, Arabis alpina . Analysing this genome revealed long-lasting and recent transposable element activity predominately driven by Gypsy long terminal repeat retrotransposons, which extended the low-recombining pericentromeres and transformed large formerly euchromatic regions into repeat-rich pericentromeric regions. This reduced capacity for long terminal repeat retrotransposon silencing and removal in A. alpina co-occurs with unexpectedly low levels of DNA methylation. Most remarkably, the striking reduction of symmetrical CG and CHG methylation suggests weakened DNA methylation maintenance in A. alpina compared with Arabidopsis thaliana . Phylogenetic analyses indicate a highly dynamic evolution of some components of methylation maintenance machinery that might be related to the unique methylation in A. alpina .
Sayou et al. (Reports, 7 February 2014, p. 645) proposed a new model for evolution of transcription factors without gene duplication, using LEAFY as an archetype. Their proposal contradicts the evolutionary history of plants and ignores evidence that LEAFY evolves through gene duplications. Within their data set, we identified a moss with multiple LEAFY orthologs, which contests their model and supports that LEAFY evolves through duplications.
Artificial microRNAs (amiRNAs) have been shown to facilitate efficient gene silencing with high specificity to the intended target gene(s). For the plant breeder, gene silencing by artificial miRNAs will certainly accelerate gene discovery, because it allows targeting of all genes in a mapping interval, independent of the genetic background. In addition, beneficial knockout phenotypes can easily be transferred between varieties and across incompatibility barriers. This chapter describes the generation and application of amiRNAs as a gene silencing tool in rice.
Recombination during meiosis shapes the complement of alleles segregating in the progeny of hybrids, and has important consequences for phenotypic variation. We examined allele frequencies, as well as crossover (XO) locations and frequencies in over 7000 plants from 17 F(2) populations derived from crosses between 18 Arabidopsis thaliana accessions. We observed segregation distortion between parental alleles in over half of our populations. The potential causes of distortion include variation in seed dormancy and lethal epistatic interactions. Such a high occurrence of distortion was only detected here because of the large sample size of each population and the number of populations characterized. Most plants carry only one or two XOs per chromosome pair, and therefore inherit very large, non-recombined genomic fragments from each parent. Recombination frequencies vary between populations but consistently increase adjacent to the centromeres. Importantly, recombination rates do not correlate with whole-genome sequence differences between parental accessions, suggesting that sequence diversity within A. thaliana does not normally reach levels that are high enough to exert a major influence on the formation of XOs. A global knowledge of the patterns of recombination in F(2) populations is crucial to better understand the segregation of phenotypic traits in hybrids, in the laboratory or in the wild.