Transposable elements (TEs) are a ubiquitous feature of plant genomes. Because of the threat they post to genome integrity, most TEs are epigenetically silenced. However, even closely related plant species often have dramatically different populations of TEs, suggesting periodic rounds of activity and silencing. Here, we show that the process of de novo methylation of an active element in maize involves two distinct pathways, one of which is directly implicated in causing epigenetic silencing and one of which is the result of that silencing. Epigenetic changes involve changes in gene expression that can be heritably transmitted to daughter cells in the absence of changes in DNA sequence. Epigenetics has been implicated in phenomena as diverse as development, stress response, and carcinogenesis. A significant challenge facing those interested in investigating epigenetic phenomena is determining causal relationships between DNA methylation, specific classes of small RNAs, and associated changes in gene expression. Because they are the primary targets of epigenetic silencing in plants and, when active, are often targeted for de novo silencing, TEs represent a valuable source of information about these relationships. We use a naturally occurring system in which a single TE can be heritably silenced by a single derivative of that TE. By using this system it is possible to unravel causal relationships between different size classes of small RNAs, patterns of DNA methylation, and heritable silencing. Here, we show that the long terminal inverted repeats within Zea mays MuDR transposons are targeted by distinct classes of small RNAs during epigenetic silencing that are dependent on distinct silencing pathways, only one of which is associated with transcriptional silencing of the transposon. Further, these small RNAs target distinct regions of the terminal inverted repeats, resulting in different patterns of cytosine methylation with different functional consequences with respect to epigenetic silencing and the heritability of that silencing.
Small RNAs are abundant in plant reproductive tissues, especially 24-nucleotide (nt) small interfering RNAs (siRNAs). Most 24-nt siRNAs are dependent on RNA Pol IV and RNA-DEPENDENT RNA POLYMERASE 2 (RDR2) and establish DNA methylation at thousands of genomic loci in a process called RNA-directed DNA methylation (RdDM). In Brassica rapa , RdDM is required in the maternal sporophyte for successful seed development. Here, we demonstrate that a small number of siRNA loci account for over 90% of siRNA expression during B. rapa seed development. These loci exhibit unique characteristics with regard to their copy number and association with genomic features, but they resemble canonical 24-nt siRNA loci in their dependence on RNA Pol IV/RDR2 and role in RdDM. These loci are expressed in ovules before fertilization and in the seed coat, embryo, and endosperm following fertilization. We observed a similar pattern of 24-nt siRNA expression in diverse angiosperms despite rapid sequence evolution at siren loci. In the endosperm, siren siRNAs show a marked maternal bias, and siren expression in maternal sporophytic tissues is required for siren siRNA accumulation. Together, these results demonstrate that seed development occurs under the influence of abundant maternal siRNAs that might be transported to, and function in, filial tissues.
Some organisms deploy small RNAs from accessory cells to maintain genome integrity in the zygote, a mechanism that has been proposed but not demonstrated in plants. Here we show that maternal mutations in the Pol IV-dependent small RNA pathway cause abortion of developing seeds in Brassica rapa . Surprisingly, small RNA production is required in maternal somatic tissues, but not in maternal gametes or the developing zygote. We propose that parental influence over zygotic genomes is a common strategy in eukaryotes and that outbreeding species such as B. rapa are key to understanding the role of small RNAs during reproduction.
While prokaryotic pan-genomes have been shown to contain many more genes than any individual organism, the prevalence and functional significance of differentially present genes in eukaryotes remains poorly understood. Whole-genome de novo assembly and annotation of 54 lines of the grass Brachypodium distachyon yield a pan-genome containing nearly twice the number of genes found in any individual genome. Genes present in all lines are enriched for essential biological functions, while genes present in only some lines are enriched for conditionally beneficial functions (e.g., defense and development), display faster evolutionary rates, lie closer to transposable elements and are less likely to be syntenic with orthologous genes in other grasses. Our data suggest that differentially present genes contribute substantially to phenotypic variation within a eukaryote species, these genes have a major influence in population genetics, and transposable elements play a key role in pan-genome evolution.
Oropetium thomaeum is a resurrection plant that can survive extreme water stress through desiccation to complete dryness, providing a model for drought tolerance; here, whole-genome sequencing and assembly of the Oropetium genome using single-molecule real-time sequencing is reported. Oropetium thomaeum is a resurrection plant that can endure extreme water stress through desiccation to complete dryness whilst retaining the ability to revive when water is available, providing a model for drought tolerance. These authors report whole-genome sequencing and assembly of the O. thomaeum genome, using only single-molecule real-time (SMRT) long-read sequencing. Understanding the genomic mechanisms of extreme desiccation tolerance in resurrection plants such as Oropetium may provide targets for engineering drought and stress tolerance in crop plants. Plant genomes, and eukaryotic genomes in general, are typically repetitive, polyploid and heterozygous, which complicates genome assembly1. The short read lengths of early Sanger and current next-generation sequencing platforms hinder assembly through complex repeat regions, and many draft and reference genomes are fragmented, lacking skewed GC and repetitive intergenic sequences, which are gaining importance due to projects like the Encyclopedia of DNA Elements (ENCODE)2. Here we report the whole-genome sequencing and assembly of the desiccation-tolerant grass Oropetium thomaeum. Using only single-molecule real-time sequencing, which generates long (>16 kilobases) reads with random errors, we assembled 99% (244 megabases) of the Oropetium genome into 625 contigs with an N50 length of 2.4 megabases. Oropetium is an example of a ‘near-complete’ draft genome which includes gapless coverage over gene space as well as intergenic sequences such as centromeres, telomeres, transposable elements and rRNA clusters that are typically unassembled in draft genomes. Oropetium has 28,466 protein-coding genes and 43% repeat sequences, yet with 30% more compact euchromatic regions it is the smallest known grass genome. The Oropetium genome demonstrates the utility of single-molecule real-time sequencing for assembling high-quality plant and other eukaryotic genomes, and serves as a valuable resource for the plant comparative genomics community.
The plant gene model remains largely an extrapolation from animals, with the cis functional unit, the gene, cast as a dynamic looping structure. Molecular genetics with model plants continues to make advances; highlighted here are quantitative-occupancy results from the Arabidopsis thaliana (Arabidopsis) Phytochrome-Interacting bHLH transcription Factors (PIF) quartet. Compared to this complex snapshot, results from chromatin occupancy and other Encyclopedia of DNA Elements (ENCODE)-like approaches increase our transcription factor-motif cognate library, but regulation cannot by itself be inferred from binding. Complementary published Arabidopsis conserved noncoding sequence lists are compared, evaluated, merged, and released. Comparative genomic approaches have identified a cis modifier of a gene's expression-hypothetically, a transposon-based 'rheostat'-that works in all cells, times and places.
In vertebrates, conserved noncoding elements (CNEs) are functionally constrained sequences that can show striking conservation over >400 million years of evolutionary distance and frequently are located megabases away from target developmental genes. Conserved noncoding sequences (CNSs) in plants are much shorter, and it has been difficult to detect conservation among distantly related genomes. In this article, we show not only that CNS sequences can be detected throughout the eudicot clade of flowering plants, but also that a subset of 37 CNSs can be found in all flowering plants (diverging similar to 170 million years ago). These CNSs are functionally similar to vertebrate CNEs, being highly associated with transcription factor and development genes and enriched in transcription factor binding sites. Some of the most highly conserved sequences occur in genes encoding RNA binding proteins, particularly the RNA splicing-associated SR genes. Differences in sequence conservation between plants and animals are likely to reflect differences in the biology of the organisms, with plants being much more able to tolerate genomic deletions and whole-genome duplication events due, in part, to their far greater fecundity compared with vertebrates.
The sequencing and analysis of the banana genome is reported; these results inform plant phylogenetic relationships and genome evolution, and provide a resource for future genetic improvement of this important crop species. Bananas (Musa spp.) are a staple food and a major source of income in many tropical and subtropical countries. This paper reports the sequencing and analysis of the banana genome. This is the first non-grass monocotyledon to have its genome sequenced, providing an important bridge for comparative genome analysis in plants. Global banana production is under threat from increasingly well-adapted pests and diseases, so the availability of the genome sequence is an important resource for future crop development and improvement. Bananas (Musa spp.), including dessert and cooking types, are giant perennial monocotyledonous herbs of the order Zingiberales, a sister group to the well-studied Poales, which include cereals. Bananas are vital for food security in many tropical and subtropical countries and the most popular fruit in industrialized countries1. The Musa domestication process started some 7,000 years ago in Southeast Asia. It involved hybridizations between diverse species and subspecies, fostered by human migrations2, and selection of diploid and triploid seedless, parthenocarpic hybrids thereafter widely dispersed by vegetative propagation. Half of the current production relies on somaclones derived from a single triploid genotype (Cavendish)1. Pests and diseases have gradually become adapted, representing an imminent danger for global banana production3,4. Here we describe the draft sequence of the 523-megabase genome of a Musa acuminata doubled-haploid genotype, providing a crucial stepping-stone for genetic improvement of banana. We detected three rounds of whole-genome duplications in the Musa lineage, independently of those previously described in the Poales lineage and the one we detected in the Arecales lineage. This first monocotyledon high-continuity whole-genome sequence reported outside Poales represents an essential bridge for comparative genome analysis in plants. As such, it clarifies commelinid-monocotyledon phylogenetic relationships, reveals Poaceae-specific features and has led to the discovery of conserved non-coding sequences predating monocotyledon–eudicotyledon divergence.
Salicylic acid (SA) is a critical mediator of plant innate immunity. It plays an important role in limiting the growth and reproduction of the virulent powdery mildew (PM) Golovinomyces orontii on Arabidopsis (Arabidopsis thaliana). To investigate this later phase of the PM interaction and the role played by SA, we performed replicated global expression profiling for wild-type and SA biosynthetic mutant isochorismate synthase1 (ics1) Arabidopsis from 0 to 7 d after infection. We found that ICS1-impacted genes constitute 3.8% of profiled genes, with known molecular markers of Arabidopsis defense ranked very highly by the multivariate empirical Bayes statistic (T 2 statistic). Functional analyses of T 2-selected genes identified statistically significant PM-impacted processes, including photosynthesis, cell wall modification, and alkaloid metabolism, that are ICS1 independent. ICS1-impacted processes include redox, vacuolar transport/secretion, and signaling. Our data also support a role for ICS1 (SA) in iron and calcium homeostasis and identify components of SA cross talk with other phytohormones. Through our analysis, 39 novel PM-impacted transcriptional regulators were identified. Insertion mutants in one of these regulators, PUX2 (for plant ubiquitin regulatory X domain-containing protein 2), results in significantly reduced reproduction of the PM in a cell death-independent manner. Although little is known about PUX2, PUX1 acts as a negative regulator of Arabidopsis CDC48, an essential AAA-ATPase chaperone that mediates diverse cellular activities, including homotypic fusion of endoplasmic reticulum and Golgi membranes, endoplasmic reticulum-associated protein degradation, cell cycle progression, and apoptosis. Future work will elucidate the functional role of the novel regulator PUX2 in PM resistance.