Tea, one of the most widely consumed beverages globally, exhibits remarkable genomic diversity in its underlying flavour and health-related compounds. In this study, we present the construction and analysis of a tea pangenome comprising a total of 11 genomes, with a focus on three newly sequenced genomes comprising the purple-leaved assamica cultivar "Zijuan", the temperature-sensitive sinensis cultivar "Anjibaicha" and the wild accession "L618" whose assemblies exhibited excellent quality scores as they profited from latest sequencing technologies. Our analysis incorporates a detailed investigation of transposon complement across the tea pangenome, revealing shared patterns of transposon distribution among the studied genomes and improved transposon resolution with long read technologies, as shown by long terminal repeat (LTR) Assembly Index analysis. Furthermore, our study encompasses a gene-centric exploration of the pangenome, exploring the genomic landscape of the catechin pathway with our study, providing insights on copy number alterations and gene-centric variants, especially for Anthocyanidin synthases. We constructed a gene-centric pangenome by structurally and functionally annotating all available genomes using an identical pipeline, which both increased gene completeness and allowed for a high functional annotation rate. This improved and consistently annotated gene set will allow for a better comparison between tea genomes. We used this improved pangenome to capture the core and dispensable gene repertoire, elucidating the functional diversity present within the tea species. This pangenome resource might serve as a valuable resource for understanding the fundamental genetic basis of traits such as flavour, stress tolerance, and disease resistance, with implications for tea breeding programmes.
BackgroundPlant immunity relies on the perception of immunogenic signals by cell-surface and intracellular receptors and subsequent activation of defense responses like programmed cell death. Under certain circumstances, the fine-tuned innate immune system of plants results in the activation of autoimmune responses that cause constitutive defense responses and spontaneous cell death in the absence of pathogens.ResultsHere, we characterized the onset of leaf death 12 (old12) mutant that was identified in the Arabidopsis accession Landsberg erecta. The old12 mutant is characterized by a growth defect, spontaneous cell death, plant-defense gene activation, and early senescence. In addition, the old12 phenotype is temperature reversible, thereby exhibiting all characteristics of an autoimmune mutant. Mapping the mutated locus revealed that the old12 phenotype is caused by a mutation in the Lectin Receptor Kinase P2-TYPE PURINERGIC RECEPTOR 2 (P2K2) gene. Interestingly, the P2K2 allele from Landsberg erecta is conserved among Brassicaceae. P2K2 has been implicated in pathogen tolerance and sensing extracellular ATP. The constitutive activation of defense responses in old12 results in improved resistance against Pseudomonas syringae pv. tomato DC3000.ConclusionWe demonstrate that old12 is an auto-immune mutant and that allelic variation of P2K2 contributes to diversity in Arabidopsis immune responses.
AbstractGene structural annotation is a critical step in obtaining biological knowledge from genome sequences yet remains a major challenge in genomics projects. Currentde novoHidden Markov Models are limited in their capacity to model biological complexity; while current pipelines are resource-intensive and their results vary in quality with the available extrinsic data. Here, we build on our previous work in applying Deep Learning to gene calling to make a fully applicable, fast and user friendly tool for predicting primary gene models from DNA sequence alone. The quality is state-of-the-art, with predictions scoring closer by most measures to the references than to predictions from otherde novotools. Helixer’s predictions can be used as is or could be integrated in pipelines to boost quality further. Moreover, there is substantial potential for further improvements and advancements in gene calling with Deep Learning.Helixer is open source and available athttps://github.com/weberlab-hhu/HelixerA web interface is available athttps://www.plabipd.de/helixer_main.html
The common foodstuff garlic produces the potent antibiotic defense substance allicin after tissue damage. Allicin is a redox toxin that oxidizes glutathione and cellular proteins and makes garlic a highly hostile environment for non-adapted microbes. Genomic clones from a highly allicin-resistant Pseudomonas fluorescens (PfAR-1), which was isolated from garlic, conferred allicin resistance to Pseudomonas syringae and even to Escherichia coli Resistance-conferring genes had redox-related functions and were on core fragments from three similar genomic islands identified by sequencing and in silico analysis. Transposon mutagenesis and overexpression analyses revealed the contribution of individual candidate genes to allicin resistance. Taken together, our data define a multicomponent resistance mechanism against allicin in PfAR-1, achieved through horizontal gene transfer.
Global warming is becoming a significant problem for food security, particularly in the Mediterranean basin. The use of molecular techniques to study gene-level responses to environmental changes in non-model organisms is increasing and may help to improve the mechanistic understanding of durum wheat response to elevated CO2 and high temperature. With this purpose, we performed transcriptome RNA sequencing (RNA-Seq) analyses combined with physiological and biochemical studies in the flag leaf of plants grown in field chambers at ear emergence. Enhanced photosynthesis by elevated CO2 was accompanied by an increase in biomass and starch and fructan content, and a decrease in N compounds, as chlorophyll, soluble proteins, and Rubisco content, in association with a decline of nitrate reductase and initial and total Rubisco activities. While high temperature led to a decline of chlorophyll, Rubisco activity, and protein content, the glucose content increased and starch decreased. Furthermore, elevated CO2 induced several genes involved in mitochondrial electron transport, a few genes for photosynthesis and fructan synthesis, and most of the genes involved in secondary metabolism and gibberellin and jasmonate metabolism, whereas those related to light harvesting, N assimilation, and other hormone pathways were repressed. High temperature repressed genes for C, energy, N, lipid, secondary, and hormone metabolisms. Under the combined increases in atmospheric CO2 and temperature, the transcript profile resembled that previously reported for high temperature, although elevated CO2 partly alleviated the downregulation of primary and secondary metabolism genes. The results suggest that there was a reprogramming of primary and secondary metabolism under the future climatic scenario, leading to coordinated regulation of C-N metabolism towards C-rich metabolites at elevated CO2 and a shift away from C-rich secondary metabolites at high temperature. Several candidate genes differentially expressed were identified, including protein kinases, receptor kinases, and transcription factors.
Natural light environments are highly variable. Flexible adjustment between light energy utilization and photoprotection is therefore of vital importance for plant performance and fitness in the field. Short-term reactions to changing light intensity are triggered inside chloroplasts and leaves within seconds to minutes, whereas long-term adjustments proceed over hours and days, integrating multiple signals. While the mechanisms of long-term acclimation to light intensity have been studied by changing constant growth light intensity during the day, responses to fluctuating growth light intensity have rarely been inspected in detail. We performed transcriptome profiling in Arabidopsis (Arabidopsis thaliana) leaves to investigate long-term gene expression responses to fluctuating light (FL). In particular, we examined whether responses differ between young and mature leaves or between morning and the end of the day. Our results highlight global reprogramming of gene expression under FL, including that of genes related to photoprotection, photosynthesis, and photorespiration and to pigment, prenylquinone, and vitamin metabolism. The FL-induced changes in gene expression varied between young and mature leaves at the same time point and between the same leaves in the morning and at the end of the day, indicating interactions of FL acclimation with leaf development stage and time of day. Only 46 genes were up- or down-regulated in both young and mature leaves at both time points. Combined analyses of gene coexpression and cis-elements pointed to a role of the circadian clock and light in coordinating the acclimatory responses of functionally related genes. Our results also suggest a possible cross talk between FL acclimation and systemic acquired resistance-like gene expression in young leaves.
The antibiotic defense substance allicin (diallylthiosulfinate) is produced by garlic ( Allium sativum L.) after tissue damage, giving garlic its characteristic odor. Allicin is a redox-toxin that oxidizes thiols in glutathione and cellular proteins. A highly allicin-resistant Pseudomonas fluorescens strain ( Pf AR-1) was isolated from garlic, and genomic clones were shotgun electroporated into an allicin-susceptible P. syringae strain ( Ps 4612). Recipients showing allicin-resistance had all inherited a group of genes from one of three similar genomic islands (GI), that had been identified in an in silico analysis of the Pf AR-1 genome. A core fragment of 8-10 congruent genes with redox-related functions, present in each GI, was shown to confer allicin-specific resistance to P. syringae , and even to an unrelated E. coli strain. Transposon mutagenesis and overexpression analyses revealed the contribution of individual candidate genes to allicin-resistance. Moreover, Pf AR-1 was unusual in having 3 glutathione reductase ( glr ) genes, two copies in two of the GIs, but outside of the core group, and one copy in the Pf AR-1 genome. Glr activity was approximately 2-fold higher in Pf AR-1 than in related susceptible Pf 0-1, with only a single glr gene. Moreover, an E. coli Δ glr mutant showed increased susceptibility to allicin, which was complemented by Pf AR-1 glr1 . Taken together, our data support a multi-component resistance mechanism against allicin, achieved through horizontal gene transfer during coevolution, and allowing exploitation of the garlic ecological niche. GI regions syntenic with Pf AR-1 GIs are present in other plant-associated bacterial species, perhaps suggesting a wider role in adaptation to plants per se .
Upon local infection, plants activate a systemic immune response called systemic acquired resistance (SAR). During SAR, systemic leaves become primed for the superinduction of defense genes upon reinfection. We used formaldehyde-assisted isolation of regulatory DNA elements coupled to next-generation sequencing to identify SAR regulators. Our bioinformatic analysis produced 10,129 priming-associated open chromatin sites in the 5' region of 3,025 genes in the systemic leaves of Arabidopsis (Arabidopsis thaliana) plants locally infected with Pseudomonas syringae pv. maculicola Whole transcriptome shotgun sequencing analysis of the systemic leaves after challenge enabled the identification of genes with priming-linked open chromatin before (contained in the formaldehyde-assisted isolation of regulatory DNA elements sequencing dataset) and enhanced expression after (included in the whole transcriptome shotgun sequencing dataset) the systemic challenge. Among them, Arabidopsis MILDEW RESISTANCE LOCUS O3 (MLO3) was identified as a previously unidentified positive regulator of SAR. Further in silico analysis disclosed two yet unknown cis-regulatory DNA elements in the 5' region of genes. The P-box was mainly associated with priming-responsive genes, whereas the C-box was mostly linked to challenge. We found that the P- or W-box, the latter recruiting WRKY transcription factors, or combinations of these boxes, characterize the 5' region of most primed genes. Therefore, this study provides a genome-wide record of genes with open and accessible chromatin during SAR and identifies MLO3 and two previously unidentified DNA boxes as likely regulators of this immune response.
Summary Recent advances in genomics technologies have greatly accelerated the progress in both fundamental plant science and applied breeding research. Concurrently, high‐throughput plant phenotyping is becoming widely adopted in the plant community, promising to alleviate the phenotypic bottleneck. While these technological breakthroughs are significantly accelerating quantitative trait locus (QTL) and causal gene identification, challenges to enable even more sophisticated analyses remain. In particular, care needs to be taken to standardize, describe and conduct experiments robustly while relying on plant physiology expertise. In this article, we review the state of the art regarding genome assembly and the future potential of pangenomics in plant research. We also describe the necessity of standardizing and describing phenotypic studies using the Minimum Information About a Plant Phenotyping Experiment (MIAPPE) standard to enable the reuse and integration of phenotypic data. In addition, we show how deep phenotypic data might yield novel trait−trait correlations and review how to link phenotypic data to genomic data. Finally, we provide perspectives on the golden future of machine learning and their potential in linking phenotypes to genomic features.
Genome sequences from over 200 plant species have already been published, with this number expected to increase rapidly due to advances in sequencing technologies. Once a new genome has been assembled and the genes identified, the functional annotation of their putative translational products, proteins, using ontologies is of key importance as it places the sequencing data in a biological context. Furthermore, to keep pace with rapid production of genome sequences, this functional annotation process must be fully automated. Here we present a redesigned and significantly enhanced MapMan4 framework, together with a revised version of the associated online Mercator annotation tool. Compared with the original MapMan, the new ontology has been expanded almost threefold and enforces stricter assignment rules. This framework was then incorporated into Mercator4, which has been upgraded to reflect current knowledge across the land plant group, providing protein annotations for all embryophytes with a comparably high quality. The annotation process has been optimized to allow a plant genome to be annotated in a matter of minutes. The output results continue to be compatible with the established MapMan desktop application.
Background Given its tolerance to stress and its richness in particular secondary metabolites, the tobacco tree, Nicotiana glauca, has been considered a promising biorefinery feedstock that would not be competitive with food and fodder crops. Results Here we present a 3.5 Gbp draft sequence and annotation of the genome of N. glauca spanning 731,465 scaffold sequences, with an N50 size of approximately 92 kbases. Furthermore, we supply a comprehensive transcriptome and metabolome analysis of leaf development comprising multiple techniques and platforms. The genome sequence is predicted to cover nearly 80% of the estimated total genome size of N. glauca. With 73,799 genes predicted and a BUSCO score of 94.9%, we have assembled the majority of gene-rich regions successfully. RNA-Seq data revealed stage-and/or tissue-specific expression of genes, and we determined a general trend of a decrease of tricarboxylic acid cycle metabolites and an increase of terpenoids as well as some of their corresponding transcripts during leaf development. Conclusion The N. glauca draft genome and its detailed transcriptome, together with paired metabolite data, constitute a resource for future studies of valuable compound analysis in tobacco species and present the first steps towards a further resolution of phylogenetic, whole genome studies in tobacco.
Summary Dinitrogen fixation by Nostoc azollae residing in specialized leaf pockets supports prolific growth of the floating fern Azolla filiculoides. To evaluate contributions by further microorganisms, the A. filiculoides microbiome and nitrogen metabolism in bacteria persistently associated with Azolla ferns were characterized. A metagenomic approach was taken complemented by detection of N2O released and nitrogen isotope determinations of fern biomass. Ribosomal RNA genes in sequenced DNA of natural ferns, their enriched leaf pockets and water filtrate from the surrounding ditch established that bacteria of A. filiculoides differed entirely from surrounding water and revealed species of the order Rhizobiales. Analyses of seven cultivated Azolla species confirmed persistent association with Rhizobiales. Two distinct nearly full‐length Rhizobiales genomes were identified in leaf‐pocket‐enriched samples from ditch grown A. filiculoides. Their annotation revealed genes for denitrification but not N2‐fixation. 15N2 incorporation was active in ferns with N. azollae but not in ferns without. N2O was not detectably released from surface‐sterilized ferns with the Rhizobiales. N2‐fixing N. azollae, we conclude, dominated the microbiome of Azolla ferns. The persistent but less abundant heterotrophic Rhizobiales bacteria possibly contributed to lowering O2 levels in leaf pockets but did not release detectable amounts of the strong greenhouse gas N2O.
A parasitic lifestyle, where plants procure some or all of their nutrients from other living plants, has evolved independently in many dicotyledonous plant families and is a major threat for agriculture globally. Nevertheless, no genome sequence of a parasitic plant has been reported to date. Here we describe the genome sequence of the parasitic field dodder, Cuscuta campestris . The genome contains signatures of a fairly recent whole-genome duplication and lacks genes for pathways superfluous to a parasitic lifestyle. Specifically, genes needed for high photosynthetic activity are lost, explaining the low photosynthesis rates displayed by the parasite. Moreover, several genes involved in nutrient uptake processes from the soil are lost. On the other hand, evidence for horizontal gene transfer by way of genomic DNA integration from the parasite’s hosts is found. We conclude that the parasitic lifestyle has left characteristic footprints in the C. campestris genome.
Updates in nanopore technology have made it possible to obtain gigabases of sequence data. Prior to this, nanopore sequencing technology was mainly used to analyze microbial samples. Here, we describe the generation of a comprehensive nanopore sequencing data set with a median read length of 11,979 bp for a self-compatible accession of the wild tomato species Solanum pennellii We describe the assembly of its genome to a contig N50 of 2.5 MB. The assembly pipeline comprised initial read correction with Canu and assembly with SMARTdenovo. The resulting raw nanopore-based de novo genome is structurally highly similar to that of the reference S. pennellii LA716 accession but has a high error rate and was rich in homopolymer deletions. After polishing the assembly with Illumina reads, we obtained an error rate of <0.02% when assessed versus the same Illumina data. We obtained a gene completeness of 96.53%, slightly surpassing that of the reference S. pennellii Taken together, our data indicate that such long read sequencing data can be used to affordably sequence and assemble gigabase-sized plant genomes.
Recent massive growth in the production of sequencing data necessitates matching improve-ments in bioinformatics tools to effectively utilize it. Existing tools suffer from limitations in both scalability and applicability which are inherent to their underlying algorithms and data structures. We identify the key requirements for the ideal data structure for sequence analy-ses: it should be informationally lossless, locally updatable, and memory efficient; requirements which are not met by data structures underlying the major assembly strategies Overlap Layout Consensus and De Bruijn Graphs. We therefore propose a new data structure, the LOGAN graph, which is based on a memory efficient Sparse De Bruijn Graph with routing information. Innovations in storing routing information and careful implementation allow sequence datasets for Escherichia coli (4.6Mbp, 117x coverage), Arabidopsis thaliana (135Mbp, 17.5x coverage) and Solanum pennellii (1.2Gbp, 47x coverage) to be loaded into memory on a desktop computer in seconds, minutes, and hours respectively. Memory consumption is competitive with state of the art alternatives, while losslessly representing the reads in an indexed and updatable form. Both Second and Third Generation Sequencing reads are supported. Thus, the LOGAN graph is positioned to be the backbone for major breakthroughs in sequence analysis such as integrated hybrid assembly, assembly of exceptionally large and repetitive genomes, as well as assembly and representation of pan-genomes.
Summary The extreme sensitivity of the microsporogenesis process to moderately high or low temperatures is a major hindrance for tomato ( Solanum lycopersicum ) sexual reproduction and hence year‐round cropping. Consequently, breeding for parthenocarpy, namely, fertilization‐independent fruit set, is considered a valuable goal especially for maintaining sustainable agriculture in the face of global warming. A mutant capable of setting high‐quality seedless (parthenocarpic) fruit was found following a screen of EMS ‐mutagenized tomato population for yielding under heat stress. Next‐generation sequencing followed by marker‐assisted mapping and CRISPR /Cas9 gene knockout confirmed that a mutation in Sl AGAMOUS ‐ LIKE 6 (Sl AGL 6 ) was responsible for the parthenocarpic phenotype. The mutant is capable of fruit production under heat stress conditions that severely hamper fertilization‐dependent fruit set. Different from other tomato recessive monogenic mutants for parthenocarpy, Sl agl6 mutations impose no homeotic changes, the seedless fruits are of normal weight and shape, pollen viability is unaffected, and sexual reproduction capacity is maintained, thus making Sl agl6 an attractive gene for facultative parthenocarpy. The characteristics of the analysed mutant combined with the gene's mode of expression imply Sl AGL 6 as a key regulator of the transition between the state of ‘ovary arrest’ imposed towards anthesis and the fertilization‐triggered fruit set.
Recent updates in sequencing technology have made it possible to obtain Gigabases of sequence data from one single flowcell. Prior to this update, the nanopore sequencing technology was mainly used to analyze and assemble microbial samples 1-3 . Here, we describe the generation of a comprehensive nanopore sequencing dataset with a median fragment size of 11,979 bp for the wild tomato species Solanum pennellii featuring an estimated genome size of ca 1.0 to 1.1 Gbases. We describe its genome assembly to a contig N50 of 2.5 MB using a pipeline comprising a Canu 4 pre-processing and a subsequent assembly using SMARTdenovo. We show that the obtained nanopore based de novo genome reconstruction is structurally highly similar to that of the reference S. pennellii LA716 5 genome but has a high error rate caused mostly by deletions in homopolymers. After polishing the assembly with Illumina short read data we obtained an error rate of <0.02 % when assessed versus the same Illumina data. More importantly however we obtained a gene completeness of 96.53% which even slightly surpasses that of the reference S. pennellii genome 5 . Taken together our data indicate such long read sequencing data can be used to affordably sequence and assemble Gbase sized diploid plant genomes. Raw data is available at http://www.plabipd.de/portal/solanum-pennellii and has been deposited as PRJEB19787.
Sustainable agriculture demands reduced input of man-made nitrogen (N) fertilizer, yet N2 fixation limits the productivity of crops with heterotrophic diazotrophic bacterial symbionts. We investigated floating ferns from the genus Azolla that host phototrophic diazotrophic Nostoc azollae in leaf pockets and belong to the fastest growing plants. Experimental production reported here demonstrated N-fertilizer independent production of nitrogen-rich biomass with an annual yield potential per ha of 1200 kg-1 N fixed and 35 t dry biomass. 15N2 fixation peaked at noon, reaching 0.4 mg N g-1 dry weight h-1. Azolla ferns therefore merit consideration as protein crops in spite of the fact that little is known about the fern's physiology to enable domestication. To gain an understanding of their nitrogen physiology, analyses of fern diel transcript profiles under differing nitrogen fertilizer regimes were combined with microscopic observations. Results established that the ferns adapted to the phototrophic N2-fixing symbionts N. azollae by (1) adjusting metabolically to nightly absence of N supply using responses ancestral to ferns and seed plants; (2) developing a specialized xylem-rich vasculature surrounding the leaf-pocket organ; (3) responding to N-supply by controlling transcripts of genes mediating nutrient transport, allocation and vasculature development. Unlike other non-seed plants, the Azolla fern clock is shown to contain both the morning and evening loops; the evening loop is known to control rhythmic gene expression in the vasculature of seed plants and therefore may have evolved along with the vasculature in the ancestor of ferns and seed plants.
Streptomyces thermoautotrophicus UBT1 has been described as a moderately thermophilic chemolithoautotroph with a novel nitrogenase enzyme that is oxygen-insensitive. We have cultured the UBT1 strain and have isolated two new strains (H1 and P1-2) of very similar phenotypic and genetic characters. These strains show minimal growth on ammonium-free media and fail to incorporate isotopically labeled N2 gas into biomass in multiple independent assays. The sdn genes previously published as the putative nitrogenase of S. thermoautotrophicus have little similarity to anything found in draft genome sequences, published here, for strains H1 and UBT1, but share >99% nucleotide identity with genes from Hydrogenibacillus schlegelii, a draft genome for which is also presented here. H. schlegelii similarly lacks nitrogenase genes and is a non-diazotroph. We propose reclassification of the species containing strains UBT1, H1 and P1-2 as a non-Streptomycete, non-diazotrophic, facultative chemolithoautotroph and conclude that the existence of the previously proposed oxygen-tolerant nitrogenase is extremely unlikely.