Natural genetic variation in photosynthesis and photoprotection within crop germplasm represents an untapped resource for crop improvement. Sorghum bicolor (sorghum) is one of the world's most widely grown crops, yet the genetic basis of photoprotection in sorghum is not well understood. This study examined genetic variation in non-photochemical quenching traits by screening a field-grown panel of 861 genetically diverse natural sorghum accessions across 2 years. Broad-sense heritability ranged between 0.3 and 0.65 across different chlorophyll fluorescence parameters. A combination of genome- and transcriptome-wide (GWAS and TWAS) identification of genetic correlates with the observed trait variation uncovered a complex genetic architecture of many significant small-effect loci. An ensemble approach based on GWAS and TWAS results and the covariance between different fluorescence parameters was used to identify 110 unique candidate genes. The resulting high-confidence candidates reveal novel genetic associations with photoprotection and offer resources for further genetic studies and crop genomic improvement efforts.
Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.
Centuries of clonal propagation in cassava (Manihot esculenta) have reduced sexual recombination, leading to the accumulation of deleterious mutations. This has resulted in both inbreeding depression affecting yield and a significant decrease in reproductive performance, creating hurdles for contemporary breeding programs. Cassava is a member of the Euphorbiaceae family, including notable species such as rubber tree (Hevea brasiliensis) and poinsettia (Euphorbia pulcherrima). Expanding upon preliminary draft genomes, we annotated 7 long-read genome assemblies and aligned a total of 52 genomes, to analyze selection across the genome and the phylogeny. Through this comparative genomic approach, we identified 48 genes under relaxed selection in cassava. Notably, we discovered an overrepresentation of floral expressed genes, especially focused at 6 pollen-related genes. Our results indicate that domestication and a transition to clonal propagation have reduced selection pressures on sexually reproductive functions in cassava leading to an accumulation of mutations in pollen-related genes. This relaxed selection and the genome-wide deleterious mutations responsible for inbreeding depression are potential targets for improving cassava breeding, where the generation of new varieties relies on recombining favorable alleles through sexual reproduction.
Centuries of clonal propagation in cassava ( Manihot esculenta ) have engaged Muller’s Ratchet, leading to the accumulation of deleterious mutations due to the absence of sexual recombination. This has resulted in both inbreeding depression affecting yield and a significant decrease in reproductive performance, creating hurdles for contemporary breeding programs. Cassava is a member of the Euphorbiaceae family, including notable species such as rubber tree (Hevea brasiliensis) and poinsettia (Euphorbia pulcherrima). Expanding upon preliminary draft genomes, we annotated 7 long-read genome assemblies and aligned a total of 52 genomes, to analyze selection across the genome and the phylogeny. Through this comparative genomic approach, we identified 48 genes under relaxed selection in cassava. Notably, we discovered an overrepresentation of floral expressed genes, especially focused at six pollen-related genes. Our results indicate that domestication and a transition to clonal propagation has reduced selection pressures on sexually reproductive functions in cassava leading to an accumulation of mutations in pollen-related genes. This relaxed selection and the genome-wide deleterious mutations responsible for inbreeding depression are potential targets for improving cassava breeding, where the generation of new varieties relies on recombining favorable alleles through sexual reproduction.### Competing Interest StatementThe authors have declared no competing interest.
Pleiotropy-when a single gene controls two or more seemingly unrelated traits-has been shown to impact genes with effects on flowering time, leaf architecture, and inflorescence morphology in maize. However, the genome-wide impact of biological pleiotropy across all maize phenotypes is largely unknown. Here, we investigate the extent to which biological pleiotropy impacts phenotypes within maize using GWAS summary statistics reanalyzed from previously published metabolite, field, and expression phenotypes across the Nested Association Mapping population and Goodman Association Panel. Through phenotypic saturation of 120,597 traits, we obtain over 480 million significant quantitative trait nucleotides. We estimate that only 1.56-32.3% of intervals show some degree of pleiotropy. We then assess the relationship between pleiotropy and various biological features such as gene expression, chromatin accessibility, sequence conservation, and enrichment for gene ontology terms. We find very little relationship between pleiotropy and these variables when compared to permuted pleiotropy. We hypothesize that biological pleiotropy of common alleles is not widespread in maize and is highly impacted by nuisance terms such as population structure and linkage disequilibrium. Natural selection on large standing natural variation in maize populations may target wide and large effect variants, leaving the prevalence of detectable pleiotropy relatively low.
The 5’ untranslated region (UTR) sequence of eukaryotic mRNAs may contain upstream open reading frames (uORFs), which can regulate translation of the main open reading frame (mORF). The current model of translational regulation by uORFs posits that when a ribosome scans an mRNA and encounters a uORF, translation of that uORF can prevent ribosomes from reaching the mORF and cause decreased mORF translation. In this study, we first observed that rare variants in the 5’ UTR dysregulate protein abundance. Upon further investigation, we found that rare variants near the start codon of uORFs can repress or derepress mORF translation, causing allelic changes in protein abundance. This finding holds for common variants as well, and common variants that modify uORF start codons also contribute disproportionately to metabolic and whole-plant phenotypes, suggesting that translational regulation by uORFs serves an adaptive function. These results provide evidence for the mechanisms by which natural sequence variation modulates gene expression, and ultimately, phenotype.
The genomic basis underlying the selection for environmental adaptation and yield-related traits in maize remains poorly understood. Here we carried out genome-wide profiling of the small RNA (sRNA) transcriptome (sRNAome) and transcriptome landscapes of a global maize diversity panel under dry and wet conditions and uncover dozens of environment-specific regulatory hotspots. Transgenic and molecular studies of Drought-Related Environment-specific Super eQTL Hotspot on chromosome 8 (DRESH8) and ZmMYBR38, a target of DRESH8-derived small interfering RNAs, revealed a transposable element-mediated inverted repeats (TE-IR)-derived sRNA- and gene-regulatory network that balances plant drought tolerance with yield-related traits. A genome-wide scan revealed that TE-IRs associate with drought response and yield-related traits that were positively selected and expanded during maize domestication. These results indicate that TE-IR-mediated posttranscriptional regulation is a key molecular mechanism underlying the tradeoff between crop environmental adaptation and yield-related traits, providing potential genomic targets for the breeding of crops with greater stress tolerance but uncompromised yield.
The need for efficient tools and applications for analyzing genomic diversity is essential for any genetics research or breeding program. One commonly used tool, TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage), provides many core methods for genomic analyses. Despite its efficiency, TASSEL has limited automation potential for reproducible research and to interact with other analytical tools. Here we present an R package, rTASSEL, that is a front-end to connect to a variety of highly used TASSEL methods and analytical tools. The goal of this package is to create a unified scripting workflow that leverages the analytical prowess of TASSEL, in conjunction with R’s data handling and visualization capabilities, without ever having the user switch between these two environments.
Sorghum is a model C4 crop made experimentally tractable by extensive genomic and genetic resources. Biomass sorghum is also studied as a feedstock for biofuel and forage. Mechanistic modelling suggests that reducing stomatal conductance (gs) could improve sorghum intrinsic water use efficiency (iWUE) and biomass production. Phenotyping for discovery of genotype to phenotype associations remain bottlenecks in efforts to understand the mechanistic basis for natural variation in gs and iWUE. This study addressed multiple methodological limitations. Optical tomography and a novel machine learning tool were combined to measure stomatal density (SD). This was combined with rapid measurements of leaf photosynthetic gas exchange and specific leaf area (SLA). These traits were then the subject of genome-wide association study (GWAS) and transcriptome-wide association study (TWAS) across 869 field-grown biomass sorghum accessions. SD was correlated with plant height and biomass production. Plasticity in SD and SLA were interrelated with each other, and productivity, across wet versus dry growing seasons. Moderate-to-high heritability of traits studied across the large mapping population supported identification of associations between DNA sequence variation, or RNA transcript abundance, and trait variation. 394 unique genes underpinning variation in WUE-related traits are described with higher confidence because they were identified in multiple independent tests. This list was enriched in genes whose orthologs in Arabidopsis have functions related to stomatal or leaf development and leaf gas exchange. These advances in methodology and knowledge will aid efforts to improve the WUE of C4 crops.
With the recent release of genome assemblies for the NAM parents, we review the impact that the maize NAM population has had on the community and discuss its utility as we move into a new era of genomics. It has been just over a decade since the release of the maize (Zea mays) Nested Association Mapping (NAM) population. The NAM population has been and continues to be an invaluable resource for the maize genetics community and has yielded insights into the genetic architecture of complex traits. The parental lines have become some of the most well-characterized maize germplasm, and their de novo assemblies were recently made publicly available. As we enter an exciting new stage in maize genomics, this retrospective will summarize the design and intentions behind the NAM population; its application, the discoveries it has enabled, and its influence in other systems; and use the past decade of hindsight to consider whether and how it will remain useful in a new age of genomics.
The need for efficient tools and applications for analyzing genomic diversity is essential for any genetics research program. One such tool, TASSEL (Trait Analysis by aSSociation, Evolution and Linkage), provides many core methods for genomic analyses. Despite its efficiency, TASSEL has limited means for reproducible research and interacting with other analytical tools. Here we present an R package rTASSEL, a front-end to connect to a variety of highly used TASSEL methods and analytical tools. The goal of this package is to create a unified scripting workflow that exploits the analytical prowess of TASSEL in conjunction with R’s popular data handling and parsing capabilities without ever having the user to switch between these two environments.
Differential gene expression (DGE) analysis is one of the most common applications of RNA-sequencing (RNA-seq) data. This process allows for the elucidation of differentially expressed genes across two or more conditions and is widely used in many applications of RNA-seq data analysis. Interpretation of the DGE results can be nonintuitive and time consuming due to the variety of formats based on the tool of choice and the numerous pieces of information provided in these results files. Here we reviewed DGE results analysis from a functional point of view for various visualizations. We also provide an R/Bioconductor package, Visualization of Differential Gene Expression Results using R, which generates information-rich visualizations for the interpretation of DGE results from three widely used tools, Cuffdiff, DESeq2 and edgeR. The implemented functions are also tested on five real-world data sets, consisting of one human, one Malus domestica and three Vitis riparia data sets.
Next-Generation Sequencing has made available substantial amounts of large-scale Omics data, providing unprecedented opportunities to understand complex biological systems. Specifically, the value of RNA-Sequencing (RNA-Seq) data has been confirmed in inferring how gene regulatory systems will respond under various conditions (bulk data) or cell types (single-cell data). RNA-Seq can generate genome-scale gene expression profiles that can be further analyzed using correlation analysis, co-expression analysis, clustering, differential gene expression (DGE), among many other studies. While these analyses can provide invaluable information related to gene expression, integration and interpretation of the results can prove challenging. Here we present a tool called IRIS-EDA, which is a Shiny web server for expression data analysis. It provides a straightforward and user-friendly platform for performing numerous computational analyses on user-provided RNA-Seq or Single-cell RNA-Seq (scRNA-Seq) data. Specifically, three commonly used R packages (edgeR, DESeq2, and limma) are implemented in the DGE analysis with seven unique experimental design functionalities, including a user-specified design matrix option. Seven discovery-driven methods and tools (correlation analysis, heatmap, clustering, biclustering, Principal Component Analysis (PCA), Multidimensional Scaling (MDS), and t-distributed Stochastic Neighbor Embedding (t-SNE)) are provided for gene expression exploration which is useful for designing experimental hypotheses and determining key factors for comprehensive DGE analysis. Furthermore, this platform integrates seven visualization tools in a highly interactive manner, for improved interpretation of the analyses. It is noteworthy that, for the first time, IRIS-EDA provides a framework to expedite submission of data and results to NCBI’s Gene Expression Omnibus following the FAIR (Findable, Accessible, Interoperable and Reusable) Data Principles. IRIS-EDA is freely available at http://bmbl.sdstate.edu/IRIS/.
Differential gene expression ( DGE ) is one of the most common applications of RNA-sequencing (RNA-seq) data. This process allows for the elucidation of differentially expressed genes ( DEGs ) across two or more conditions. Interpretation of the DGE results can be non-intuitive and time consuming due to the variety of formats based on the tool of choice and the numerous pieces of information provided in these results files. Here we present an R package, ViDGER (Visualization of Differential Gene Expression Results using R), which contains nine functions that generate information-rich visualizations for the interpretation of DGE results from three widely-used tools, Cuffdiff , DESeq2 , and edgeR .