Understanding and predicting complex traits in plants remains a fundamental challenge due to the emergent nature of most phenotypes and their dependence on genetic, regulatory, and environmental interactions. Accurate prediction of traits and identification of underlying genetic elements has broad applications for plant breeding, systems biology, and biotechnology. Here, we tested if multi-omic datasets could improve predictive accuracy of 129 diverse maize phenotypes across nine environments using genomic markers, field based transcriptomic data from two locations, and drone-derived phenomic data of vegetative indices. We trained and compared linear (rrBLUP) and nonlinear (support vector regression) models using single- and multi-omics inputs. Multi-omics models consistently outperformed single-omics models for most traits, with genomic and transcriptomic inputs contributing distinct biological features. Phenomic features alone yielded the lowest predictive power but improved predictions for specific trait categories like root architecture. Transcriptomic datasets enabled cross-environment prediction, demonstrating that gene expression patterns from one field site could accurately predict traits measured in another. Environment-specific expression of benchmark flowering time genes highlighted the value of transcriptomics in capturing genotype-by-environment (G×E) interactions not detectable through genomic data alone. These findings demonstrate that integrating transcriptomic and phenomic data with genotypes enhances trait prediction, improves model generalizability across environments, and provides deeper insight into the genetic and regulatory architecture of agriculturally important traits in maize.
The increasing frequency and intensity of droughts present significant challenges to global food security. In this study, we examined the genetic and physiological mechanisms underlying drought tolerance and resilience in sorghum (Sorghum bicolor L.) by phenotyping the Sorghum Association Panel (SAP; n = 397) for a broad suite of traits. These included leaf anatomical characteristics (stomatal density [SD], stomatal size, pore area, stomatal pore area per leaf area, and anatomical maximum stomatal gas conductance), physiological traits [net photosynthetic rate (An), stomatal gas conductance (gsw), and intrinsic water-use efficiency (iWUE)], and functional traits (leaf width, leaf thickness, leaf mass area, and chlorophyll content). Substantial natural variation was detected within the SAP, and correlation analyses indicated that leaf anatomical and functional characteristics play key roles in regulating physiological traits, including An, gsw, and iWUE. Genome-wide association studies identified a genomic hotspot on chromosome 1 (77.5-78.6 Mb) region associated with 3 key single-nucleotide polymorphisms (S01_77550396, S01_78561058, and S01_78619413). Haplotype analysis of these loci uncovered 8 distinct allele combinations influencing SD, An, gsw, and iWUE. Application of the Ball-Woodrow-Berry gsw model to these haplotypes demonstrated that accessions from haplotypes 1 to 5 exhibited greater stomatal plasticity, displaying more dynamic responses under well-watered conditions. In contrast, accessions from haplotypes 6 to 8 showed more conservative stomatal behavior under water-limited conditions. These results provide insights into the coordinated genetic control of leaf traits underlying drought resilience in sorghum and offer a predictive framework for breeding cultivars with stable performance across diverse water regimes.
A simple and effective method to identify genetic markers of yield response to nitrogen (N) fertilizer among maize hybrids is urgently needed. In this article, we describe a detailed methodology to identify genetic markers and develop associated assays for the prediction of yield N-response in maize. We first outline an in silico workflow to identify high-priority single-nucleotide polymorphism (SNP) markers from genome-wide association studies (GWAS). We then describe a detailed methodology to develop cleaved amplified polymorphic sequences (CAPS) and derived CAPS (dCAPS)-based assays to quickly and effectively test genetic marker subsets. This protocol is expected to provide a robust approach to determine N-response type among maize germplasm, including elite commercial varieties, allowing more appropriate on-farm N application rates, minimizing N fertilizer waste.
Sorghum is emerging as an ideal genetic model for designing high-biomass bioenergy crops. Biomass yield, a complex trait influenced by various plant architectural features, is typically regulated by numerous genes. This study aims to dissect the genetic mechanisms underlying fourteen plant architectural and ten biomass yield traits in a sorghum association panel (SAP) across two growing seasons. We identified 321 associated loci via genome-wide association studies involving 234,264 single nucleotide polymorphisms (SNPs). These loci encompass both genes with a priori links to biomass traits, such as maturity, dwarfing (Dw), leafbladeless1, cryptochrome, and several loci not previously linked to roles in determining these traits. We identified 22 pleiotropic loci associated with variation in multiple phenotypes. Three of these loci, located on chromosomes 3 (S03_15463061), 6 (S06_42790178; Dw2), and 9 (S09_57005346; Dw1), exert significant and consistent effects on multiple traits. Additionally, we identified three genomic hotspots on chromosomes 6, 7, and 9, containing multiple SNPs associated with variation in plant architecture and biomass yield traits. Positive correlations were observed among linked SNPs close to or within the same genomic regions. Thirteen haplotypes were identified from these positively correlated SNPs on chr 6, with haplotypes 8 and 11 emerging as optimal combinations, exhibiting pronounced effects on the traits. Lastly, network analysis revealed that loci associated with flowering, plant heights, leaf characteristics, plant number, and tiller number per plant were highly interconnected with other genetic loci linked to plant architecture and biomass yield traits. The pyramiding of favorable alleles related to these traits holds promise for enhancing the future development of bioenergy sorghum crops.
Maize (Zea mays) production systems are heavily reliant on the provision of managed inputs such as fertilizers to maximize growth and yield. Hence, the effective use of nitrogen (N) fertilizer is crucial to minimize the associated financial and environmental costs, as well as maximize yield. However, how to effectively utilize N inputs for increased grain yields remains a substantial challenge for maize growers that requires a deeper understanding of the underlying physiological responses to N fertilizer application. We report a multiscale investigation of five field-grown maize hybrids under low or high N supplementation regimes that includes the quantification of phenolic and prenyl-lipid compounds, cellular ultrastructural features, and gene expression traits at three developmental stages of growth. Our results reveal that maize perceives the lack of supplemented N as a stress and, when provided with additional N, will prolong vegetative growth. However, the manifestation of the stress and responses to N supplementation are highly hybrid-specific. Eight genes were differentially expressed in leaves in response to N supplementation in all tested hybrids and at all developmental stages. These genes represent potential biomarkers of N status and include two isoforms of Thiamine Thiazole Synthase involved in vitamin B1 biosynthesis. Our results uncover a detailed view of the physiological responses of maize hybrids to N supplementation in field conditions that provides insight into the interactions between management practices and the genetic diversity within maize.
Transcriptome-wide association studies (TWAS) can provide single gene resolution for candidate genes in plants, complementing genome-wide association studies (GWAS) but efforts in plants have been met with, at best, mixed success. We generated expression data from 693 maize genotypes, measured in a common field experiment, sampled over a 2-h period to minimize diurnal and environmental effects, using full-length RNA-seq to maximize the accurate estimation of transcript abundance. TWAS could identify roughly 10 times as many genes likely to play a role in flowering time regulation as GWAS conducted data from the same experiment. TWAS using mature leaf tissue identified known true-positive flowering time genes known to act in the shoot apical meristem, and trait data from a new environment enabled the identification of additional flowering time genes without the need for new expression data. eQTL analysis of TWAS-tagged genes identified at least one additional known maize flowering time gene through trans-eQTL interactions. Collectively these results suggest the gene expression resource described here can link genes to functions across different plant phenotypes expressed in a range of tissues and scored in different experiments.
AbstractPlant breeding relies on information gathered from field trials to select promising new crop varieties for release to farmers and to develop genomic prediction models that can enhance the efficiency of genetic improvement in future breeding cycles. However, generating the genetic marker data required to apply genomic prediction at the early stages of a breeding program remains costly for many public‐sector breeding programs as well as for many plant breeders operating in developing countries. As the pace of climate change intensifies, the time lag of developing and deploying new crop varieties requires plant breeders to make selection decisions without knowing the future environments those crop varieties will encounter in farmers’ fields. Therefore, both lower cost and higher accuracy methods for prediction of crop performance are essential for creating and maintaining resilient agricultural systems in the latter half of the 21st century. To address this challenge, we conducted linked yield trials of 752 public maize (Zea mays) genotypes in two distinct environments. We developed and trained a phenotypic prediction model to predict yield from manually scored plant traits. The phenotypic prediction approach we employed outperformed genomic prediction in predicting yields in a second environment, with 8.7%–63% higher R2 and 4%–13% less root mean square error than the genomic prediction. The phenotypic prediction has the potential to be applied to a wider range of breeding programs, including those that lack the resources to genotype large populations, such as programs in the developing world, breeding programs for specialty crops, and public sector programs.
Comparing the genomes of two species offers a robust approach to unveil potentially overlooked genes that wield substantial influence on observed agronomic traits. Extensive phenotypic data from maize and sorghum were previously utilized in genome-wide association studies. This project aims to harness the potential of comparative genomics to strengthen confidence in these marker-trait associations and to suggest previously unexplored relationships. To achieve this goal, insights from studies on the genetically diverse Sorghum Association Panel and the maize Wisconsin Diversity Panel were leveraged to classify candidate genes as either shared between the two species or unique to one or the other. Candidate orthologous genes found in both species, associated with shared phenotypes, enhance the reliability of the associations within each species. Additionally, genes unique to each species provide parameters to inform future predictive models. Finally, given the ancient tetraploidy of maize and the biased loss of genes over time, the candidate genes identified in sorghum provide valuable information for understanding the orthologs found in the maize subgenomes.
Enhancing photosynthesis for increased sorghum grain yield has become a key focus in sorghum breeding efforts. Phenotyping, involving the measurement of various morpho-physiological and physical traits associated with photosynthesis and grain yield, is a time-intensive process. However, the potential of non-invasive leaf-level hyperspectral imaging to swiftly detect plant performance, optimizing photosynthesis and grain yield, is promising. This study aimed to evaluate the feasibility of utilizing hyperspectral reflectance in the 350–950 nm range for the rapid estimation of these traits in intact sorghum leaves. Multiple machine learning regression algorithms were developed using leaf-level hyperspectral reflectance data from nearly 400 sorghum accessions within an association panel. The best-performing prediction models were then considered as potential methods for constructing a prediction model targeting multiple other physiological and yield traits in sorghum accessions. The results indicate that this approach enables the early detection of leaf photosynthetic and yield traits through leaf-level hyperspectral reflectance without the need for a full-range, high-cost leaf spectrometer.
Nannochloropsis oceanica, as other stramenopile microalgae, is rich in long-chain polyunsaturated fatty acids (LC-PUFA) such as eiconsapentaenoic acid (EPA). We observed that fatty acid desaturases (FAD) involved in LC-PUFA biosynthesis were among the strongest blue light induced genes in N. oceanica CCMP1779. Blue light was also necessary for maintaining LC-PUFA levels in CCMP1779 cells, and growth under red light led to a reduction in EPA content. Aureochromes are stramenopile specific proteins that contain a light-oxygen-voltage-sensing (LOV) domain that associates with a flavin mononucleotide and is able to sense blue light. These proteins also contain a bZIP DNA binding motif and can act as blue light regulated transcription factors by associating with a E-box like motif, which we found enriched in the promoters of blue light induced genes. We demonstrated that, in vitro, two CCMP1779 aureochromes were able to absorb blue light. Moreover, the loss or reduction of any of the three aureochromes led to a decrease in the blue light specific induction of several FADs in CCMP1779. EPA content was also significantly reduced in NoAureo 2 and NoAureo 4 mutants. Taken together, our results indicate that aureochromes mediate blue light dependent regulation of LC-PUFA content in N. oceanica CCMP1779 cells.
Classical genetic studies have identified many cases of pleiotropy where mutations in individual genes alter many different phenotypes. Quantitative genetic studies of natural genetic variants frequently examine one or a few traits, limiting their potential to identify pleiotropic effects of natural genetic variants. Widely adopted community association panels have been employed by plant genetics communities to study the genetic basis of naturally occurring phenotypic variation in a wide range of traits. High-density genetic marker data—18M markers—from 2 partially overlapping maize association panels comprising 1,014 unique genotypes grown in field trials across at least 7 US states and scored for 162 distinct trait data sets enabled the identification of of 2,154 suggestive marker-trait associations and 697 confident associations in the maize genome using a resampling-based genome-wide association strategy. The precision of individual marker-trait associations was estimated to be 3 genes based on a reference set of genes with known phenotypes. Examples were observed of both genetic loci associated with variation in diverse traits (e.g., above-ground and below-ground traits), as well as individual loci associated with the same or similar traits across diverse environments. Many significant signals are located near genes whose functions were previously entirely unknown or estimated purely via functional data on homologs. This study demonstrates the potential of mining community association panel data using new higher-density genetic marker sets combined with resampling-based genome-wide association tests to develop testable hypotheses about gene functions, identify potential pleiotropic effects of natural genetic variants, and study genotype-by-environment interaction.
Plant growth and development is impacted by the ability to capture resources including sunlight, determined in part by the arrangement of plant parts throughout the canopy. This is a very complex t...
Chiococca alba (L.) Hitchc. (snowberry), a member of the Rubiaceae, has been used as a folk remedy for a range of health issues including inflammation and rheumatism and produces a wealth of specialized metabolites including terpenes, alkaloids, and flavonoids. We generated a 558Mb draft genome assembly for snowberry which encodes 28,707 high-confidence genes. Comparative analyses with other angiosperm genomes revealed enrichment in snowberry of lineage-specific genes involved in specialized metabolism. Synteny between snowberry and Coffea canephora Pierre ex A. Froehner (coffee) was evident, including the chromosomal region encoding caffeine biosynthesis in coffee, albeit syntelogs of N-methyltransferase were absent in snowberry. A total of 27 putative terpene synthase genes were identified, including 10 that encode diterpene synthases. Functional validation of a subset of putative terpene synthases revealed that combinations of diterpene synthases yielded access to products of both general and specialized metabolism. Specifically, we identified plausible intermediates in the biosynthesis of merilactone and ribenone, structurally unique antimicrobial diterpene natural products. Access to the C. alba genome will enable additional characterization of biosynthetic pathways responsible for health-promoting compounds in this medicinal species.
Potato ( L.) breeders often use dihaploids, which are 2× progeny derived from 4× autotetraploid parents. Dihaploids can be used in diploid crosses to introduce new genetic material into breeding germplasm that can be integrated into tetraploid breeding through the use of unreduced gametes in 4× by 2× crosses. Dihaploid potatoes are usually produced via pollination by haploid inducer lines known as in vitro pollinators (IVP). In vitro pollinator chromosomes are selectively degraded from initially full hybrid embryos, resulting in 2× seed. During this process, somatic translocation of IVP DNA may occur. In this study, a genome-wide approach was used to identify such events and other chromosome-scale abnormalities in a population of 95 dihaploids derived from a cross between potato cultivar Superior and the haploid inducing line IVP101. Most Superior dihaploids showed translocation rates of <1% at 16,947,718 assayable sites, yet two dihaploids showed translocation rates of 1.86 and 1.60%. Allelic ratios at translocation sites suggested that most translocations occurred in individual cell lineages and were thus not present in all cells of the adult plants. Translocations were enriched in sites associated with high gene expression and H3K4 dimethylation and H4K5 acetylation, suggesting that they tend to occur in regions of open chromatin. The translocations likely result as a consequence of double-stranded break repair in the dihaploid genomes via homologous recombination during which IVP chromosomes are used as templates. Additionally, primary trisomy was observed in eight individuals. As the trisomic chromosomes were derived from Superior, meiotic nondisjunction may be common in potato.
Genome mining is a routine technique in microbes for discovering biosynthetic pathways. In plants, however, genomic information is not commonly used to identify novel biosynthesis genes. Here, we present the genome of the medicinal plant and oxindole monoterpene indole alkaloid (MIA) producer Gelsemium sempervirens (Gelsemiaceae). A gene cluster from Catharanthus roseus, which is utilized at least six enzymatic steps downstream from the last common intermediate shared between the two plant alkaloid types, is found in G. sempervirens, although the corresponding enzymes act on entirely different substrates. This study provides insights into the common genomic context of MIA pathways and is an important milestone in the further elucidation of the Gelsemium oxindole alkaloid pathway.
Circadian clocks allow organisms to predict environmental changes caused by the rotation of the Earth. Although circadian rhythms are widespread among different taxa, the core components of circadian oscillators are not conserved and differ between bacteria, plants, animals and fungi. Stramenopiles are a large group of organisms in which circadian rhythms have been only poorly characterized and no clock components have been identified. We have investigated cell division and molecular rhythms in Nannochloropsis species. In the four strains tested, cell division occurred principally during the night period under diel conditions, however, rhythms dampened within 2-3 days after transfer to constant light. We developed firefly luciferase reporters for long-term monitoring of in vivo transcriptional rhythms in two Nannochlropsis species, N. oceanica CCMP1779 and N. salina CCMP537. The reporter lines express free-running bioluminescence rhythms with periods of ~21-31 h that dampen within ~3-4 days under constant light. Using different entrainment regimes, we demonstrate that these rhythms are regulated by a circadian-type oscillator. In addition, the phase of free-running luminescence rhythms can be modulated pharmacologically using a CK1 ε/δ inhibitor, suggesting a role of this kinase in the Nannochloropsis clock. Together with the molecular and genomic tools available for Nannochloropsis species, these reporter lines represent an excellent system for future studies on the molecular mechanisms of stramenopile circadian oscillators. Significance statement Stramenopiles are a large and diverse line of eukaryotes in which circadian rhythms have been only poorly characterized and no clock components have been identified. We have developed bioluminescence reporter lines in Nannochloropsis species and provide evidence for the presence of a circadian oscillator in stramenopiles; these lines will serve as tools for future studies to uncover the molecular mechanisms of circadian oscillations in these species.
The evolution of chemical complexity has been a major driver of plant diversification, with novel compounds serving as key innovations. The species-rich mint family (Lamiaceae) produces an enormous variety of compounds that act as attractants and defense molecules in nature and are used widely by humans as flavor additives, fragrances, and anti-herbivory agents. To elucidate the mechanisms by which such diversity evolved, we combined leaf transcriptome data from 48 Lamiaceae species and four outgroups with a robust phylogeny and chemical analyses of three terpenoid classes (monoterpenes, sesquiterpenes, and iridoids) that share and compete for precursors. Our integrated chemical-genomic-phylogenetic approach revealed that: (1) gene family expansion rather than increased enzyme promiscuity of terpene synthases is correlated with mono- and sesquiterpene diversity; (2) differential expression of core genes within the iridoid biosynthetic pathway is associated with iridoid presence/absence; (3) generally, production of iridoids and canonical monoterpenes appears to be inversely correlated; and (4) iridoid biosynthesis is significantly associated with expression of geraniol synthase, which diverts metabolic flux away from canonical monoterpenes, suggesting that competition for common precursors can be a central control point in specialized metabolism. These results suggest that multiple mechanisms contributed to the evolution of chemodiversity in this economically important family.
Cultivated potato (Solanum tuberosum L.) is a highly heterozygous autotetraploid that presents challenges in genome analyses and breeding. Wild potato species serve as a resource for the introgression of important agronomic traits into cultivated potato. One key species is Solanum chacoense and the diploid, inbred clone M6, which is self-compatible and has desirable tuber market quality and disease resistance traits. Sequencing and assembly of the genome of the M6 clone of S. chacoense generated an assembly of 825 767 562 bp in 8260 scaffolds with an N50 scaffold size of 713 602 bp. Pseudomolecule construction anchored 508 Mb of the genome assembly into 12 chromosomes. Genome annotation yielded 49 124 high-confidence gene models representing 37 740 genes. Comparative analyses of the M6 genome with six other Solanaceae species revealed a core set of 158 367 Solanaceae genes and 1897 genes unique to three potato species. Analysis of single nucleotide polymorphisms across the M6 genome revealed enhanced residual heterozygosity on chromosomes 4, 8 and 9 relative to the other chromosomes. Access to the M6 genome provides a resource for identification of key genes for important agronomic traits and aids in genome-enabled development of inbred diploid potatoes with the potential to accelerate potato breeding.
Relative to homozygous diploids, the presence of multiple homologs or homeologs in polyploids affords greater tolerance to mutations that can impact genome evolution. In this study, we describe sequence and structural variation in the genomes of six accessions of cultivated potato (Solanum tuberosum L.), a vegetatively propagated autotetraploid and their impact on the transcriptome. Sequence diversity was high with a mean single nucleotide polymorphisms (SNP) rate of approximately 1 per 50 bases suggestive of high levels of allelic diversity. Additive gene expression was observed in leaves (3605 genes) and tubers (6156 genes) that contrasted the preferential allele expression of between 2180 and 3502 and 3367 and 5270 genes in the leaf and tuber transcriptome, respectively. Preferential allele expression was significantly associated with evolutionarily conserved genes suggesting selection of specific alleles of genes responsible for biological processes common to angiosperms during the breeding selection process. Copy number variation was rampant with between 16 098 and 18 921 genes in each cultivar exhibiting duplication or deletion. Copy number variable genes tended to be evolutionarily recent, lowly expressed, and enriched in genes that show increased expression in response to biotic and abiotic stress treatments suggestive of a role in adaptation. Gene copy number impacts on gene expression were detected with 528 genes having correlations between copy number and gene expression. Collectively, these data suggest that in addition to allelic variation of coding sequence, the heterogenous nature of the tetraploid potato genome contributes to a highly dynamic transcriptome impacted by allele preferential and copy number-dependent expression effects.
Background Meiotic recombination is the foundation for genetic variation in natural and artificial populations of eukaryotes. Although genetic maps have been developed for numerous plant species since the late 1980s, few of these maps have provided the necessary resolution needed to investigate the genomic and epigenomic features underlying meiotic crossovers. Results Using a whole genome sequencing-based approach, we developed two high-density reference-based haplotype maps using diploid potato clones as parents. The vast majority (81%) of meiotic crossovers were mapped to less than 5 kb. The fine-scale accuracy of crossover detection was validated by Sanger sequencing for a subset of ten crossover events. We demonstrate that crossovers reside in genomic regions of “open chromatin”, which were identified based on hypersensitivity to DNase I digestion and association with H3K4me3-modified nucleosomes. The genomic regions spanning crossovers were significantly enriched with the Stowaway family of miniature inverted-repeat transposable elements (MITEs). The occupancy of Stowaway elements in gene promoters is concomitant with an increase in recombination rate. A generalized linear model identified the presence of Stowaway elements as the third most important genomic or chromatin feature behind genes and open chromatin for predicting crossover formation over 10-kb windows. Conclusions Collectively, our results suggest that meiotic crossovers in potato are largely determined by the local chromatin status, marked by accessible chromatin, H3K4me3-modified nucleosomes, and the presence of Stowaway transposons.