Genome-wide association studies (GWAS) use statistical models to correlate single nucleotide polymorphisms (SNPs) to a phenotype of interest. This scan of the entire genome identifies regions of association with a phenotype, but due to linkage disequilibrium (LD), GWAS on their own cannot identify single genes responsible for phenotypic variation. Rather, fine-mapping of GWAS regions is required, necessitating the use of additional tools and software. With the introduction of more pangenomic resources in a number of crops (Guo et al. 2025; Hufford et al. 2021), the fidelity of these fine-mapping efforts is growing, presenting the opportunity to leverage new information about allelic variation towards gene discovery (Shi et al. 2023; Della Coletta et al. 2021). Panvar is a tool developed to integrate existing software and resources to perform GWAS and fine-mapping in one seamless step. For each identified GWAS peak, panvaR outputs information about LD and SNP effect prediction for each SNP and by layering locations of nearby genes, creates a refined list of possible candidate genes. We have implemented Panvar as an R package, "panvaR", which runs the analysis functions, creates interactive and static visualizations, and outputs results tables. This tool seeks to bridge the gap between GWAS and gene speeding up an important step of quantitative genetic studies.
Although the green revolution adapted a handful of crops to homogeneous and high-input industrialized agriculture, much of the global population still relies on the local production of variable crop cultivars by low-input smallholder farms. This diversity of unhomogenized crops1, like that of the grain and bioenergy crop sorghum2-5, offers raw materials for genetic gain and cultivar improvement. However, breeding efforts can be constrained by highly specialized traits and breeding targets6. Here, to bridge this diversity, we constructed a 33-member pangenome reference and a diversity panel across 1,984 cultivars and landraces. We leveraged these resources to explore the complex interplay among historical contingency, ongoing adaptation and previously uncharacterized structural diversity. Specifically, our analyses conclusively demonstrated multiple nested and deeply diverged structural variants in the domestication gene SHATTERING1, which distinguish the previously established multicentric origin of sorghum. We then applied landscape genomics to reveal how gene flow and secondary contact created the complex genetic mosaic in contemporary breeding networks. As proof of concept for pangenome-accelerated trait discovery, we connected biosynthetic gene cluster structural variation to phenotypic leaf concentration of the cyanogenic glucoside dhurrin. Combined, these approaches will accelerate breeding and trait discovery and provide a framework for similar applications in other crops.
While the green revolution adapted a handful of crops to homogenous and high-input industrialized agriculture, much of the global population still relies on local food production from low-input smallholder farms that grow highly variable crop cultivars. The high diversity of the grain and bioenergy crop sorghum [1][1]–[4][2], and many other crops that were not homogenized during the green revolution [5][3], not only provides the raw materials for breeders to make substantial gains in cultivar improvement, but also constrains breeding efforts due to highly specialized locally adapted plant phenotypes [6][4]. Here, we construct a 33-member pangenome and identify trait-associated variants in 1,988 cultivars and landraces. We then apply these resources to explore the complex interplay between historical contingency, ongoing adaptation, and the potential for future gains through climate-aware genome-enabled breeding. Specifically, our analyses conclusively demonstrate that multiple nested, deeply diverged, and previously uncharacterized structural variants in the domestication gene SHATTERING1 distinguish the previously established multicentric origin of sorghum. We then apply landscape genomics tests to reveal how gene flow, adaptation, and secondary contact created the complex genetic mosaic in current global breeding networks. Further analysis of climate-gene associations highlights candidate loci underlying adaptation, including the biosynthetic gene cluster for the cyanogenic glucoside dhurrin. Combined, the pangenome-informed variants developed here will enable both trait discovery and subsequent marker assays to accelerate breeding and provide a framework for similar applications in other diverse and non-model crops. ### Competing Interest Statement The authors have declared no competing interest. United States Department of Energy, DE-AC02-05CH11231 [1]: #ref-1 [2]: #ref-4 [3]: #ref-5 [4]: #ref-6
Fonio (Digitaria exilis), known as the "grain of life," is a vital food security crop grown by smallholder farmers across West Africa. Renowned for its short maturity period of 6-8 weeks, fonio provides essential harvests during periods of food scarcity. Its nutritional value, drought resilience, and ability to thrive on marginal soils make it a crucial staple in the region. However, fonio’s yield, averaging less than 0.5 metric tons per hectare, remains significantly lower than major cereals, due to minimal breeding efforts aimed at enhancing agronomic traits. In this study, we generated a high-quality chromosome-resolved, tetraploid (2n=4x=36) genome of a fonio line (PI 349688) available through the United States Germplasm Resource Information Network (GRIN). Time-of-day (TOD) expression profiling revealed sub-genome dominance in pathways associated with key agronomic traits that were also identified in a selective sweep analysis across a panel of 90 fonio accessions that originated from 21 different villages over five regions of Senegal representing eight distinct ethnic groups. Leveraging phenotyping across 14 key agronomic traits, we identified genotype-trait associations for plant height, tillering and panicle yield with genes involved in the circadian clock, light signaling and plant architecture. We demonstrated that fonio is amenable to transformation and gene editing, as evidenced by successful CRISPR-Cas9 modification of the Green Revolution gene REDUCED HEIGHT 1 (RHT1), resulting in lines with increased height, vegetative area, and enhanced tillering. Collectively, the genomic resources developed here pave new avenues for molecular breeding and trait discovery, laying a robust foundation for accelerating fonio improvement to bolster food security, climate resilience, and sustainable agriculture in West Africa.
Water availability is a major determinant of crop production, and rising temperatures from climate change are leading to more extreme droughts. To combat the effects of climate change on crop yields, we need to develop varieties that are more tolerant to water-limited conditions. We aimed to determine how diverse crop types (winter/spring oilseed, tuberous, and leafy) of the allopolyploid Brassica napus, a species that contains the economically important rapeseed oilseed crop, respond to prolonged water limitation. We exposed plants to an 80% reduction in water and assessed growth and color on a high-throughput phenotyping system over 4 weeks and ended the experiment with tissue collection for a time course transcriptomic study. We found an overall reduction in growth across cultivars but to varying degrees. Diel transcriptome analyses revealed significant accession-specific changes in time-of-day regulation of photosynthesis, carbohydrate metabolism, and sulfur metabolism. Interestingly, there was extensive variation in which homoeologs from the two parental subgenomes responded to water limitation across crop types that could be due to differences in regulatory regions in these allopolyploid lines. Follow-up experiments on select cultivars confirmed that plants maintained photosynthetic health during the prolonged water limitation while slowing growth. In two cultivars examined, we found significant time of day changes in levels of glucosinolates, sulfur- and nitrogen -rich specialized metabolites, consistent with the diel transcriptomic responses. These results suggest that these lines are adjusting their sulfur and nitrogen stores under water-limited conditions through distinct time of day regulation.
Accurate and efficient estimation of crop biophysical traits, such as leaf chlorophyll concentrations (LCC) and average leaf angle (ALA), is an important bridge between intelligent crop breeding and precision agriculture. While Unmanned Aerial Vehicle (UAV)-based hyperspectral sensors and advanced machine learning models offer high-throughput solutions, collecting sufficient ground truth data for machine learning training can be challenging, leading to models that lack generalizability for practical uses. This study proposes a transfer learning based dual stream neural network (DSNN) called PROSAIL-Net, which leverages the knowledge gained from PROSAIL simulation and improves the estimation of corn LCC and ALA from UAV-borne hyperspectral images. In addition to hyperspectral data, the DSNN also includes solar-sensor geometry data, which was automatically extracted from a cross-grid UAV flight. The hyperspectral branch in the DSNN was also tested with multi-layer perceptron (MLP), long short-term memory (LSTM), gated recurrent unit (GRU), and 1D convolutional neural network (CNN) architectures. The results suggest that the 1D CNN architecture exhibits superior performance compared to MLP, LSTM, and GRU networks when used in the spectral branch of DSNN. PROSAIL-Net outperforms all other modeling scenarios in predicting LCC (R2 0.66, NRMSE 8.81%) and ALA (R2 0.57, NRMSE 24.32%) and the use of multi-angular UAV observations significantly improves the prediction accuracy of both LCC (R2 improved from 0.52 to 0.66) and ALA (R2 improved from 0.35 to 0.57). This study highlights the importance of utilizing large amounts of PROSAIL-simulated data in conjunction with transfer learning and multi-angular UAV observations in precision agriculture.
IntroductionSorghum bicolor is a promising cellulosic feedstock crop for bioenergy due to its high biomass yields. However, early growth phases of sorghum are sensitive to cold stress, limiting its planting in temperate environments. Cold adaptability is crucial for cultivating bioenergy and grain sorghum at higher latitudes and elevations, or for extending the growing season. Identifying genes and alleles that enhance biomass accumulation under early cold stress can lead to improved sorghum varieties through breeding or genetic engineering.MethodsWe conducted image-based phenotyping on 369 accessions from the sorghum Bioenergy Association Panel (BAP) in a controlled environment with early cold treatment. The BAP includes diverse accessions with dense genotyping and varied racial, geographical, and phenotypic backgrounds. Daily, non-destructive imaging allowed temporal analysis of growth-related traits and water use efficiency (WUE). A genome-wide association study (GWAS) was performed to identify genomic intervals and genes associated with cold stress response.ResultsThe GWAS identified transient quantitative trait loci (QTL) strongly associated with growth-related traits, enabling an exploration of the genetic basis of cold stress response at different developmental stages. This analysis of daily growth traits, rather than endpoint traits, revealed early transient QTL predictive of final phenotypes. The study identified both known and novel candidate genes associated with growth-related traits and temporal responses to cold stress.DiscussionThe identified QTL and candidate genes contribute to understanding the genetic mechanisms underlying sorghum's response to cold stress. These findings can inform breeding and genetic engineering strategies to develop sorghum varieties with improved biomass yields and resilience to cold, facilitating earlier planting, extended growing seasons, and cultivation at higher latitudes and elevations.
Gene functional descriptions offer a crucial line of evidence for candidate genes underlying trait variation. Conversely, plant responses to environmental cues represent important resources to decipher gene function and subsequently provide molecular targets for plant improvement through gene editing. However, biological roles of large proportions of genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, an update able data resource consisting of transcript abundance assays spanning 18 diverse species. To integrate across these diverse genotypes, we anal yzed e xpression pr ofiles, b uilt gene c lusters that exhibited tissue / condition specific expression, and tested for transcriptional response to environmental queues. We discovered extensive phylogenetically constrained and condition-specific expression profiles f or genes without an y previously documented functional annotation. Such conserved expression patterns and tightly co-expressed gene clusters let us assign expression derived additional biological information to 64 495 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas ( https://plantgeneatlas.jgi.doe.gov ) and Phytozome ( https://phytozome .jgi.doe .gov/), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentiall y e xpressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.
The demand for agricultural production is becoming more challenging as climate change increases global temperature and the frequency of extreme weather events. This study examines the phenotypic variation of 149 accessions of Brachypodium distachyon under drought, heat, and the combination of stresses. Heat alone causes the largest amounts of tissue damage while the combination of stresses causes the largest decrease in biomass compared to other treatments. Notably, Bd21-0, the reference line for B. distachyon, did not have robust growth under stress conditions, especially the heat and combined drought and heat treatments. The climate of origin was significantly associated with B. distachyon responses to the assessed stress conditions. Additionally, a GWAS found loci associated with changes in plant height and the amount of damaged tissue under stress. Some of these SNPs were closely located to genes known to be involved in responses to abiotic stresses and point to potential causative loci in plant stress response. However, SNPs found to be significantly associated with a response to heat or drought individually are not also significantly associated with the combination of stresses. This, with the phenotypic data, suggests that the effects of these abiotic stresses are not simply additive, and the responses to the combined stresses differ from drought and heat alone.
BACKGROUND:Daylength is a key seasonal cue for animals and plants. In cereals, photoperiodic responses are a major adaptive trait, and alleles of clock genes such as PHOTOPERIOD1 (PPD1) and EARLY FLOWERING3 (ELF3) have been selected for in adapting barley and wheat to northern latitudes. How monocot plants sense photoperiod and integrate this information into growth and development is not well understood.RESULTS:We find that phytochrome C (PHYC) is essential for flowering in Brachypodium distachyon. Conversely, ELF3 acts as a floral repressor and elf3 mutants display a constitutive long day phenotype and transcriptome. We find that ELF3 and PHYC occur in a common complex. ELF3 associates with the promoters of a number of conserved regulators of flowering, including PPD1 and VRN1. Consistent with observations in barley, we are able to show that PPD1 overexpression accelerates flowering in short days and is necessary for rapid flowering in response to long days. PHYC is in the active Pfr state at the end of the day, but we observe it undergoes dark reversion over the course of the night.CONCLUSIONS:We propose that PHYC acts as a molecular timer and communicates information on night-length to the circadian clock via ELF3.
Irrigation of crops accounts for a significant portion of fresh water consumption. In order to utilize this resource more efficiently, it is necessary to engineer crops that can more efficiently use water. Water use efficiency, defined as the ratio of plant growth to water used, is a complex property of plants affected by many different factors. Despite this complexity, genetic variability has been able to be identified in a number of different crops. The C4 model species Setaria viridis remains under-studied in this regard and consequently we sought to identify promising genetic loci contributing to variation in water use efficiency. In order to accomplish this goal we leveraged the high-throughput phenotyping platform at the Donald Danforth Plant Science center to grow S. viridis in well-watered and water-limited conditions. This automated system enables strict control of watering regimes as well as measures of plant traits extracted from photographs using computer vision. Combining these two sets of data allows for direct measurement of whole-plant water-use efficiency on a daily basis which was used as a response variable in a genome wide association study. Significant associations were found for water-use efficiency and related traits. These loci were then prioritized further by pooling information across each day of an experiment and across multiple experiments to zero in on the most likely locations of genes responsible for driving water-use efficiency in S. viridis.
Near-earth hyperspectral big data present both huge opportunities and challenges for spurring developments in agriculture and high-throughput plant phenotyping and breeding. In this article, we present data-driven approaches to address the calibration challenges for utilizing near-earth hyperspectral data for agriculture. A data-driven, fully automated calibration workflow that includes a suite of robust algorithms for radiometric calibration, bidirectional reflectance distribution function (BRDF) correction and reflectance normalization, soil and shadow masking, and image quality assessments was developed. An empirical method that utilizes predetermined models between camera photon counts (digital numbers) and downwelling irradiance measurements for each spectral band was established to perform radiometric calibration. A kernel-driven semiempirical BRDF correction method based on the Ross Thick-Li Sparse (RTLS) model was used to normalize the data for both changes in solar elevation and sensor view angle differences attributed to pixel location within the field of view. Following rigorous radiometric and BRDF corrections, novel rule-based methods were developed to conduct automatic soil removal; and a newly proposed approach was used for image quality assessment; additionally, shadow masking and plot-level feature extraction were carried out. Our results show that the automated calibration, processing, storage, and analysis pipeline developed in this work can effectively handle massive amounts of hyperspectral data and address the urgent challenges related to the production of sustainable bioenergy and food crops, targeting methods to accelerate plant breeding for improving yield and biomass traits.
Osmotic adjustment (OA) is a major component of drought resistance in crops. The genetic basis of OA in wheat and other crops remains largely unknown. In this study, 248 field-grown durum wheat elite accessions grown under well-watered conditions, underwent a progressively severe drought treatment started at heading. Leaf samples were collected at heading and 17 days later. The following traits were considered: flowering time (FT), leaf relative water content (RWC), osmotic potential (ψs), OA, chlorophyll content (SPAD), and leaf rolling (LR). The high variability (3.89-fold) in OA among drought-stressed accessions resulted in high repeatability of the trait (h2 = 72.3%). Notably, a high positive correlation (r = 0.78) between OA and RWC was found under severe drought conditions. A genome-wide association study (GWAS) revealed 15 significant QTLs (Quantitative Trait Loci) for OA (global R2 = 63.6%), as well as eight major QTL hotspots/clusters on chromosome arms 1BL, 2BL, 4AL, 5AL, 6AL, 6BL, and 7BS, where a higher OA capacity was positively associated with RWC and/or SPAD, and negatively with LR, indicating a beneficial effect of OA on the water status of the plant. The comparative analysis with the results of 15 previous field trials conducted under varying water regimes showed concurrent effects of five OA QTL cluster hotspots on normalized difference vegetation index (NDVI), thousand-kernel weight (TKW), and/or grain yield (GY). Gene content analysis of the cluster regions revealed the presence of several candidate genes, including bidirectional sugar transporter SWEET, rhomboid-like protein, and S-adenosyl-L-methionine-dependent methyltransferases superfamily protein, as well as DREB1. Our results support OA as a valuable proxy for marker-assisted selection (MAS) aimed at enhancing drought resistance in wheat.
Rapid environmental change can lead to population extinction or evolutionary rescue. The global staple crop sorghum (Sorghum bicolor) has recently been threatened by a global outbreak of an aggressive new biotype of sugarcane aphid (SCA; Melanaphis sacchari). We characterized genomic signatures of adaptation in a Haitian breeding population that had rapidly adapted to SCA infestation, conducting evolutionary population genomics analyses on 296 Haitian lines versus 767 global accessions. Genome scans and geographic analyses suggest that SCA adaptation has been conferred by a globally rare East African allele of RMES1, which spread to breeding programs in Africa, Asia, and the Americas. De novo genome sequencing revealed potential causative variants at RMES1. Markers developed from the RMES1 sweep predicted resistance in eight independent commercial and public breeding programs. These findings demonstrate the value of evolutionary genomics to develop adaptive trait technology and highlight the benefits of global germplasm exchange to facilitate evolutionary rescue.
Calculating solar-sensor zenith and azimuth angles for hyperspectral images collected by UAVs are important in terms of conducting bi-directional reflectance function (BRDF) correction or radiative transfer modeling-based applications in remote sensing. These applications are even more necessary to perform high-throughput phenotyping and precision agriculture tasks. This study demonstrates an automated Python framework that can calculate the solar-sensor zenith and azimuth angles for a push-broom hyperspectral camera equipped in a UAV. First, the hyperspectral images were radiometrically and geometrically corrected. Second, the high-precision Global Navigation Satellite System (GNSS) and Inertial Measurement Unit (IMU) data for the flight path was extracted and corresponding UAV points for each pixel were identified. Finally, the angles were calculated using spherical trigonometry and linear algebra. The results show that the solar zenith angle (SZA) and solar azimuth angle (SAA) calculated by our method provided higher precision angular values compared to other available tools. The viewing zenith angle (VZA) was lower near the flight path and higher near the edge of the images. The viewing azimuth angle (VAA) pattern showed higher values to the left and lower values to the right side of the flight line. The methods described in this study is easily reproducible to other study areas and applications.
ABSTRACTGene functional descriptions, which are typically derived from sequence similarity to experimentally validated genes in a handful of model species, offer a crucial line of evidence when searching for candidate genes that underlie trait variation. Plant responses to environmental cues, including gene expression regulatory variation, represent important resources for understanding gene function and crucial targets for plant improvement through gene editing and other biotechnologies. However, even after years of effort and numerous large-scale functional characterization studies, biological roles of large proportions of protein coding genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, a public and updateable data resource consisting of transcript abundance assays from 2,090 samples derived from 604 tissues or conditions across 18 diverse species. We integrated across these diverse conditions and genotypes by analyzing expression profiles, building gene clusters that exhibited tissue/condition specific expression, and testing for transcriptional modulation in response to environmental queues. For example, we discovered extensive phylogenetically constrained and condition-specific expression profiles across many gene families and genes without any functional annotation. Such conserved expression patterns and other tightly co-expressed gene clusters let us assign expression derived functional descriptions to 64,620 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas (https://plantgeneatlas.jgi.doe.gov) and Phytozome (https://phytozome-next.jgi.doe.gov), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentially expressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.
We explore the use of deep convolutional neural networks (CNNs) trained on overhead imagery of biomass sorghum to ascertain the relationship between single nucleotide polymorphisms (SNPs), or groups of related SNPs, and the phenotypes they control. We consider both CNNs trained explicitly on the classification task of predicting whether an image shows a plant with a reference or alternate version of various SNPs as well as CNNs trained to create data-driven features based on learning features so that images from the same plot are more similar than images from different plots, and then using the features this network learns for genetic marker classification. We characterize how efficient both approaches are at predicting the presence or absence of a genetic markers, and visualize what parts of the images are most important for those predictions. We find that the data-driven approaches give somewhat higher prediction performance, but have visualizations that are harder to interpret; and we give suggestions of potential future machine learning research and discuss the possibilities of using this approach to uncover unknown genotype × phenotype relationships.
An indoor wireless fixed camera network was developed for an efficient, cost-effective method of extracting informative plant phenotypes in a controlled greenhouse environment. Deployed at the Donald Danforth Plant Science Center (DDPSC), this fixed camera platform implements rapid and automated plant phenotyping. The platform uses low-cost Raspberry Pi computers and digital cameras to monitor aboveground morphological and developmental plant phenotypes. The Raspberry Pi is a readily programmable, credit card-sized computer board with remote accessibility. A standard camera module connects to the Raspberry Pi computer board and generates eight-megapixel resolution images. With a fixed array, or "bramble," of Raspberry Pi computer boards and camera modules placed strategically in a greenhouse, we can capture automated, high-resolution images for 3D reconstructions of individual plants on timescales ranging from minutes to hours, capturing temporal changes in plant phenotypes.
Sorghum and maize share a close evolutionary history that can be explored through comparative genomics1,2. To perform a large-scale comparison of the genomic variation between these two species, we analysed ~13 million variants identified from whole-genome resequencing of 499 sorghum lines together with 25 million variants previously identified in 1,218 maize lines. Deleterious mutations in both species were prevalent in pericentromeric regions, enriched in non-syntenic genes and present at low allele frequencies. A comparison of deleterious burden between sorghum and maize revealed that sorghum, in contrast to maize, departed from the domestication-cost hypothesis that predicts a higher deleterious burden among domesticates compared with wild lines. Additionally, sorghum and maize population genetic summary statistics were used to predict a gene deleterious index with an accuracy greater than 0.5. This research represents a key step towards understanding the evolutionary dynamics of deleterious variants in sorghum and provides a comparative genomics framework to start prioritizing these variants for removal through genome editing and breeding.
Photosynthesis is one of the most important biological reactions on earth providing oxygen and food for humanity. As global populations rise and arable land decreases, crops need to become more efficient at photosynthetic processes, particularly utilizing absorbed light energy. Chlorophyll fluorescence imaging is a rapid, non-destructive measurement that can provide information on the efficiency of the light-dependent reactions and carbon reactions of photosynthesis. Over the years chlorophyll fluorescence imaging systems have been developed and improved to capture two critical measurements: minimum fluorescence (F 0) and maximum fluorescence (F M). These systems have primarily been utilized in controlled chamber or greenhouse settings focused at the single leaf or small plant level. To improve plant photosynthesis, fluorescence imaging data needs to be obtained from field-grown plants to capture canopy spatial effects. Previously developed software to extract F 0 and F M from controlled, leaf level images do not capture the complexity of the light-dependent reactions from field-grown plants. New software is needed that accounts for the canopy spatial effects from images of field-grown plants. FLIP: fluorescence imaging pipeline, was designed specifically for the TERR-REF field scanalyzer located at the University of Arizona's Maricopa Agricultural Center located in Maricopa, Arizona but could be adapted for other field deployed fluorescence imaging systems. FLIP utilizes open source tools to convert binary images, apply a multi-threshold to extract the fluorescence data from the plant canopy, calculate photosynthetic efficiency, and assign those values to the appropriate experimental plot.