Standardized microhaplotype databases for eight diverse crops enable multiallelic analyses, comparative genetics, and breeding decisions. Microhaplotypes are short genomic segments that contain multiple tightly linked variants, providing multi-allelic data that can enhance genetic resolution compared to traditional biallelic single nucleotide polymorphism (SNP) markers. Here, we present the creation and utilization of separate microhaplotype databases for eight crop species representing diverse genome sizes, ploidy levels, and breeding systems. We developed a standardized, species-agnostic pipeline for processing, filtering, and databasing microhaplotypes generated using the DArTag targeted genotyping platform. To enhance user accessibility, we developed a no-code, user-friendly application, HapApp, that uses an R Shiny front-end interface to allow breeders and researchers to add unique, standardized microhaplotype identities from raw DArTag reports and iteratively update the existing crop-specific database with the newly discovered microhaplotypes. Selected case studies with these databases highlight the operational advantages of microhaplotypes, especially for challenging, highly heterozygous, or polyploid species. They offer an informative alternative to traditional biallelic SNP analyses for resolving population structures and improving linkage map ordering. This integrated framework provides a reproducible and scalable foundation for managing and exploiting microhaplotype data in plant breeding and genetic research, enabling robust cross-project comparisons and facilitating trait discovery in both simple and complex crop genomes, while enabling comparative genomics and cross-species functional transfer that accelerates genetic gains across all crop species.
Background and Aims Agricultural phosphorus (P) management is essential for maintaining crop yields, but practices aimed at preventing deficiency often result in overuse and soil P saturation. The subsequent runoff contributes to eutrophication of nearby water bodies. Phytoremediation, using plants to absorb excess soil nutrients, offers a sustainable solution but requires engineering crops that hyperaccumulate P. Recent genetic studies show that mutations in the PHO2 -like genes, which help maintain phosphate (Pi) homeostasis, can result in substantial Pi hyperaccumulation in mutant plants. Methods In this study, mutated pho2 alleles from the six-haplotype mutant PhotM65-2 were introgressed into elite genotypes of tetraploid alfalfa ( Medicago sativa L.), resulting in seven F1 lines with mutations in PHO2 -like genes ( PHO2-B and PHO2-C ). The study characterized these F1 mutants, evaluated their Pi accumulation, and identified the most effective genetic combinations for Pi accumulation. Results Specifically, double and triple mutants with edits in PHO2-B and PHO2-C accumulated significantly more Pi than control genotypes under both high and low P conditions. Notably, some novel mutants maintained enhanced Pi content in their leaves without a biomass penalty. Conclusion This study demonstrates the successful introgression of CRISPR/Cas-edited PHO2-B/C alleles in alfalfa, identifying specific mutated pho2 allele combinations in F1 lines to create P-hyperaccumulator plants for remediating P-saturated soils and advancing agricultural sustainability.
Alfalfa (Medicago sativa L.) is a critical forage crop whose improvement depends on resolving the complex genetic architecture of agronomic traits. While genome-wide association studies (GWAS) effectively identify statistically associated markers, they often fail to distinguish direct genetic effectors from indirect or pleiotropic signals arising from linkage disequilibrium and population structure. Here, we present a causal graph based genomic discovery framework that integrates de-confounded feature screening with causal graph learning to infer directional dependency structures from observational genomic data. Using Double Machine Learning to control for confounding and the PC algorithm for structural learning, we construct directed acyclic graphs that distinguish Direct Parent SNPs (DPSs), representing local effectors within the Markov Blanket of a trait, from Upstream Hub SNPs (UHSs), representing pleiotropic regulators with broad network connectivity. Applied to four stem-related traits in alfalfa, the framework reduces genome-wide associations to compact, interpretable causal-consistent networks. Predictive validation demonstrates that DPSs consistently outperform both upstream UHSs and random controls, confirming their role as precise trait-specific biomarkers, while UHSs exhibit limited direct predictive power consistent with signal dilution along causal pathways. Together, these results demonstrate that causal graph learning can act as a biologically grounded regularizer for GWAS in polyploid crops, enabling principled marker prioritization and providing a structural foundation for future multi-omics integration.
Microhaplotypes are short genomic segments that contain multiple tightly linked variants, providing multiallelic data that can enhance genetic resolution compared to traditional biallelic single nucleotide polymorphism (SNP) markers. Here, we present the creation and utilization of separate microhaplotype databases for eight crop species representing diverse genome sizes, ploidy levels, and breeding systems. We developed a standardized, species-agnostic pipeline for processing, filtering, and databasing microhaplotypes generated using the DArTag targeted genotyping platform. To enhance user accessibility, we developed a no-code, user-friendly application, HapApp, that uses an R Shiny front-end interface to allow breeders and researchers to add unique, standardized microhaplotype identities from raw DArTag reports and iteratively update the existing crop-specific database with the newly discovered microhaplotypes. Comparative analyses of these databases highlighted the advantages of microhaplotypes in capturing greater allelic diversity, resolving fine-scale population structures, and improving linkage map construction. This integrated framework provides a reproducible and scalable foundation for managing and exploiting microhaplotype data in plant breeding and genetic research, enabling robust cross-project comparisons and facilitating trait discovery in both simple and complex crop genomes, while enabling comparative genomics and cross-species functional transfer that accelerates genetic gains across all crop species.
Plant genebanks contain large numbers of germplasm accessions that likely harbor useful alleles or genes absent in commercial plant breeding programs. Broadening the genetic base of commercial alfalfa germplasm with these valuable genetic variations can be achieved by screening the extensive genetic diversity in germplasm collections and enabling maximal recombination among selected genotypes. In this study, we assessed the genetic diversity and differentiation of germplasm pools selected in northern U.S. latitudes (USDA Plant Hardiness Zone 7 or below) originating from Eurasian germplasm. The germplasm evaluated included four BASE populations (C0) from different geographical origins (Central Asia, Northeastern Europe, Balkans-Turkey-Black Sea, and Siberia/Mongolia), 20 cycle-one populations (C1) derived from each of the four BASE populations selected across five locations in the U.S. and Canada, and four commercial cultivars. Using a panel of 3,000 Diversity Array Technologies (DArTag) marker loci, we retrieved 2,994 target SNPs and approximately 12,000 microhaplotypes. Microhaplotypes exhibited higher genetic diversity values than target SNPs. Principal component analysis and discriminant analysis of principal components revealed significant population structure among the alfalfa populations based on geographical origin, while the check cultivars formed a central cluster. Inbreeding coefficients (FIS) ranged from - 0.1 to 0.006, with 27 out of 28 populations showing negative FIS values, indicating an excess of heterozygotes. Interpopulation genetic distances were calculated using Rho pairwise distances (FST adapted for autotetraploid species) and analysis of molecular variance (AMOVA) parameters. All BASE populations showed lower Rho values compared to C1 populations and check cultivars. AMOVA revealed that most of the genetic diversity was among individuals within populations, especially in BASE populations (92.7%). This study demonstrates that individual plants in BASE populations possess high genetic diversity, low interpopulation distances, and minimal inbreeding, characteristics that are essential for base-broadening selection. The populations developed in this project serve as valuable sources of novel alleles for North American alfalfa breeding programs, offering breeders access to diverse, regionally adapted pools for improving various alfalfa traits.
Developing drought-resistant alfalfa (Medicago sativa L.) that maintains high biomass yield is a key breeding goal to enhance productivity in water-limited areas. In this study, 424 alfalfa breeding families were analyzed to identify molecular markers associated with biomass yield under drought stress and to predict high-merit plants. Biomass yield was measured from 18 harvests from 2020 to 2023 in a field trial with deficit irrigation. A total of 131 significant markers were associated with biomass yield, with 80 markers specifically linked to yield under drought stress; among these, 19 markers were associated with multiple harvests. Finally, genomic best linear unbiased prediction (GBLUP) was employed to obtain predictive accuracies (PAs) and genomic estimated breeding values (GEBVs). Removing low-informative SNPs [SNPs with p-values > 0.05 from the additive Genome-Wide Association (GWAS) model] for GBLUP increased PA by 47.3%. The high number of markers associated with yield under drought stress and the highest PA (0.9) represent a significant achievement in improving yield under drought stress in alfalfa.
Alfalfa (Medicago sativa L.), known as the queen of forages, is a versatile and valuable forage crop that holds significant importance in agriculture due to its myriad benefits for livestock. Ruminants benefit from alfalfa's digestible fiber and protein, contributing to improved feed efficiency and milk production. However, alfalfa protein is rapidly and extensively degraded in rumen, and it is a challenge to maximize the efficiency of the forage crude protein utilized as metabolizable protein by ruminant livestock. In this study, the phenotypic data of 14 traits related to forage digestibility were collected from 200 alfalfa accessions planted at three different locations for 2 years. The performance of these accessions showed dramatic variations by location, indicating that environmental factors play important roles in alfalfa digestibility. Twenty-two significant genetic markers associated with 12 traits related to forage digestibility were identified by genome-wide association study. Among them, seven markers were associated with more than one trait, although the significant markers varied by year and location. Putative candidate genes associated with these loci were also identified. The digestibility-related markers and associated genes identified in this study will help to better understand the genetic basis of forage digestibility and its interaction with environments. After validation, the closely linked markers and associated genes can be used for marker-assisted selection of alfalfa with improved forage quality.
Traditional alfalfa stem phenotyping is labor-intensive and susceptible to bias from subjective ratings. Computer vision and machine learning present a promising solution for objectively assessing stem morphology. This study proposed an AI-driven image analysis to replace manual phenotyping methods with high efficiency and accuracy. We developed a novel pipeline that combines YOLOv8n with Otsu’s thresholding and K-means clustering to identify the medoids of internal and external polygons of the stem, thereby quantifying stem traits using pixel-based morphometric masks. The approach achieved an F1 score of 0.91 in detecting and classifying hollow or solid stems across plots and genotypes. Further analysis measured stem area and the proportions of stem tissue versus hollow regions, generating traits like hollowness score and percentage of hollowness. These stem-level metrics provide novel, objective, and quantitative phenotypic measurements, supporting ongoing chemical digestibility analyses and enabling real-time, image-based digestibility assessments through a field-deployable mobile application.
Roots are essential for acquiring water and nutrients to sustain and support plant growth and anchorage. However, they have been studied less than the aboveground traits in phenotyping and plant breeding until recent decades. In modern times, root properties such as morphology and root system architecture (RSA) have been recognized as increasingly important traits for creating more and higher quality food in the “Second Green Revolution”. To address the paucity in RSA and other root research, new technologies are being investigated to fill the increasing demand to improve plants via root traits and overcome currently stagnated genetic progress in stable yields. Artificial intelligence (AI) is now a cutting-edge technology proving to be highly successful in many applications, such as crop science and genetic research to improve crop traits. A burgeoning field in crop science is the application of AI to high-resolution imagery in analyses that aim to answer questions related to crops and to better and more speedily breed desired plant traits such as RSA into new cultivars. This review is a synopsis concerning the origins, applications, challenges, and future directions of RSA research regarding image analyses using AI.
Background: Root system architecture (RSA) is of growing interest in implementing plant improvements with belowground root traits. Modern computing technology applied to images offers new pathways forward to plant trait improvements and selection through RSA analysis (using images to discern/classify root types and traits). However, a major stumbling block to image-based RSA phenotyping is image label noise, which reduces the accuracies of models that take images as direct inputs. To address the label noise problem, this study utilized an artificial intelligence model capable of classifying the RSA of alfalfa (Medicago sativa L.) directly from images and coupled it with downstream label improvement methods. Images were compared with different model outputs with manual root classifications, and confident machine learning (CL) and reactive machine learning (RL) methods were tested to minimize the effects of subjective labeling to improve labeling and prediction accuracies. Results: The CL algorithm modestly improved the Random Forest model’s overall prediction accuracy of the Minnesota dataset (1%) while larger gains in accuracy were observed with the ResNet-18 model results. The ResNet-18 cross-population prediction accuracy was improved (~8% to 13%) with CL compared to the original/preprocessed datasets. Training and testing data combinations with the highest accuracies (86%) resulted from the CL- and/or RL-corrected datasets for predicting taproot RSAs. Similarly, the highest accuracies achieved for the intermediate RSA class resulted from corrected data combinations. The highest overall accuracy (~75%) using the ResNet-18 model involved CL on a pooled dataset containing images from both sample locations. Conclusions: ResNet-18 DNN prediction accuracies of alfalfa RSA image labels are increased when CL and RL are employed. By increasing the dataset to reduce overfitting while concurrently finding and correcting image label errors, it is demonstrated here that accuracy increases by as much as ~11% to 13% can be achieved with semi-automated, computer-assisted preprocessing and data cleaning (CL/RL).
Alfalfa biomass can be fractionated into leaf and stem components. Leaves comprise a protein-rich and highly digestible portion of biomass for ruminant animals, while stems constitute a high fiber and less digestible fraction, representing 50 to 70% of the biomass. However, little attention has focused on stem-related traits, which are a key aspect in improving the nutritional value and intake potential of alfalfa. This study aimed to identify molecular markers associated with four morphological traits in a panel of five populations of alfalfa generated over two cycles of divergent selection based on 16-h and 96-h in vitro neutral detergent fiber digestibility in stems. Phenotypic traits of stem color, presence of stem pith cells, winter standability, and winter injury were modeled using univariate and multivariate spatial mixed linear models (MLM), and the predicted values were used as response variables in genome-wide association studies (GWAS). The alfalfa panel was genotyped using a 3K DArTag SNP markers for the evaluation of the genetic structure and GWAS. Principal component and population structure analyses revealed differentiations between populations selected for high- and low-digestibility. Thirteen molecular markers were significantly associated with stem traits using either univariate or multivariate MLM. Additionally, support vector machine (SVM) and random forest (RF) algorithms were implemented to determine marker importance scores for stem traits and validate the GWAS results. The top-ranked markers from SVM and RF aligned with GWAS findings for solid stem pith, winter standability, and winter injury. Additionally, SVM identified additional markers with high variable importance for solid stem pith and winter injury. Most molecular markers were located in coding regions. These markers can facilitate marker-assisted selection to expedite breeding programs to increase winter hardiness or stem palatability.
Alfalfa is widely recognized as an important forage crop. To understand the morphological characteristics and genetic basis of seed morphology in alfalfa, we screened 318 Medicago spp., including 244 Medicago sativa subsp. sativa (alfalfa) and 23 other Medicago spp., for seed area size, length, width, length-to-width ratio, perimeter, circularity, the distance between the intersection of length & width (IS) and center of gravity (CG), and seed darkness & red-green-blue (RGB) intensities. The results revealed phenotypic diversity and correlations among the tested accessions. Based on the phenotypic data of M. sativa subsp. sativa, a genome-wide association study (GWAS) was conducted using single nucleotide polymorphisms (SNPs) called against the Medicago truncatula genome. Genes in proximity to associated markers were detected, including CPR1, MON1, a PPR protein, and Wun1(threshold of 1E-04). Machine learning models were utilized to validate GWAS, and identify additional marker-trait associations for potentially complex traits. Marker S7_33375673, upstream of Wun1, was the most important predictor variable for red color intensity and highly important for brightness. Fifty-two markers were identified in coding regions. Along with strong correlations observed between seed morphology traits, these genes will facilitate the process of understanding the genetic basis of seed morphology in Medicago spp.
The bacterial stem blight of alfalfa (Medicago sativa L.), first reported in the United States in 1904, has emerged recently as a serious disease problem in the western states. The causal agent, Pseudomonas syringae pv. syringae, promotes frost damage and disease that can reduce first harvest yields by 50%. Resistant cultivars and an understanding of host-pathogen interactions are lacking in this pathosystem. With the goal of identifying DNA markers associated with disease resistance, we developed biparental F1 mapping populations using plants from the cultivar ZG9830. Leaflets of plants in the mapping populations were inoculated with a bacterial suspension using a needleless syringe and scored for disease symptoms. Bacterial populations were measured by culture plating and using a quantitative PCR assay. Surprisingly, leaflets with few to no symptoms had bacterial loads similar to leaflets with severe disease symptoms, indicating that plants without symptoms were tolerant to the bacterium. Genotyping-by-sequencing identified 11 significant SNP markers associated with the tolerance phenotype. This is the first study to identify DNA markers associated with tolerance to P. syringae. These results provide insight into host responses and provide markers that can be used in alfalfa breeding programs to develop improved cultivars to manage the bacterial stem blight of alfalfa.
Winter injury of alfalfa [Medicago sativa (L.)] in the northern United States decreases its economic and ecosystem benefits. Therefore, continued improvement in alfalfa cultivar winter survival (WS) is crucial for sustaining the productivity of this perennial crop. The North American Alfalfa Improvement Conference (NAAIC) standard test for WS recommends measuring the WS of spaced plants established in rows the previous spring. Measurement of WS of alfalfa grown in sward plots used by plant breeders would increase data collection and better reflect the potential for WS when grown in production fields. We conducted trials at seven location-year environments spanning from Wisconsin to South Dakota in the northern United States. These trials involved six check cultivars and followed protocols from the NAAIC standard test. The objectives were to determine (1) if WS and biomass yield assessment from sward plots were similar to those from the standard spaced planted row ratings and (2) if location-dependent environmental conditions affected the usefulness of alternative approaches for measuring WS. Estimation of WS using spaced plants and sward measurements was highly correlated, while correlations between the WS of the spaced planted rows and biomass yields were less. The number of locations required for spaced and sward plantings to determine cultivar differences was at least two, with four replications per location. Measuring WS from swards can enhance data collection and its relevance to on-farm alfalfa production, as sward plots serve a dual purpose by allowing both WS testing and evaluation of yield, making them a practical choice in comparison to the exclusive use of spaced plants in rows for WS testing. Availability of sward-plot WS descriptions of alfalfa cultivars will enhance decision making by producers.
The low digestibility of fiber in alfalfa (Medicago sativa L.) limits dry matter intake and energy availability in ruminant animal production systems. Previously, alfalfa plants were identified for low or high rapid (16 h) and low or high potential (96 h) in vitro neutral detergent fiber digestibility (IVNDFD) of plant stems. Here, two cycles of bidirectional selection for 16 h and 96 h IVNDFD were carried out. The resulting populations were evaluated for total herbage, percentage of stems to total biomass, IVNDFD, neutral detergent fiber (NDF), and acid detergent lignin as a proportion of NDF (ADL/NDF) at three maturity stages. Within these populations, 96 h IVNDFD was highly heritable (h(2) = 0.71), while 16 h IVNDFD had lower heritability (h(2) = 0.46). Selection for high IVNDFD reduced NDF and ADL/NDF in plant stems at the late flowering and green pod maturity stages and reduced seasonal variability in stem digestibility but did not alter the percentage of stems. Stability analyses across 12 harvest environments found that selection for high IVNDFD had little effect on environmental stability of the trait compared to the unselected population. Thus, selection for stem IVNDFD was a highly effective strategy for developing alfalfa populations with improved nutritional quality without changing the percentage of stems to total biomass.
EDITORIAL article Front. Plant Sci., 27 November 2023Sec. Technical Advances in Plant Science Volume 14 - 2023 | https://doi.org/10.3389/fpls.2023.1331918
Active breeding programs specifically for root system architecture (RSA) phenotypes remain rare; however, breeding for branch and taproot types in the perennial crop alfalfa is ongoing. Phenotyping in this and other crops for active RSA breeding has mostly used visual scoring of specific traits or subjective classification into different root types. While image-based methods have been developed, translation to applied breeding is limited. This research is aimed at developing and comparing image-based RSA phenotyping methods using machine and deep learning algorithms for objective classification of 617 root images from mature alfalfa plants collected from the field to support the ongoing breeding efforts. Our results show that unsupervised machine learning tends to incorrectly classify roots into a normal distribution with most lines predicted as the intermediate root type. Encouragingly, random forest and TensorFlow-based neural networks can classify the root types into branch-type, taproot-type, and an intermediate taproot-branch type with 86% accuracy. With image augmentation, the prediction accuracy was improved to 97%. Coupling the predicted root type with its prediction probability will give breeders a confidence level for better decisions to advance the best and exclude the worst lines from their breeding program. This machine and deep learning approach enables accurate classification of the RSA phenotypes for genomic breeding of climate-resilient alfalfa.
Categorical data derived from qualitative classifications or countable quantitative data are common in biological scientific work and crop breeding. Categorical data analyses are important for drawing correct inferences from experiments. However, categorical data can introduce unique issues in data analysis. This paper discusses common problems arising from categorical variable analysis and modeling, demonstrates the issues or risks of misapplying analysis, and suggests approaches to address data analysis challenges using two data sets from alfalfa breeding programs. For each data set, we present several analysis methods, e.g., simple t-test, analysis of variance (ANOVA), split plot analysis, generalized linear model (glm), generalized linear mixed model (glmm) using R with R markdown, and with the standard statistical analysis software SAS/JMP. The goal is to demonstrate good analysis practices for categorical data by comparing the potential ‘bad’ analyses with better ones, avoiding too much reliance on reaching a significant p-value of 0.05, and navigating the morass of ever-increasing numbers of potential R functions. The three main aspects of this research focus on choosing the right data distribution to use, using the correct error terms for hypothesis test p-values including the right type of sum of the squares (Type I, II, and III), and proper statistical models for categorical data analysis. Our results show the importance of good statistical analysis practice to help agronomists, breeders, and other researchers apply appropriate statistical approaches to draw more accurate conclusions from their data.