
The ability to efficiently discriminate genotypes is a critical step in genomics-assisted breeding, population genomics, biodiversity studies, traceability along food chains, and germplasm management. However, identifying the minimal and most informative subset of SNPs capable of uniquely distinguishing a large set of individuals remains a computationally challenging task. Here, we present SNPoptimizer, a user-friendly Shiny application that uses a genetic algorithm–based framework to optimally select discriminatory SNPs from large-scale genotyping datasets. By leveraging the evolutionary principles of selection, mutation, and crossover, SNPoptimizer iteratively identifies compact SNP panels that maximize genotype resolution. The application supports HapMap-formatted and VCF genotype files and includes an optional second-round optimization for resolving putative duplicates. We benchmarked SNPoptimizer across three independent datasets, including a tomato diversity panel, 820 Cauliflower genotypes, and a soybean diversity panel comprising 30 million variants across 1,511 samples. Across the three datasets, panels of 17–22 SNPs yielded R-VDP values ranging from 0.8744 to 0.9973, with complete discrimination obtained in Dataset III, demonstrating robust performance across different datasets. Cross-tool comparisons revealed complementary trade-offs among discriminatory power, panel size, runtime, and run-to-run reliability. SNPoptimizer provides a flexible solution for researchers seeking to reduce genotyping costs while maintaining high discriminative power.
TaDEP1, a heterotrimeric G protein γ subunit gene, was characterized as a positive regulator of grain size and thousand-grain weight (TGW) in bread wheat. Association analysis in 492 global wheat accessions and validation in a recombinant inbred line population identified a missense SNP (C/G) in TaDEP1-5 A that causes a cysteine to tryptophan substitution, with the TaDEP1-5 A-C allele significantly associated with increased grain width and TGW. This favorable allele has been preferentially selected during breeding, increasing from 15.10
Soil salinity is a major constraint on rice (Oryza sativa L.) production, particularly during the yield-determining reproductive stage. Utilizing a high-throughput RGB platform, we non-destructively phenotyped a diverse panel of 294 rice accessions over two consecutive years. By extracting 60 dynamic image-based traits (i-traits) reflecting canopy architecture and stay-green capacity, and four seed-setting rate-related traits, our genome-wide association study (GWAS) identified 95 significant loci, 35.8
Sustainable crop improvement is urgently needed to ensure global food security, particularly for developing and densely populated countries. The integration of artificial intelligence (AI) and machine learning (ML) into crop science tri typing is reshaping the conventional agriculture practices into an era of high-throughput phenotyping (HTPP) data-driven modern agriculture. AI tools accelerate data generation, mining, imputation, storage, transfer, and optimal decision-making within agricultural systems. AI tools are paving the way for modern plant breeding strategies by uncovering genetic variability and bridging the genotype-to-phenotype (G2P) gap, thus enabling the future of predictive breeding. Plant genetic gains or phenotype (P), by and large, depend on the genotype (G), environment (E), and their interaction (GEI). This review will provide a comprehensive overview of the historical background, current status, and prospects for integrating AI and ML tools in agricultural tri-typing, encompassing genotyping, phenotyping, and envirotyping. We explore AI-driven tools for genome analysis, HTPP platforms, and environmental data integration, emphasizing how these technologies overcome persistent bottlenecks in predictive breeding. Furthermore, this review will offer the reader key insight into modern trends, including the paradigm shift in phenomics patent filings, global distribution of HTP phenomics facilities, the publications volume and related research over the last two decades, and individual institutions currently leading or prospectively will lead the world in plant phenomics. Similar to plant phenotyping, we also try to address the integration and application of AI/ML algorithms in plant genotyping and envirotyping.
Recurrent crop disease outbreaks linked to global warming pose challenges to sustainable food production. Conventional plant breeding techniques may become less effective at addressing these threats, as improving disease resistance often causes yield reduction. Amid these challenges, CRISPR/Cas-based gene editing offers targeted and tractable solutions. This review synthesizes recent approaches to uncoupling immunity from productivity in cereals. We show that susceptibility (S) gene disruption can provide resistance without activating costly defense mechanisms. We also discuss the generation of new alleles through targeted modifications that mitigate autoimmunity-associated fitness costs. Here, we propose CRISPR-mediated de-moonlighting, an approach for decoupling multifunctional protein activities. Multiplex editing of minor resistance loci, especially in polyploids such as wheat, offers long-lasting, broad-spectrum protection. These strategies converge on the manipulation of canonical moonlighting proteins, multifunctional signaling hubs, and pleiotropic regulators, which serve as regulatory nodes that link development and immunity. CRISPR-mediated precision modification of these regulators can fine-tune the growth-defense balance. Combining these approaches with systems biology, AI-driven design and advanced breeding pipelines can help develop high-yielding, disease-resistant cereals for sustainable agriculture.
The brown planthopper (BPH, Nilaparvata lugens) and rice blast (Magnaporthe oryzae) are major constraints on rice production, causing extensive yield losses worldwide. Pyramiding multiple resistance genes is a key strategy for achieving durable and broad-spectrum protection against these pests and pathogens. We pyramided 13 resistance genes/QTLs (Bph10, Bph14, Bph15, Bph17, Pi2, Bph20, Bph21, Bph6, QBph3, QBph4, Bph18, Bph9, and Bph24) in the 9311 backgrounds through pairwise crossing, generating 30 pyramided lines, and true hybrids were identified by marker-assisted selection (MAS). Seedling-stage BPH resistance of the pyramided families was evaluated using the Standard Seedbox Screening Test (SSST), whereas leaf blast resistance of the selected lines was assessed by field evaluation of leaf blast. Most pyramided lines exhibited stronger resistance than their corresponding single-gene parental lines, and some also showed improved agronomic traits. These lines provide valuable breeding material for developing “Green Super Rice” with combined BPH and blast resistance and high, stable yield.
Flowering time and vegetative vigour are central adaptive traits in chickpea, shaping the balance between stress escape, biomass accumulation, and yield. Major phenology loci have been identified in this species, including ELF3a, CaLG3b and a cluster of FT homologues on chromosome 3 (CaLG3a), but their individual and interactive contributions to field adaptation remain unclear. Here, we developed near-isogenic lines (NILs) to dissect these loci under controlled conditions and across eight field environments in Southern and Western Australia. In a phytotron, NILs revealed strong photoperiod dependence: ELF3a suppressed flowering under short days, while early alleles at CaLG3b and the FT cluster promoted flowering under long days, with additive effects across loci. In the field, allelic substitutions consistently shifted phenology by a few days or nodes but rarely conferred direct yield benefits. Yield responses were variable and environment-dependent, with one NIL pair exhibiting a consistent penalty in environments with yields above 255 g m− 2. In contrast, canopy traits showed clearer divergence between NIL pairs, aligning with pleiotropic effects of the FT cluster. These results highlight the utility of NILs for attributing modest but adaptive phenological shifts to defined genomic regions, and suggest that the primary breeding value of these loci lies in fine-tuning crop development to local stress patterns rather than delivering robust yield gains. Stacking early and late alleles may offer breeders a genetic tool to broaden adaptive flexibility, while yield improvement will likely require integration with physiological traits related to resource capture, efficiency and higher reproductive allocation.
Soybean seed length (SL) is a key agronomic trait affecting both seed yield and appearance quality and is controlled by a complex interplay of multiple genes and environmental factors. Therefore, identifying major loci or genes associated with SL is critical for soybean molecular breeding. Phenotypic measurements of SL were conducted across three environments using a four‑way recombinant inbred line (FW‑RIL) population of 144 lines, along with their parental accessions.Linkage analysis identified 26 quantitative trait loci (QTLs) for SL. Candidate genes were predicted within the qSL-5-6 interval, which was consistently detected across multiple environments and analytical methods. Functional annotation, sequence variation, and haplotype analyses identified the tightly linked genes Glyma.05G238900 and Glyma.05G239100 as promising candidates for soybean SL that require further functional validation.
Despite their substantial therapeutic value, medicinal plants have undergone limited genetic improvement through breeding because of the scarcity of expert breeders. Moreover, quantifying bioactive compounds is expensive. Genomic selection, which leverages genome-wide markers to predict breeding values and assemble favorable alleles, offers a practical way to unlock latent genetic potential. As a model case, we evaluated genomic selection in red perilla (Perilla frutescens). Building on previous work, we implemented a cross-selection strategy that prioritized segregation variance by selecting crosses based on predicted additive genotypic values of the progeny, and evaluated its effectiveness through actual crossing experiments targeting three key medicinal compounds. Several genomic selection-based crosses produced G₂ progeny with high perillaldehyde and rosmarinic acid contents and broad phenotypic variation under the tested field condition. The best individual derived from the genomic selection-based crosses exhibited nearly twofold higher levels of two target compounds relative to the existing cultivar ‘Sekiho’. Because only one phenotypically selected cross was included, this experiment was not designed to provide a statistical comparison between genomic and phenotypic selection. Instead, it serves as an empirical case study demonstrating the application of progeny-based cross selection in an underutilized medicinal plant breeding program.
Spray-induced gene silencing (SIGS) is an emerging, time-efficient, and eco-friendly strategy for crop trait manipulation. Heat stress severely threatens global maize production; however, whether SIGS can be effectively applied to improve maize heat tolerance remains largely unexplored. The transcription factor ZmHSF20, a class B2a heat shock factor, has been identified as a negative regulator of maize thermotolerance that represses the ZmHSF4–ZmCESA2 module (Li et al. 2024), yet whether its silencing via a non-transgenic RNA-based approach can confer heat protection is unknown. Here, we designed a specific artificial small RNA atsRNA-ZmHSF20 targeting ZmHSF20 and evaluated its potential in improving maize thermotolerance via the SIGS approach. Compared with untreated plants, foliar application of atsRNA-ZmHSF20 significantly enhanced maize heat tolerance, leading to a remarkable increase in survival rate from 7.94
Flowering and maturity time in soybean are highly sensitive to photoperiod and influence adaptation and yield potential. In Ethiopia, soybean yield remains stagnant at about 2.4 t/ha. Precocious flowering and maturity induced by tropical short-day environments are the primary contributors to the low yields. This study dissects the genetic basis of flowering and maturity under Ethiopian conditions. A diverse panel of 323 accessions, from 13 U.S. maturity groups (MG 000 to X), was evaluated across four mid-altitude production environments. A multi-locus genome-wide association study was conducted using the mrMLM framework by integrating phenotypic data with 38,281 high-quality SNPs. Analysis of variance for the phenotypic traits revealed significant differences among the genotypes. Population structure analysis clustered the accessions into four groups, and linkage disequilibrium decayed at approximately 229 kb. Eight stable quantitative trait nucleotides (QTNs) with LOD > 3 were identified across five chromosomes; six co-localized with known loci and two were novel. Several candidate genes involved in controlling flowering and maturity act as regulators of phase transitions and circadian rhythms. For instance, Glyma.13g274900 encodes an SBP-type protein Glyma.13g274300 encodes a NAM domain, Glyma.13g273300 encodes a PHD-type domain, Glyma.12g009100, a cytochrome b5 domain, and Glyma.19g206100 encodes an auxin response factor. Allelic effect analysis demonstrated that homozygous alleles significantly delay flowering and extend maturity. The identified QTNs should be considered for advanced molecular breeding applications aimed at developing high-yielding soybean lines suited to Ethiopia’s unique tropical mid-altitude environments. Future research should validate these candidate genes and their gene interactions.
Bipolaris species are major fungal pathogens responsible for foliar blights in cereals, notably B. sorokiniana, B. maydis, and B. oryzae. Despite their substantial negative impact in crops, molecular markers such as Simple Sequence Repeats (SSR) have remained scarce for this genus, limiting progress in population genetics, epidemiological surveillance, virulence which indirectly could affect breeding for disease resistance. Here, we present BipolarisSSRDb, an indigenously developed, web-based SSR marker database for the Bipolaris genus. The database integrates 12,382 SSR loci mined from genome assemblies of three key species, along with locus-specific primers, motif details, and genomic context. The architecture of the database supports intuitive browsing, visualization, and primer retrieval. Experimental validation confirmed specificity and polymorphism for many of the selected markers. BipolarisSSRDb enables high-throughput marker selection for diagnostics including strain differentiation, and population studies. To the best of our knowledge, BipolarisSSRDb represents the first dedicated SSR marker database developed exclusively for Bipolaris species, providing a genome-wide SSR loci and primer information, for genomics research in India. The database is freely accessible at www.biopolarisdb.fwh.is .
Blackleg (Phoma, Leptosphaeria maculans) is one of the most important diseases affecting rapeseed (Brassica napus, 2n = AACC) worldwide. Although a number of major blackleg resistance genes have been identified in B. napus, finding novel sources of resistance is critical in dealing with this rapidly evolving fungal pathogen. Black mustard (Brassica nigra, 2n = BB) is a highly resistant wild relative species for which at least one blackleg resistance gene has been described (Rlm10 on chromosome B4), but which has been rarely investigated as a source of blackleg resistance for introgression breeding (probably because it is self-incompatible, heterozygous, difficult to hybridize with B. napus and until recently lacked the genetic and genomic resources available for the Brassica A- and C-genome species). We produced allohexaploid hybrids (2n = AABBCC) between B. napus and B. nigra followed by two generations of backcrossing to develop a BC2F1 population segregating for inheritance of specific B-genome chromosomes from B. nigra. This population was phenotyped for blackleg resistance using a cotyledon test with blackleg isolate “JN2” (AvrLm4-7, 5–9, 6, 8, 10 A-B, 11, S-Lep2) and genotyped with B-genome chromosome-specific markers. Resistance segregated with the presence of chromosome B2 (p < 0.0001). Future work will target the identification or production of introgression lines, and recovery of this genetic locus from the B genome in a rapeseed (2n = AACC) genomic background. Our work highlights the potential of B. nigra and Brassica wild relatives as sources of novel biotic stress resistances.
The flavonoid content of peaches has gradually decreased as peach cultivars have been domesticated as part of their evolutionary process, and this reduction is driven by the fact that some flavonoids impart a bitter taste to the fruit. This has resulted in a gradual decrease in flavonoid content as the peach fruit matures during the developmental stages across most current peach cultivars. In this study, we first analyzed changes in flavonoid content among three peach varieties during different fruit developmental stages. We found that 75 days after full blooming was the stage when flavonoid content exhibited the most significant increase. We also performed transcriptome analysis on the fruit samples, screening out 77 differentially expressed genes between the two varieties with different flavonoid accumulation patterns. Combined with the results of previous association analysis, a glycosyltransferase gene, PpUGT73C3, was identified to be a candidate involving flavonoid synthesis. Transient transformation assays demonstrated that both overexpression and silencing of PpUGT73C3 significantly alter flavonoid accumulation, confirming that PpUGT73C3 positively regulates flavonoid biosynthesis in peach fruits. Additionally, PpZFP6, a gene encoding a zinc finger protein and mapping to chromosome 3, can influence the glycosylation modification of flavonoids by regulating the promoter activity of PpUGT73C3, ultimately controlling flavonoid accumulation in peach fruit. This study preliminarily clarified the molecular regulatory network of PpUGT73C3 in peach flavonoid metabolism, laying a research foundation for the breeding of functional fruit cultivars and the development of relevant regulatory strategies.
Mitogen-activated protein kinase kinase kinase kinase (MAP4K) cascade components are versatile signaling integrators in plants, yet their agronomic relevance in crops remains largely unexplored. Here, we systematically characterized 11 maize MAP4K genes through genome-wide identification, expression profiling, candidate-gene association mapping, and network-based analyses. Genomic analysis suggested that the current composition of the ZmMAP4K family was shaped by both whole-genome duplication/segmental duplication and dispersed duplication, accompanied by substantial divergence in spatiotemporal and stress-responsive expression patterns. Association analyses across 23 agronomic and stress-related traits revealed significant functional diversification within the family. Notably, ZmMAP4K2 emerged as a multi-stress-associated locus, as different variants within this gene were associated with salt tolerance and southern leaf blight (SLB) resistance, respectively. Furthermore, natural variation in ZmMAP4K4 was strongly associated with drought survival, whereas a favorable allele of ZmMAP4K10, characterized by higher transcriptional activity, was associated with increased kernel thickness and other yield traits. Haplotype analyses identified favorable allelic combinations for stress-related traits, while multi-locus evaluation of stress-associated variants revealed genotype combinations with enhanced resilience to multiple stresses and no evident penalties on major yield-related traits. Finally, predictive regulatory and protein–protein interaction networks provided putative molecular contexts for key ZmMAP4Ks in both stress adaptation and yield formation. Together, this study highlights the agronomic potential of the maize MAP4K family, offering specific candidate loci and allelic combinations for future functional validation and trait pyramiding.
The peanut seed coat serves as a vital protective layer, but the occurrence of fine cracks breaches this integrity, thereby impairing visual quality and elevating the risk of pathogen infection and postharvest mycotoxin contamination. However, knowledge on genetic factors underlying seed coat cracking (SCC) tolerance remains limited. In this study, a population of 521 recombinant inbred line (RIL) was used to identify genetic regions associated with SCC tolerance across three environments. Quantitative trait locus (QTL) mapping identified two stable QTLs on chromosomes 7 and 16, qSCCA07 and qSCCA16, explaining 3.91
Flower color is a crucial morphological trait and evolutionary marker in soybean (Glycine max). However, its complex genetic architecture and domestication history remain incompletely understood. Here, we performed a genome-wide association study (GWAS) on 741 soybean accessions, comprising 411 cultivated and 330 wild accessions, using three distinct statistical models (GLM, MLM, and FarmCPU) with over 7.95 million SNPs, integrated with selective sweep and transcriptomic analyses. We identified two significant quantitative trait loci (QTLs) governing floral pigmentation. The major locus, qFC13, explaining 48.63
Phytophthora sojae and Pythium spp. cause important root diseases of soybeans resulting in millions of dollars of loss for producers every year. Our previous studies on mapping Rps (Resistance to Ph. sojae) genes and quantitative disease resistance loci (QDRL) towards Ph. sojae and Pythium species, respectively, identified key regions for resistance to these diseases on chromosomes 3, 6, 7, 8, 13, 14 and 18. The development of R gene enrichment and sequencing (RenSeq) have allowed for the expedited discovery of new nucleotide-binding, leucine-rich repeat (NLR)-encoding genes. This study focused on exploring the co-localization of Rps loci or QDRL regions with single nucleotide polymorphism (SNP) markers derived from sequences identified from a RenSeq study of germplasm varying for resistance to Ph. sojae or Pythium species. Another aspect of this study was to saturate the QDRL regions on chromosomes 6, 8 and 14 for resistance to multiple Pythium species utilizing conventional molecular and SNP markers developed from RenSeq analysis. Nineteen RenSeq SNP markers were selected for evaluation in six RIL populations related to resistance to Ph. sojae and/or Pythium as allele-specific KASP markers. Most RenSeq SNP markers co-localized within the Rps loci or QDRL regions in these RIL populations. The development of RenSeq SNPs and their co-localization with Rps loci or QDRL regions indicate their potential use in designing resistance gene-specific markers and their application for marker-assisted selection in soybean resistance breeding programs.
Cabbage is severely threatened by clubroot disease worldwide. To improve the clubroot resistance of cabbage, clubroot resistance from ‘9LQ18-5’ was introgressed into the susceptible parent ‘F416’ of commercial cabbage variety ‘Xiyuan 4’. During the backcrossing program, clubroot resistant individuals were selected by CCR1 and CCR2 markers, whereas the background of backcrossing progenies was detected by whole genome re-sequencing in BC3F1 individuals. Among them, 21 − 9 with the highest recurrent parent genomic recovery (83.93