Objective:Seed weight is a key factor in soybean yield and value, but its genetic basis and environmental stability are not fully understood. Despite many QTL studies, there's a lack of integration between bi-parental linkage mapping and diverse germplasm association analysis. We hypothesized that combining high-resolution QTL mapping in recombinant inbred lines with GWAS in natural populations could identify both population-specific and broadly segregating seed weight loci, aiding in candidate gene discovery for breeding. Methods:We integrated biparental QTL mapping with genome-wide association studies (GWAS) to comprehensively dissect the genetics of hundred-seed weight (HSW). A recombinant inbred line population of 325 F2:5 lines from Qihuang 34 × Dongsheng 16 was phenotyped across three environments and genotyped using SLAF-seq, generating a high-density genetic map with 6,297 SNP markers spanning 2,945.26 cM (0.47 cM resolution). Simultaneously, 348 diverse soybean accessions underwent whole-genome resequencing (10× coverage), yielding 1,882,531 SNPs for association analysis across two years. Results:QTL mapping identified 11 significant loci explaining 2.47-8.59% of phenotypic variance, with broad-sense heritability of 0.78. The major-effect QTL qHSW-19-4 (44.84-44.85 Mb, LOD = 9.72) demonstrated unprecedented 11.4 kb mapping precision. GWAS independently detected six genome-wide significant associations (P < 1 × 10-8), including a stable chromosome 19 peak at 45.28 Mb (P = 2.06 × 10-²³) explaining 15.3-18.7% of variance. Remarkably, this GWAS signal co-localized within 580 kb of qHSW-19-4, providing robust cross-population validation of chromosome 19 as a major seed weight regulatory region. Functional analysis of 44 candidate genes, validated by quantitative RT-PCR across seed developmental stages, identified four high-priority candidates: Glyma.19G195400 (cell wall invertase, 2.7-fold upregulation in large-seeded parent, r = 0.68 with HSW), Glyma.19G194300 (PEBP/Dt1 family protein), Glyma.19G193400 (bZIP transcription factor), and Glyma.06G095100 (Myb DNA-binding domain). Novelty and conclusions:This first integrated QTL-GWAS analysis for soybean seed weight reveals both major-effect loci and polygenic architecture, providing validated molecular markers and candidate genes for breeding programs targeting yield improvement.
Shade stress affects soybean yield in intercropping, but the molecular basis of cultivar-specific tolerance is unclear. We analyzed shade-tolerant (Guru) and sensitive (Heinong 53) soybeans under 30% and 70% shade using transcriptomic, physiological, and biochemical methods. Moderate shade (30%) initially promoted growth in Heinong 53, while 70% shade restricted growth in both. Guru maintained better physiology, with higher chlorophyll (2.8 vs. 2.1 mg/g), photosynthesis, and stable stomatal conductance. RNA-seq identified 2596-7841 differentially expressed genes (DEGs), with severe shade causing more changes. Down-regulated DEGs are linked to development and shade response; up-regulated DEGs are involved in metabolism, including glucosinolate biosynthesis and protein export. Core shade responses included 279 up- and 388 down-regulated DEGs across treatments. Shade tolerance involved metabolic reprogramming: Guru had higher sucrose content (30.2 vs. 13.8 mg/g) and sucrose synthase activity and increased nitrogen enzyme activity. Antioxidant enzymes showed cultivar-specific strategies, with Guru having higher peroxidase and lower oxidative stress markers. Gene-trait analysis linked 37 DEGs to photosynthesis and 28 to transpiration, indicating water use regulation. Key genes included calcium-dependent kinases and histone deacetylases. Overall, shade tolerance involves maintaining photosynthesis, metabolic shifts favoring carbs, and stress responses, guiding development of tolerant cultivars and understanding plant adaptation in intercropping systems.
IntroductionNitrogen form and concentration are key environmental regulators that mediate symbiotic nitrogen fixation and root development in legumes.MethodsTo understand the metabolic and molecular mechanisms underlying the effects of distinct nitrogen sources (nitrate and ammonium) on soybean nodulation and root development, this study evaluated root and nodulation phenotypes, and their corresponding transcriptional and metabolomic responses under different concentrations of NH₄Cl or KNO₃.ResultsResults showed that both high concentrations of NH₄Cl and KNO₃ significantly suppressed nodulation and promoted root growth, with nitrate exerting a stronger effect than ammonium. Metabolomic analysis revealed that ammonium treatment enhanced nitrogen assimilation and primary metabolism while suppressing symbiosis-related flavonoids. Nitrate specifically activated chemical defense pathways and inhibited parts of central carbon metabolism. Integrated multi-omics analysis indicated that the nitrogen sources differentially regulated key genes and metabolites involved in nitrogen metabolism, flavonoid/isoflavonoid biosynthesis, and arginine metabolism, leading to distinct metabolic fluxes.DiscussionOur results demonstrate that soybean perceives different nitrogen forms to orchestrate a metabolic trade-off between autonomous growth, defense, and symbiosis, thereby providing new insights into the mechanistic basis of nitrogen-form adaptation in legumes.
Agricultural biological breeding is rapidly accelerating its paradigm shift from traditional, empirical breeding toward a new era of Artificial Intelligence (AI) design, characterized by the deep integration of Biotechnology (BT) and Information Technology (IT). This transition serves as a crucial strategic lever for ensuring national food security and propelling the green transformation of the agricultural sector in China. Soybean, functioning as a vital multi-purpose crop for food, edible oil, and livestock feed, has long confronted significant bottlenecks in the enhancement of its unit yield and broad environmental adaptability. These enduring challenges fundamentally stem from the highly complex polygenic basis governing essential agronomic traits-such as yield potential, nutritional quality, and stress tolerance. Furthermore, these traits are heavily influenced by profound genotype-by-environment (G & times;E) interactions and are strictly constrained by the imperative need for the efficient utilization of resources, including water, soil, fertilizers, and pesticides. Consequently, traditional breeding methodologies struggle to achieve synergistic breakthroughs across multiple desirable traits while simultaneously maintaining yield stability. To address these limitations, this comprehensive review systematically synthesizes the frontier progress of BT and IT integration within soybean research. The paper is structured around a whole-chain trajectory: "multi-omics trait dissection-engineered germplasm creation-intelligent G & times;E prediction-the integrative application of superior varieties, optimal agronomic practices, and smart equipment". In the realm of Biotechnology (BT), we outline the unprecedented value of pan-genomics and multi-dimensional omics in deciphering core genetic regulatory networks, alongside pivotal technological breakthroughs such as gene editing, tissue-culture-free transformation methodologies, and the structural reconstruction of metabolic pathways via synthetic biology. Concurrently, we examine IT-empowered pathways- specifically encompassing AI, high-throughput phenomics, advanced envirotyping, and smart agriculture. Through this dual-lens approach, we elucidate the core transformative role of BT+IT integration in dissecting complex agronomic traits, intelligently formulating novel germplasm, facilitating highly efficient G & times;E predictions, and enabling the synergistic deployment of elite cultivars with optimized management and intelligent hardware. Furthermore, this article meticulously traces the ongoing technological evolution of AI-driven G & times;E prediction models, whole-genome intelligent crossing and prediction (WPICP) pipelines, and agricultural digital twin systems. It summarizes the integrative innovation paradigms surrounding green prevention and control, highly efficient nutrient utilization, and intelligent agronomic equipment. Ultimately, the paper highlights that China has already established distinctive, cumulative advantages in the specialized domains of soybean stress tolerance, biological nitrogen fixation, and intelligent factory-based breeding. Looking forward, the critical pathways requiring imminent breakthroughs include generative intelligent design propelled by large foundational models in life sciences, the systematic reconstruction of complex metabolic networks utilizing synthetic biology, and the multidimensional synergistic optimization of genotype-by-environment-by-management (G & times;E & times;M) interactions. Collectively, these advancements will provide robust theoretical foundations and technological support for the precision breeding and sustainable green productivity enhancement of soybeans.
Seed weight is a key agronomic trait determining soybean yield and quality, yet only a few of genes regulating this trait have been functionally characterized to date. In this study, we identified 155 homologous genes in the soybean genome through BLAST searches using 78 functionally validated rice grain weight-related genes as queries. Haplotype analysis prioritized 40 candidate genes exhibiting significant differences in seed weight between haplotypes. To further refine the candidate list, we integrated haplotype frequency analysis, expression-trait association mapping, and tissue-specific expression profiling, ultimately delineating eight key genes. Given the established role of ubiquitination in seed development, we focused on homologs of OsUBP15 and identified three candidate genes, GmUBP5, GmUBP11, and GmUBP33, that exhibited significant haplotype-dependent variation in seed weight. Subcellular localization assays confirmed their nuclear localization. Haplotype frequency analysis revealed that the superior haplotypes of these genes have been preferentially retained during modern breeding and are widely distributed across major soybean-producing regions. Leveraging non-synonymous SNP variants, we developed and validated robust KASP markers that efficiently discriminate germplasm with contrasting seed weight phenotypes. Collectively, our study provides not only high-confidence genetic targets and actionable molecular markers but also insights into pyramiding breeding strategies for improving seed weight in soybean.
Cultivated soybean (Glycine max [L.] Merr.) is a major global source of vegetable protein and edible oil, yet its production is severely constrained by soybean cyst nematode (SCN, Heterodera glycines), with race 3 being the most prevalent and destructive pathotype in China. Zhongpin 03-5373 (ZP03) is an elite Chinese soybean line characterized by stable resistance to SCN3 (race 3, HG type 0) and superior agronomic performance. Here, we report a near-gapless de novo genome assembly of ZP03 generated using PacBio Revio (HiFi) long-read sequencing combined with Hi-C chromatin interaction mapping. The assembled genome spans 1,067.68 Mb with a Contig N50 of 35.36 Mb, anchoring 95.24% of sequences onto 20 chromosomes. Comparative genomic analyses revealed an extensive structural variation (SV) landscape associated with SCN3 resistance in ZP03. A 393-bp deletion-type SV located in the promoter region of GmSNAP11 was identified and was associated with increased gene expression following SCN3 infection. This allelic variation may contribute to differential transcriptional regulation under pathogen challenge. Using an SV-based genetic map constructed from a ZP03×Zhonghuang 13 (ZH13) recombinant inbred line population, we further identified a novel QTL qSCN3-16 contributing to SCN3 resistance, which explaining 7.51% of the phenotypic variance (LOD=3.42). Within this locus, Glyma.16G045900 (SYP16), encoding a t-SNARE protein, was prioritized as the candidate gene. Haplotype analysis across 2,214 diverse soybean accessions demonstrated that a favorable promoter SV in SYP16 is significantly associated with enhanced SCN3 resistance. Together, these results highlight the critical role of structural variation in shaping soybean immune responses and provide high-resolution genomic resources and functional markers, specifically the novel QTL qSCN3-16, to support precision breeding of SCN3-resistant soybean cultivars.
Seed weight is a key determinant of soybean yield; however, its underlying genetic regulation remains poorly understood. Here, besides the previously reported positive seed weight regulator GmST05/GmSW5, which encodes a protein homologue of Arabidopsis Mother of TFL1 and FT (MFT), we identify GmSW19 that exert negative effect on seed weight. GmSW19 belongs to the basic leucine zipper (bZIP) transcription factor family; it encoded bZIP transcription factor can represses GmSW5 expression. We further discover that the glycogen synthase kinase 3 (GSK3)-like kinase GmSK21 physically interacts with and phosphorylates GmSW19. A single nucleotide polymorphism (A to C), resulting in a T to P substitution in GmSW19, which alters its phosphorylation by GmSK21 and consequently its protein abundance. Functional analyses reveal that the GmSW19A variant represses GmSW5 and seed weight stronger than that of GmSW19C. Population analyses show that the heavy-seed allele GmSW19C has not been fully utilized in modern soybean breeding. These findings elucidate the genetic components and their possible interaction in determining seed weight and highlight the potential for enhancing yield in soybean.
Salt stress is an abiotic constraint on crop production, impairing growth while reducing yield and quality. As a source of edible oil and protein, soybean is particularly vulnerable to salt stress during emergence and seedling establishment, so the genetic basis of salt tolerance bears directly on productive cultivation. Here, a natural population of 256 soybean accessions was phenotyped for salt tolerance at emergence and seedling stages and genotyped using the Zhongdouxin No.1 (ZDX1) single nucleotide polymorphism (SNP) array. Genome-wide association analysis identified 60 salt tolerance-associated SNPs consolidated into 19 quantitative trait loci (QTL), 16 associated with emergence-stage and 3 with seedling-stage salt tolerance. Five QTLs mapped to chromosomal regions harboring previously reported stress-tolerance genes, including GmSALT3, approximately 192 kb from qSTG-SSB-03. Four putative candidate genes were prioritized: Glyma.10G040000, encoding a glutathione S-transferase (GST), was identified at the emergence stage, while Glyma.10G148700, Glyma.10G149200, and Glyma.10G149600, encoding a calmodulin-binding protein (CaM), drought-induced protein 19 (Di19), and a protein phosphatase 2C (PP2C), respectively, were identified at the seedling stage, all with reported roles in salt-stress responses.
[Objective]The"stay-green"trait can prolong the effective photosynthesis duration in soybeans and increase dry matter accumulation,thereby holding significant potential for improving yield.Mining stay-green related QTL and elucidating their molecular mechanisms can provide a theoretical basis and technical support for enhancing soybean yield.[Method]A soybean nested association mapping population was evaluated for stay-green traits across multiple environments.Genome-wide association study was conducted using genotyping data.Candidate genes were screened via SNP variation,tissue-specific expression,and functional annotation analyses,haplotype,promoter cis-acting elements,and protein structure prediction analyses were performed to characterize the candidate genes.Additionally,the application effect of genomic selection for the stay-green trait was evaluated.[Result]Six significant QTL intervals were co-localized on chromosomes 3,4,5,and 16.Among these,qSG5-1(Chr.5:41600128..42273303,613.18 kb)was repeatedly mapped across multiple environments and represents a novel QTL for stay-green regulation in soybean.Linkage disequilibrium analysis allocated two significantly associated regions within qSG5-1:qSG5-1.1(Chr.5:41798499..41996276,197.78 kb)and qSG5-1.2(Chr.5:41996989..42273303,276.32 kb),containing 29 and 37 genes,respectively.SNP variation analysis identified 53 genes containing variants that cause nonsynonymous mutations,alternative splicing,stopgain,or stoploss.Of these,eight genes were transcriptionally active in stems and leaves.Functional annotation suggested that Glyma.05G245200 and Glyma.05G247900 were involved in protein folding and oxidative metabolism,respectively,which highlights they might regulate cell cycle,growth metabolism,and nutrient remobilization during senescence.Besides,two major haplotypes of these genes exhibited highly significant phenotypic differences as Glyma.05G245200 harbored nonsynonymous mutations which changed C617T into A206V and C44T into P15L,and caused subtle alterations in its protein structure.Likewise,Glyma.05G247900 also contained a nonsynonymous mutation which changed A275G into D92G that did not alter its protein conformation.Analysis of cis-acting elements revealed that the presence of light and abscisic acid(ABA)-responsive elements in their promoters hints they might regulate soybean growth,senescence,and the stay-green trait by participating in light and hormonal signaling.These genes may serve as candidate genes for soybean stay-green and the prediction accuracy of genome-wide selection for stay-green across different marker sets ranged from 0.27 to 0.36.[Conclusion]This study identified a novel QTL,qSG5-1,and two candidate genes,Glyma.05G245200 and Glyma.05G247900,associated with the stay-green trait in soybean.
Stem canker caused by Diaporthe caulivora (syn. D. phaseolorum var. caulivora) was first reported in North America in the mid-20th century and is referred to as northern stem canker. Since it was reported in Uruguay in 2013, it is considered one of the most important diseases in the crop, causing up to 24
Introduction:Seed shape and hundred-seed weight (HSW) are yield-related traits, but the interrelationships between them and a systematic genetic analysis are lacking. Methods:In this study, phenotypic evaluation of seed shape-related traits and HSW was performed in a recombinant inbred line (RIL) population derived from the cross between "ZD41" and "ZYD02878." All seed shape-related traits (SW, seed width; SL, seed length; ST, seed thickness; SP, seed perimeter; SA, seed area; SLW, seed length-to-width ratio; SLT, seed length-to-thickness ratio; SWT, seed width to thickness ratio) were evaluated. Broad-sense heritability was calculated. Correlation analysis, path analysis, QTL mapping (using ICIM, EMMAX, and TASSEL), and genomic selection (using rrBLUP) were performed. Results:All seed shape-related traits exhibited approximately normal distributions in the RIL population. The broad-sense heritability of seed shape-related traits ranged from 0.67 to 0.90, indicating that additive genes played a major role in expressing all studied traits. Correlation analysis revealed that SW showed the highest positive correlation with HSW (r = 0.86-0.87, P < 0.001) among the three measured seed shape-related traits (SL, SW, and ST), which agrees with the finding that SW was the only trait showing a positive direct effect on HSW in path analysis. A total of 98 QTL corresponding to 59 unique regions were identified by ICIM, EMMAX, and TASSEL. Three genes, Glyma.02G255900, Glyma.02G270100, and Glyma.02G275200, were identified as candidates for seed shape regulation. Genomic selection for seed shape-related traits using rrBLUP revealed that the prediction accuracy of SL, SW, ST, SP, and SA reached 0.65, 0.68, 0.65, 0.67, and 0.68, respectively, based on genome-wide SNP markers. Discussion:This study provided a systematic dissection of seed shape-related traits in terms of both phenotypic and genetic aspects, which laid a solid foundation for gene cloning and could be used in breeding programs in the future.
Pod number is an important factor influencing soybean yield. In this study, a recombinant inbred line population derived from the cross between improved cultivar (Glycine max) and wild soybean (Glycine soja) was subjected to multi-environment phenotypic evaluation for pod number. Quantitative trait locus (QTL) mapping associated with pod number was performed utilizing both linkage analysis (LA) and genome-wide association study (GWAS) by integrating the phenotypic and genotypic data. The results showed that a total of 16 QTLs associated with pod number, notably, overlapping genomic intervals were observed between qPN10-1 (by linkage analysis) and qPN10-3 (by EMMAX), as well as between qPN11-2 (by LA) and qPN11-3 (by EMMAX), and these intervals were regarded as major-effect Quantitative trait locus regions. Furthermore, two candidate genes (Glyma.10G206300, and Glyma.10G207500) encode a basic Helix-Loop-Helix (bHLH) transcription factor and an armadillo (ARM)-repeat protein respectively, as determined through linkage disequilibrium analysis, SNP variant analysis, haplotype analysis, and gene functional annotation. These two candidates are possibly involved in the regulation of flower number, flowering time, and pod number (siliques number in Arabidopsis thaliana). These findings pave the way for gene cloning related to pod number, and provide novel insights into the genetic architecture underlying pod number variation in soybean.
Soybean is a primary global vegetable oil source, yet modern South American cultivars often exhibit superior oil content compared to those from China, the center of origin. Elucidating the genetic basis of this differentiation is crucial for enhancing production efficiency. In this study, we systematically evaluated 98 representative accessions, comprising Chinese germplasm (CN) and Uruguayan germplasm. The latter included Uruguayan conventional germplasm (UY_N, where ‘N’ indicates ‘Normal’, meaning non-transgenic) and Uruguayan transgenic germplasm (UY_T). Using the “Zhongdouxin No. 1” SNP array and multi-environment phenotypic data. Uruguayan germplasm exhibited significantly higher mean oil content (21.48%) than Chinese germplasm (19.42%, p < 0.001), with high heritability (H2 ranging from 0.78 to 0.92). Genetic analysis revealed significant differentiation (mean FST = 0.14), with Uruguayan lines showing reduced diversity due to breeding bottlenecks. Genome-wide scans identified differentiation in genomic regions harboring known lipid biosynthesis genes; notably, the high-oil allele frequency of GmDGAT1 was 78.3% in Uruguayan germplasm versus 25.7% in Chinese lines, and the favorable GmbZIP123 haplotype was fixed in the Uruguayan population. Uruguayan accessions also carried significantly more favorable alleles (18.3) than Chinese accessions (14.8). We conclude that high-oil traits in Uruguayan soybean result from the systematic stacking of favorable haplotypes at key loci via directional selection. Consequently, we propose incorporating South American high-oil allelic modules into the broadly adapted genetic backgrounds of Chinese cultivars to bridge the oil content gap.
Salt stress is a primary abiotic constraint on crop production, impairing growth and reducing yield and quality in high-salinity environments. As a critical source of edible oil and protein, soybean is particularly vulnerable to salt stress during emergence and seedling establishment, making it essential to understand the genetic basis of salt tolerance at these early developmental stages for productive cultivation in saline soils. Here, a natural population of 256 soybean accessions was phenotyped for salt tolerance at emergence and seedling stages and genotyped using the ZDX1 SNP array. Genome-wide association analysis identified 60 salt tolerance-associated SNPs consolidated into 19 QTLs, 13 associated with emergence-stage and three with seedling-stage salt tolerance. Five QTLs co-localized with previously reported stress-tolerance genes, including seedling-stage QTL qSTG-SSB-03, located 192 kb from the major salt tolerance gene GmSALT3. Four candidate genes were prioritized: Glyma.10G040000, encoding a glutathione S-transferase, was identified at the emergence stage, while Glyma.10G148700, Glyma.10G149200, and Glyma.10G149600, encoding a calmodulin-binding protein (CaM), drought-induced protein 19 (Di19), and a protein phosphatase 2C (PP2C), respectively, were identified at the seedling stage, all with established roles in salt-stress responses.
Isopentenyltransferase (IPT) is the rate-limiting enzyme in cytokinin biosynthesis and plays a critical role in plant acclimation to abiotic stress. To explore soybean IPT genes, we performed genome-wide identification, bioinformatics analysis, and molecular experimental validation to systematically characterize the features and functions of the soybean IPT (GmIPT) gene family. We identified 15 GmIPT genes in the soybean genome, which are unevenly distributed across 12 chromosomes; their evolutionary expansion is primarily driven by whole-genome duplication events. Phylogenetic analysis of soybean IPT proteins with those from Arabidopsis, rice and maize clustered them into four groups, exhibiting lineage-specific functional specialization. GmIPT genes exhibit significant variations in conserved motifs, gene structure, and cis-acting elements; their promoter regions are enriched in light-responsive, abiotic stress-responsive, and hormone-responsive elements, indicating their involvement in complex transcriptional regulatory networks. Tissue expression profiling revealed that GmIPT7 and GmIPT10 are highly expressed in various tissues, whereas GmIPT14 shows specific expression in flowers and the shoot apical meristem. Transcriptomic analysis and qRT-PCR validation demonstrated that GmIPT7, GmIPT10 and GmIPT15 respond differentially to drought, salt and low-temperature stress, with GmIPT15 exhibiting a transient upregulation at 3 h (p < 0.01) followed by a gradual decline to levels close to the pre-treatment control at 6-12 h under low-temperature stress. We further performed haplotype analysis of GmIPT15 and identified a putative elite haplotype (hap1) associated with cold tolerance based on low-temperature germination index assessment. This study provides useful insights for the future functional characterization of plant IPT genes and offers potential genetic resources and molecular markers that may support molecular-assisted breeding for soybean abiotic stress tolerance.
Soybean is one of the major sources of high-quality protein and oil and plays a crucial role in ensuring nutritional and economic security. The rapidly increasing population, soybean supply–demand imbalance, and yield losses due to soil salinity collectively necessitate the development of salt-tolerant cultivars. Perennial wild Glycine species harbor-rich genetic diversity and constitute a reservoir of stress-tolerant genes. This study systematically characterized the COBL gene family in the salt-resilient Glycine tabacina (Labill.) Benth., a wild soybean that tolerates 500 mM NaCl with 100
Introduction Seed weight and nutritional composition (protein and oil content) are critical agronomic traits that collectively determine the yield and quality in soybean. However, the genetic architecture and regulatory mechanisms governing these traits remain poorly understood. Objectives This study aimed to identify the key genes and molecular mechanisms governing 100-seed weight and nutritional quality in soybean, providing a theoretical framework for the dual-enhancement of seed weight and protein content. Methods A genome-wide association study (GWAS) was conducted using 1,702 diverse soybean cultivars to identify candidate loci associated with seed weight. Functional characterization was conducted through CRISPR-mediated knockout and overexpression analyses. Population genomic analyses were further performed to elucidate the evolutionary history and selection signals of the candidate gene. Results We identified GmSW6 (Seed Weight 6), encoding a 2-oxoglutarate Fe(II)-dependent dioxygenase (2OGD), as a master regulator of 100-seed weight. Knockout of GmSW6 markedly enhanced seed weight and protein content while simultaneously reducing oil content. Mechanistically, we demonstrated that the transcription factor GmSW13 directly activates GmSW6 expression. This GmSW13-GmSW6 module, in turn, upregulates GmOLEO1 to coordinately modulate both seed weight and quality. Population genomic analysis revealed that the elite allele, GmSW6G, is significantly associated with increased seed weight and has undergone intense positive selection during soybean domestication and modern improvement. Conclusion Our findings elucidate a hierarchical genetic pathway governing soybean seed development and provide potent molecular targets for simultaneous improvement of seed weight and protein content in future ’designer’ varieties.
Understanding the relationship between genomic variation and phenotype is fundamental to deciphering the genetic architecture underlying complex traits. Yet, existing statistical models struggle to balance massive genomic datasets with biological interpretability. Here, we introduce GP-WAITER, a deep learning framework integrating GWAS-derived SNP weights into a hybrid convolutional neural network and Transformer architecture. By utilizing a weighted embedding mechanism and multi-head self-attention, GP-WAITER effectively captures long-range dependencies across ultra-long genomic sequences. The model consistently outperforms seven state-of-the-art genomic prediction models across six datasets, achieving up to a 77.5% improvement in prediction accuracy, a 78% reduction in mean squared error, and a 1.8-2.4fold increase in computational efficiency. Furthermore, GP-WAITER offers biological transparency by pinpointing key genetic variants driving specific traits. This scalable, interpretable framework provides a powerful tool for precision breeding and the functional interpretation of trait-associated variants.
Small RNAs (sRNAs) are essential for regulating plant growth and development, and they possess the notable ability to travel long distances within organisms to regulate target gene expression. Our study examined the dcl2 mutant, a key enzyme in sRNA biogenesis, to determine the role of the DCL2 protein in sRNA synthesis and to identify mobile sRNAs under DCL2 regulation. Through grafting experiments between dcl2 mutants and wild-type soybean plants, coupled with sRNA sequencing, we identified 14,105 sRNAs significantly affected by DCL2 and discovered 375 mobile sRNAs under its regulation. Degradome analysis provided deeper insights into the regulatory effects of these mobile sRNAs on their target genes, enabling us to understand their potential influences on plant development and stress responses. Leveraging the systemic movement of sRNAs from roots to shoots, we propose a novel strategy for manipulating gene expression in aboveground tissues. Overall, our research findings not only deepen our understanding of the complex regulatory networks involving mobile sRNAs regulated by DCL2, but also provide a new strategy for gene regulation, which could have a positive impact on agricultural biotechnology.