
Barley (Hordeum vulgare L.) is an important crop in the world, and its seed dormancy is primarily controlled by a mitogen-activated protein kinase kinase 3 (MKK3) gene. Although kinase activity of MKK3 and its roles in barley post-domestication have been widely studied, the pre-domestication evolution of MKK3 and the spread of nondormant alleles among global barley varieties remain largely unexplored. In this study, we analyzed MKK3 sequences in barley and its wild progenitor (Hordeum spontaneum K. Koch) and identified two polymorphic miniature inverted-repeat transposable elements (MITEs). Comparative analyses indicated that the insertions/excision of the MITEs predated the current estimates of barley domestication. Examination of the barley pangenomes coupled with droplet digital polymerase chain reaction revealed extensive copy number variation of MKK3 and suggested that transposons likely contributed to tandem amplification of the MKK3 gene on chromosome 5H. Additionally, approximately 1-Kb MKK3 sequences were found on chromosomes 1H and 6H. Further analysis indicated that these short MKK3 sequences were captured by a CACTA transposon that also contained fragments from four other expressed genes. The acquisition of MKK3 was estimated to be between 1.9 and 2.5 million years ago. Together, these findings illuminate the dynamic pre-domestication evolution of the MKK3 gene and identify three divergent MKK3 haplotype groups including a unique lineage predominant in Ethiopian germplasm. This study highlights the contribution of transposons to structural diversification and evolutionary differentiation of the MKK3 locus and provides helpful information for understanding the complex history of MKK3 gene in barley and also for improving preharvest sprouting tolerant varieties under distinct natural conditions.
DnaJ proteins (Hsp40s) are essential components of the cellular proteostasis network, functioning as molecular co-chaperones that regulate protein folding, stability, and stress-responsive homeostasis in plants. Beyond their classical role as Hsp70 partners, accumulating evidence demonstrates that DnaJ proteins participate in diverse biological processes by integrating proteostasis with hormonal, developmental, and defense signaling pathways. This review synthesizes recent advances in the structural diversity, evolutionary distribution, and functional specialization of DnaJ proteins across model plants and crops, including Arabidopsis, rice, maize, tomato, soybean, grapevine, citrus, cucumber, and other economically important species. We summarize mechanistic evidence demonstrating their involvement in abiotic stress tolerance, including heat, drought, salinity, osmotic, oxidative, and cold stress, through regulation of protein quality control, reactive oxygen species detoxification, ion homeostasis, abscisic acid and melatonin signaling, and transcriptional stress networks. Emerging studies further reveal their roles in biotic stress responses by modulating immune signaling, pathogen resistance, and host-microbe interactions, while additional evidence links DnaJ proteins to reproductive development, root architecture, chloroplast stability, fruit ripening, and postharvest quality maintenance. Advances in genome-wide identification, transcriptomic profiling, gene editing, transgenic validation, and functional genomics have substantially expanded understanding of DnaJ-mediated molecular networks and their contributions to climate resilience. By integrating mechanistic biology with translational breeding strategies, this review highlights DnaJ proteins as promising molecular targets for crop improvement and climate-smart agriculture aimed at increasing productivity, stress tolerance, and postharvest performance under changing environmental conditions.
Genomic selection (GS) is a powerful tool for accelerating genetic gain in potato (Solanum tuberosum L.) breeding, particularly for complex traits. In this study, three practical aspects of GS implementation in a potato breeding program were examined. First, the predictive ability of GS models was evaluated for three key traits (total yield, marketable yield, and specific gravity) using two elite potato populations with shared ancestry, tested across seven location-year environments. Two cross-validation strategies were used to reflect practical breeding scenarios: predicting unphenotyped lines in known environments and predicting clonal performance in unknown environments. Four models were evaluated, two of which included genotype-by-environment interactions. Tuber specific gravity showed higher and more consistent prediction accuracy across environments, supporting the evidence that it is a more stable trait. Second, the impact of genotyping platforms and marker density on GS performance were examined, as the two populations were genotyped using two different targeted sequencing platforms: Flex-seq (22K loci) and DArTag (4K loci), sharing ∼4K common loci, that allowed direct comparison. Prediction accuracies were comparable across platforms, indicating that both are suitable for GS implementation, with the choice depending on breeding goals, cost, and throughput considerations. Finally, the long-term impact of GS on genetic gain was assessed through stochastic simulation of a 30-year breeding pipeline, comparing conventional phenotypic selection with GS-assisted selection scenarios. GS scenarios achieved higher long-term genetic gains, though practical deployment should consider both cost and breeding objectives. Our findings for the three aspects of this study support the integration of GS into potato breeding programs, while highlighting key considerations for its effective implementation.
Bacterial wilt, caused by Ralstonia spp., poses a major threat to blueberry (Vaccinium corymbosum) production due to its persistence and rapid spread through soil and infected stock, highlighting the need for genetic insights to guide breeding strategies. This study investigated the genetic basis of bacterial wilt resistance in blueberry using a genome-wide association study (GWAS) across two populations comprising 401 advanced selections from the University of Florida Blueberry Breeding and Genomics Program. A high-throughput screening assay was developed to evaluate southern highbush blueberry responses to bacterial wilt based on leaf wilting severity and stem necrosis. Capture sequencing identified 38,379 single-nucleotide polymorphisms. Moderate narrow-sense heritability estimates were observed for leaf severity (0.26) and stem necrosis (0.20), and GWAS identified five small-effect quantitative trait loci on chromosomes 1, 2, 5, and 11, each explaining 4.0%-7.4% of the phenotypic variance. Candidate gene analysis revealed putative pentatricopeptide repeat (PPR), serine/threonine protein kinase, and MYB-related proteins for leaf severity, and Mlo genes and polysaccharide biosynthesis genes for stem necrosis. Genomic selection (GS) analyses demonstrated potential for improving bacterial wilt resistance, with the GS de novo GWAS approach achieving the highest predictive ability by leveraging two key markers on chromosomes 1 and 11. These results elucidate the genetic architecture of bacterial wilt resistance in blueberries and provide resources for molecular breeding strategies to enhance resistance and ensure sustainable production.
Camptothecin (CPT), a plant-derived monoterpene indole alkaloid first identified in Camptotheca acuminata, is a drug precursor widely used for cancer chemotherapeutics. However, the full set of genes responsible for CPT biosynthesis remains unclear, hindering efforts to elucidate the complete pathway or establish biosynthetic production of CPT in heterologous hosts. In this study, we engineered an experimental callus system for inducible production of CPT, which enabled multi-omics and deep learning analyses to identify candidate genes in CPT biosynthesis. We first generated an improved genome assembly and gene annotation for C. acuminata. We then leveraged the natural variation of CPT levels in C. acuminata tissues and performed transcriptomic analysis of multiple callus and tissue types to shortlist candidate enzymes responsible for CPT biosynthesis. Finally, we conducted large-scale deep learning-enabled protein-ligand complex structure prediction to prioritize 117 candidate enzymes for studies that map their roles in CPT biochemical reactions. By integrating experimental, genomic, transcriptomic, and deep learning approaches, this study provides a valuable foundation for the complete elucidation of the CPT biosynthetic pathway.
Improving seed size and weight is a major breeding goal in mungbean (Vigna radiata (L.) R. Wilczek). Improved genomic resources and precision phenotyping may enable more efficient selection for seed trait improvement. In this study, we integrated a deep learning-based segment anything model phenotyping pipeline with genome-wide association studies (GWAS), comparative mapping, and genomic prediction (GP) to dissect the genetic architecture of seed size, shape, and weight traits in the Iowa mungbean diversity panel. The zero-shot segmentation approach reliably captured seed size traits, which exhibited high heritability (H2 > 0.90) and strong correlation with seed weight (r > 0.91). A multi-model GWAS identified 82 unique single nucleotide polymorphisms (SNPs) across all seven traits, of which 13 were major pleiotropic SNPs governing multiple seed dimensions, including high-confidence regions on chromosomes 1, 4, and 6 that explained over 20% of the phenotypic variance. Within these SNP regions, comparative mapping highlighted candidate genes including an ABC transporter (Virad01G0084400) and two colocated candidates, NPGR1 (Virad06G0255600) and a RING-type E3 ubiquitin-protein ligase (Virad06G0255800), presented as hypothesis-generating candidates for seed size regulation. GP using genomic best linear unbiased prediction (gBLUP) produced moderate to high accuracies for seed size and weight traits (r = 0.76-0.84). Incorporating significant GWAS SNPs (gBLUP + SNPs) yielded slight improvements, suggesting the standard gBLUP model is sufficiently robust for selection. Collectively, this study provides the most comprehensive genomic dissection of seed size and weight traits in mungbean to date, providing candidate loci and genomic prediction models that can accelerate genetic improvement for seed yield and quality.
The rapid expansion of genomic, environmental, phenomic, and other high-dimensional data sources has transformed genomic prediction in plant breeding. However, the terms multimodal, interaction modeling, and multimodule architecture are often used inconsistently, generating ambiguity regarding whether they refer to biological assumptions, data integration strategies, or computational design. These dimensions are conceptually independent in the sense that none logically requires or implies the others. Interaction modeling may be implemented within a multimodule architecture, but modular computation does not inherently encode biological interaction. Multimodality refers strictly to the joint use of heterogeneous biological data sources; interaction modeling reflects explicit assumptions about biological dependencies such as genotype-by-environment effects; and multimodularity describes how computation is architecturally organized. We illustrate the proposed framework using conceptual and literature-based examples from wheat breeding, emphasizing interpretation rather than introducing new experimental results. By clarifying terminology and model design principles, this framework aims to improve methodological transparency, facilitate fair comparison among prediction approaches, and strengthen communication between quantitative geneticists, data scientists, and breeding practitioners. While often grouped under the umbrella of artificial intelligence, the approaches used here are more precisely framed as statistical learning methods designed to model and predict measurable genotype-environment-phenotype relationships rather than to generate synthetic or human-like outputs.
Zoysiagrass (Zoysia spp.) is an important warm-season turfgrass cultivated across tropical, subtropical, and temperate regions of the world. The genus is characterized by the presence of salt-secreting glands on the adaxial leaf surface, which contribute to its high salt tolerance. In this study, we analyzed an interspecific F2 population, derived from selfing an F1 from a cross between Z. japonica acc. Meyer and Z. matrella acc. PI 231146, for variation in adaxial salt gland density, leaf width, and vein count. Using composite interval mapping with a previously constructed genetic map as a framework, we identified three quantitative trait loci (QTL) for leaf width, two QTL for vein count, and two QTL for salt gland density. We complemented the QTL analysis with bulked segregant RNA-seq (BSR-seq) to identify shared genomic regions and candidate genes for leaf width and salt gland density. BSR-seq identified four trait-associated regions, but only a single region identified for leaf width on Chr08 overlapped with a QTL for the same trait. We highlight putative candidate genes underlying the leaf width and salt gland density QTL and discuss their potential roles in leaf development. Together, the QTL and candidate genes provide an important resource for breeding stress-resilient Zoysia germplasm.
Heterosis is critical to high maize (Zea mays L.) yields; however, its genetic mechanism remains poorly understood because different molecular markers reflect distinct genetic components. This study uses a North Carolina II mating design to evaluate grain yield per plant of 87 hybrids derived from 29 recombinant inbred lines and three testers from different heterotic groups across two environments. The correlations between heterosis, combining ability, and heterozygous functional variant sites located in upstream regions, exons, and splice sites (HEUES-single nucleotide polymorphism [SNPs] and HEUES-insertion and deletions [InDels]) were analyzed. Results showed that the heterotic group-specific and general combining ability had the strongest correlation with heterosis (r = 0.560, p < 0.001). Functional HEUES markers correlated more closely with heterosis than genome-wide genetic distance. HEUES-SNPs and HEUES-InDels showed similar associations (r = 0.469 and 0.455, respectively) with better-parent heterosis. The two marker types presented high collinearity (r = 0.987), indicating overlapping genetic information and indistinguishable independent genetic mechanisms. Residual analysis further confirmed no significant difference of performance between the two markers in different environments. Accordingly, a conceptual weighted multi-kernel genomic prediction framework integrating marker types, functional contexts and genetic architecture were proposed. This framework innovates the traditional single-marker method and provides a theoretical reference for developing advanced genomic prediction models to improve parental selection efficiency in maize breeding.
Heat stress is an increasingly serious threat to rice (Oryza sativa L.) productivity, yet the genetic and regulatory architecture underlying thermotolerance remain poorly resolved and fragmented across studies. Earlier research focused on individual pathways or specific developmental stages; however, recent advances now support an integrated understanding of heat stress adaptation in rice. This review synthesizes emerging insights into molecular physiology, regulatory signaling, epigenetic memory, and genome-scale variation associated with thermotolerance. We highlight the interconnected roles of calcium reactive oxygen species signaling, heat shock transcription factor networks, translational regulation, and chromatin-based stress memory in shaping reproductive-stage tolerance and maintaining grain quality under elevated temperatures. The review also emphasizes the value of pangenome analyses and structural variant discovery for identifying heat-responsive genes and regulatory elements absent from single-reference genomes. In addition, genome-wide association studies, haplotype-based breeding, genomic selection, and CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-based genome editing are discussed as promising approaches for functional validation and deployment of favorable alleles controlling polygenic heat resilience. Despite these advances, several challenges continue to hinder translation into breeding-ready outcomes, including limited field-based validation of candidate genes, poor integration of multi-omics datasets into predictive breeding frameworks, and insufficient understanding of reproductive-stage regulatory networks. Furthermore, genotype × environment interactions, together with trade-offs among yield, grain quality, and stress resilience, strongly influence the stability and transferability of thermotolerance traits across diverse agroecological environments. By integrating mechanistic insights with genome-scale diversity and predictive breeding tools, this review outlines a genomics-enabled roadmap for developing heat-resilient rice cultivars under intensifying global warming and supporting sustainable global rice production.
Southern corn leaf blight (SCLB) is caused by the fungal pathogen Bipolaris maydis (syn. Cochliobolus heterostrophus Drechsler) and is a common disease of fall crops of sweet corn. Phenotyping for SCLB resistance is performed through visual scoring, which is subjective and may limit genetic gain for this quantitative trait. As an alternative, we integrated computer vision (CV)-based phenotyping, genome-wide association studies (GWASs), and predictive breeding approaches to dissect the genetic basis of SCLB resistance. We utilized a sweet corn diversity panel with 693 genotypes, for which whole-genome resequencing produced a high-density single-nucleotide polymorphism (SNP) dataset. Broad-sense heritability for visual scoring ranged from 0.44 to 0.73, while CV-based phenotyping produced estimates ranging from 0.56 to 0.73 in multi-environment resistance trials conducted across 5 years and three locations. We performed GWAS using 16,755,210 SNPs and identified 41 associated SNPs. Genomic selection (GS) models on visual scoring phenotypes achieved moderate prediction accuracies under cross-validation of untested genotypes across characterized environments (0.22-0.47) and high prediction accuracies when predicting tested genotypes in uncharacterized environments (0.49-0.68). Using CV-based phenotypes for GS, we observed prediction accuracies of 0.45-0.47 under the untested genotypes in the characterized environments cross-validation scheme and 0.59-0.62 under the tested genotypes in the uncharacterized environments scheme. GS demonstrated reliability for ranking the individuals across a gradient of environments. These findings identify candidate loci and predictive breeding strategies to accelerate the development of resistant sweet corn cultivars.
Fusarium head blight (FHB) is a devastating disease that severely impacts global wheat (Triticum aestivum L.) production. Sumai 3, a wheat cultivar widely used in breeding programs for its strong FHB resistance, has not been fully resolved at the chromosome level. Here, we present a high-quality chromosome-scale assembly of Sumai 3 using PacBio HiFi reads and chromosome conformation capture sequencing. The 14.6 Gb assembly consists of 832 contigs, with the longest contig being 245.2 Mb and a contig N50 of 41.90 Mb, which were scaffolded into 21 pseudomolecules. De novo annotation identified 104,620 high-confidence protein-coding genes and found 92.67% of the genome to consist of repetitive sequences. Synteny analysis showed strong collinearity between Sumai 3 and the wheat reference sequence Chinese Spring IWGSC RefSeq v2.1 (CS). Structural variant analysis identified chromosome 2A with the highest number of deletions (2773) and insertions (2645), while chromosome 3B had the most inversions (359). Duplications were most frequent on 2A (337), and contractions on 5B (122). The gene content of the major resistance quantitative trait loci on 3B, Fhb1, largely validates previous annotations for CS, although we discovered two new genes at approximately 12.4 Mb, including an additional copy of a terpene synthase, further suggesting homology with CS 3D over 3B. Differential expression analysis highlighted up-regulation of three genes coding pore-forming toxin-like protein, sina superfamily protein, and plastid-lipid-associated proteins potentially involved in FHB resistance. This assembly provides critical insights into FHB resistance and offers a valuable genomic resource for wheat breeding programs.
Red clover (Trifolium pratense L.) is a globally important temperate forage legume. Its symbiosis with soil-borne rhizobia enables nitrogen fixation, and its ability to produce quality forage under diverse soil conditions enhances pasture productivity, particularly during water deficits. With increasing climate-related stresses, harnessing adaptive traits absent in current cultivars is critical. Genebanks conserve diverse red clover germplasm, providing genetic variation for agronomic and adaptive traits. In this study, we introgressed novel germplasm into locally adapted cultivars to track the inheritance of allelic variants using genotyping-by-sequencing. Multi-location, multi-year trials evaluated half-sib families (generation two [Gen 2]) two generations removed from the exotic germplasm (genereation zero [Gen 0]) against local cultivars. Several Gen 2 populations matched or outperformed local cultivars and exhibited a moderate family mean heritability (h2 > 0.40) for most traits. Integrating genomic, phenotypic, and environmental data, 77 bioclimatic-associated single nucleotide polymorphisms (SNPs) were identified, of which 35 SNPs and 27 associated genes were significantly linked to trait expression. By using the original germplasm (Gen 0) as a training population and the derived half-sib families (Gen 2) as a validation population, genomic prediction models were developed to calculate prediction accuracies for key agronomic traits. Biomass and plot density traits showed high predictive abilities and the highest prediction accuracies across generations. This study demonstrates a route by which genetic diversity from genebanks can be successfully incorporated into local populations, enabling evaluation and selection of key traits. The identified molecular markers and genomic prediction models provide a pathway to efficiently develop climate-adaptive red clover cultivars.
The International Weed Genomics Consortium (IWGC) has sequenced and annotated the genomes of over 30 weed species, generating genomic resources to understand their biology, evolution, and adaptation. The objective of this study was to evaluate the semi-automated, isoform sequencing (Iso-seq)-based, IWGC genome annotation pipeline by reannotating the genome of the model species Arabidopsis thaliana with various amounts and types of extrinsic data and to measure the impact that varying inputs had on the annotation completeness and quality. Annotations were run comparing the effects of (1) the quantity and source of Iso-seq reads, (2) annotated proteins from botanically closely related or distantly related species, and (3) the number of proteins provided to the annotation program "MAKER-P." Reannotations were compared to each other and to the published annotation of the A. thaliana genome. The IWGC annotation pipeline annotated almost all the genes without manual curation when informed with an Iso-seq dataset and proteins of related species. In general, the pipeline produced more accurate, annotated genes with more input proteins, especially from closely related species, in the gene model prediction step. Furthermore, the combination of proteins from several closely related species increased the number of annotated genes. The number or source of Iso-seq reads did not have a significant effect if many proteins from closely related species were utilized. The annotation pipeline annotated nearly 90% of genes from additional crop species genomes. The IWGC genome annotation pipeline is robust in reannotating the A. thaliana genome and therefore is most likely performing well in the several non-model weed species it has been used on so far.
Induced mutagenesis is a cornerstone of crop functional genomics, yet the extent to which distinct radiation sources reshape the spatial distribution of mutations remains difficult to evaluate in reduced-representation datasets. Here, we analyze a published genotyping-by-sequencing (GBS) panel (192,040 loci) to compare proton-beam and gamma-ray mutagenesis in Sorghum bicolor. Because GBS sampling is nonuniform, all analyses were conducted within an explicitly defined GBS-callable sequence space. Within this callable space, 96-channel trinucleotide spectra were broadly similar between radiation types, whereas spatial summaries differed. Macroscale analysis using the Gini coefficient indicated that proton-treated lines exhibit a highly unequal, spike-like distribution of mutations, whereas gamma-treated lines show a more diffuse window-level distribution. Microscale spatial statistics were consistent with clustering patterns that were more prominent in proton-treated lines, including an aggregation scale of ∼500 kb, with a substantial fraction of the mutational burden falling into high-density tracks. Within the callable locus set, coding- and promoter-proximal categories were not depleted of induced mutation events (single-nucleotide variants) across treatments. Furthermore, we did not detect a negative association between total mutational load and the coding-region mutation fraction in this dataset. These findings suggest that, within this dataset, proton mutagenesis is characterized not by unique chemical signatures but by a distinct spatial geometry that concentrates detectable mutation events. Because proton irradiation was represented by a single dose, whether proton treatment produces stronger clustering than gamma irradiation at equal mutational burden remains to be directly tested.
Abstract Deep learning, as a pivotal branch of machine learning, has demonstrated remarkable potential in advancing crop science by effectively integrating genomics and phenomics. This review systematically outlines the application of diverse deep learning architectures—such as convolutional neural networks, recurrent neural networks, and transformers—across key crop genomic tasks, including gene expression prediction, alternative splicing analysis, cis‐regulatory element identification, epigenomic profiling, and genome‐based trait prediction. In phenomics, these models facilitate high‐throughput extraction of crop phenotypic traits from multispectral, unmanned aerial vehicle, and ground‐based imagery, supporting yield forecasting, disease diagnosis, and stress response monitoring. We critically evaluate the performance and limitations of each model type across tasks, considering trade‐offs between complexity, accuracy, and interpretability, to offer practical guidance for crop researchers. Additionally, the review addresses major challenges in deploying deep learning—such as data scarcity, model transparency, and computational demands—and proposes future pathways to enhance model generalizability, multimodal data integration, and applications in intelligent breeding and sustainable agriculture.
Abstract A framework beyond single‐reference genomes is needed to understand transcription factor evolution. This study employed an integrated haplotype‑resolved genomes–transcriptome atlas–machine learning to characterize the transcription factors of autotetraploid green jujube (Ziziphus mauritiana). The first haplotype‑resolved comparison of transcription factor superfamilies from eight haplotype genomes (HapGenome) representing a specific wild and a specific cultivated green jujube accession, encompassing 42 superfamilies and 12,123 gene copies. Evolutionary analyses revealed high structural conservation with minimal copy number variation, gene presence/absence variations, and strong purifying selection (Ka/Ks < 1). Dispersed duplication (47.24%), not whole‐genome duplication (36.10%), was the most frequently observed duplication event in the expansion of transcription factor superfamily. A haplotype‑resolved transcriptome atlas demonstrated that tissue‐specific expression divergence occurred at the superfamily level and between the core/dispensable genes. Integrating transcriptomic and metabolomic data, support vector machine classification with leave‑one‑out cross‑validated distinguished three wild fruits from six cultivated fruits with the accuracy of 89% using orthologous gene groups (OGGs) expression profiles. The eXtreme gradient boosting was employed as an exploratory tool to prioritize OGGs related to metabolite changes. Finally, OGG‐95, a Lesion Simulating Disease Zn finger transcription factor, was screened out, which was significantly upregulated in cultivated fruits, and its expression was significantly correlated with differential accumulation of nucleotides and organic acids that need further functional validation. This integrative study provided novel insights into the genomic architecture and regulatory evolution of transcription factors in a polyploid fruit crop, highlighting the power of multi‐dimensional analyses for gene discovery.
Pecan [Carya illinoinensis (Wangenh.) K. Koch] is the fifth-largest tree nut in global cultivation, with 80% of production occurring in the southern states of the United States. Despite the economic and health benefits of pecans, there is a lack of genomic tools available to breeders for crop improvement. The pecan breeding community is small, and most breeding programs have many barriers to adopting technology, particularly in the cost and know-how needed to create and use genetic marker panels for genomic-based decisions in selection. Here, we report the creation of a DArTag (Diversity Array Technology [DArT]) panel of 3100 loci distributed across the diploid pecan genome for use in molecular breeding and genomic prediction. Here, we show the panel's ability to distinguish key parents and founders of cultivated pecan from other Carya species that may be used in pre-breeding. We also demonstrate that the resultant data can create a linkage map in a biparental population. The creation of this marker panel brings cost-effective and rapid genotyping capabilities to pecan breeding programs, making routine genotyping a reality for any pecan breeder. Furthermore, the open access provided by this platform enables the comparison and integration of genetic datasets generated on the marker panel across projects, institutions, and countries.
Abstract The processing of phenotypic information prior to training genomic selection (GS) models is a key factor that is frequently overlooked. Several approaches have been proposed to isolate the genetic signal from the field variability. However, in most cases, the estimated genetic signal still carries the field variability print. In addition, the statistical metrics are not conclusive about the model that isolates the signal the best since the breeding values are unknown. In this study, we evaluate the effects of different spatial models for separating the genetic from the field variability components, and their repercussions implementing GS models. A real soybean (Glycine max L. Merr.) data and a simulation study under controlled conditions were analyzed. Three standard models were implemented accounting for different field variability components (M1: block, M2: block + row + column, and M3: block + row + column + row × column). Results derived from the real dataset showed that accounting for field variability reduces predictive ability of the isolated genetic signals. In the simulated data, however, it was found that field variability corrections improved the predictive ability of breeding values. We conclude that training GS models with isolated genetic signals improves the predictability of breeding values and that the current benchmarks, relying on the correlation between predicted and observed values, can be misleading due to the lack of comparability between phenotypes and breeding values.