Genomic selection holds the potential to serve as a strategic tool to enhance the genetic gain of complex traits in Miscanthus breeding programs. The development of improved cultivars requires their assessment for various traits across diverse environments to ensure suitable overall performance. Hence, the multi-trait multi-environment (MTME) genomic prediction (GP) models offer an opportunity to improve selection accuracy. This study aims to evaluate the potential of five GP models: (1) three MTME models including genotype-by-trait-by-environment interaction (G×E×T) and (2) two single-trait multi-environment (STME) models (with and without G×E interaction). A Miscanthus sacchariflorus population comprising 336 genotypes evaluated in three environments and scored for four traits (biomass yield YDY, total culm number TCM, average internode length AIL, and culm node number CNN) was analyzed. The predictive ability of the models was evaluated considering three cross-validation schemes resembling realistic scenarios (CV1: predicting new genotypes, CVP: predicting missing traits in a given environment, and CV2: predicting partially observed genotypes). On average, in all cross-validation schemes compared to the STME the predictive ability of the MTME models was 10% to 70% higher for TCM and AIL. On the other hand, for YDY and CNN, both STME models performed similarly or slightly better (between 5 to 64%) than the MTME models in most environments. While the MTME models were not successful for all traits when compared to their STME counterparts, MTME models improved the prediction of the performance of genotypes that were untested across environments or lacked trait information in a specific environment. Overall, our study suggests that MTME GP models can be implemented in Miscanthus breeding programs to improve the predictive ability of the complex traits, shorten breeding cycles, and accelerate selection decisions.
Genomic selection has accelerated genetic gain in many breeding programs worldwide but genotype-by-environment-by-management (GxExM) hampers further progress for systems where these interactions are important and not well represented in the training data. Process-based crop growth models (CGMs), which encode physiological relationships between plants and their environments, can extrapolate to novel conditions but cannot directly leverage genomic information. Coupling these complementary approaches in the Crop Growth Model–Whole Genome Prediction (CGM-WGP) framework addresses both limitations, yet applications in horticultural crops remain scarce. In this study, we apply CGM-WGP to predict flowering time in broccoli ( Brassica oleracea var. italica ) and common bean ( Phaseolus vulgaris L.), two horticultural species with contrasting physiological responses to temperature and photoperiod. Genotype-specific thermal time requirements and photoperiod parameters were jointly estimated with genome-wide marker effects and predictions were compared against a Reaction Norm Genomic Best Linear Unbiased Prediction (RN-GBLUP) benchmark across four cross-validation scenarios of increasing predictive difficulty. RN-GBLUP achieved the highest accuracy under sparse-testing scenarios where training data covered all target environments, while CGM-WGP outperformed RN-GBLUP when predicting untested environments and untested genotype-environment combinations (broccoli: Pearson r = 0.66, RMSE = 9.4 days; bean: r = 0.86, RMSE = 5.2 days). These results demonstrate that CGM-WGP can be applied to horticultural crops using genome-wide markers alone, but without requiring prior identification of quantitative trait loci. CGM-WGP also provides a modular foundation that can be extended to predict the timing of other developmental transitions and output traits such as biomass and yield.
ABSTRACT Genomic prediction models that account genotype-by-environment (G×E) have the potential to accelerate the rate of genetic gain for yield and agronomic performance, yet relatively few studies have applied G×E prediction in public soft red winter wheat ( Triticum aestivum ) breeding programs. In this study, we extended a reaction norm-based genomic prediction framework by integrating weather-based environmental covariates to more effectively capture genotype– environment interactions. Key agronomic traits, including seed yield, plant height, test weight, and heading date, were evaluated across 33 environments (location–year) using over 3,200 breeding lines from the North Carolina State University small grains breeding program. Multiple genomic prediction models were compared using several cross-validation (CV) schemes representing common breeding scenarios. Across traits, the reaction norm M5 model, which incorporates both G×E and genotype-by-environmental covariate interactions (G×O), achieved the highest prediction accuracy (PA) in CV2 (predicting incomplete field trials) and CV1 for yield and test weight (predicting new lines). The highest PA was observed for test weight under CV2 (0.54) and for yield under CV1 (0.41). Under CV0 (predicting new environments), the M3 model incorporating G×E produced highest PA across traits, with the greatest accuracy for plant height (0.45), although differences among M2, M3, and M4 were small. Prediction under CV00 (predicting new lines in new environments) remained more challenging, with PA values 0.10 – 0.20 across traits. Overall, our results demonstrate that integrating environmental covariates into genomic prediction models can improve predictive performance across diverse wheat-growing environments in North Carolina, supporting their utility for applied breeding efforts. CORE IDEAS Integrating genotype-by-environment (G×E) interactions with environmental covariates improves prediction accuracy across environments. Model performance varies by prediction scenario, with different approaches performing best for new lines, incomplete trials, or new environments. Prediction of new lines in new environments remains challenging. PLAIN LANGUAGE SUMMARY This study explores how adding environmental information to genomic prediction models can improve prediction accuracy in a public winter wheat breeding program. Using data from multi-environment trials conducted across diverse conditions in North Carolina, we evaluated statistical models that capture how different wheat lines respond to changing environments. By incorporating weather data, we improved the ability to predict performance across locations and years. These findings provide practical insights for refining selection strategies and accelerating genetic gain in wheat breeding.
Abstract The integration of digital technologies for high-throughput field phenotyping is critical for accelerating crop improvement in agriculture. However, extracting traits from remote sensing data remains constrained by fragmented workflows, manual intervention, and limited interoperability among existing tools, resulting in delays that hinder timely biological insight and decision-making. To address these challenges, we present PhenoStream (Phenotyping Streaming), a scalable, end-to-end cyberinfrastructure designed to automate the full lifecycle of aerial imagery-based phenotyping, from data acquisition to plot- and genotype-level inference. The framework integrates automated data ingestion from distributed field sites, geospatial processing, and AI-enabled trait extraction within a unified, user-accessible graphical interface. Its modular and extensible architecture supports adaptable trait modeling and seamless integration of new data sources, enabling deployment across diverse crops, environments, and experimental designs. We demonstrate the system across a large multi-location field trial network of bioenergy crops, where it enables high-throughput characterization of spatiotemporal growth dynamics, genotype-by-environment (G×E) interactions, and predictive modeling of key agronomic traits. By significantly reducing processing latency and manual effort, the platform facilitates near-real-time analysis and reproducible workflows. This work establishes a generalizable and scalable pathway for operationalizing very-high-spatial resolution aerial phenotyping in agricultural research. By bridging data acquisition and analytics, the end-to-end cyberinfrastructure provides a foundation for integrating heterogeneous and unstructured data streams-including remote sensing, environmental, and management data - toward data-driven decision making in agriculture.
Optimizing biomass partitioning is essential for achieving sustainable yield improvement in wheat, particularly under increasing environmental stress. Traits such as spike partitioning index (SPI), harvest index (HI), and fruiting efficiency (FE) are central to understanding how assimilates are allocated between vegetative and reproductive organs. However, their complex physiology and the difficulty of manual phenotyping have limited their routine use in breeding programs. This study assessed the potential of unmanned aerial vehicle (UAV)-based hyperspectral reflectance data to predict biomass partitioning traits and related yield components in wheat. Three trials of facultative soft wheat lines (2022-2024) and an independent validation set of advanced breeding lines were used to develop genomic prediction (GP), phenomic prediction (PP), and integrated multi-omic models combining genomic, phenomic, and environmental covariates (ECs). Kernel-based best linear unbiased prediction (BLUP), and machine-learning based, random forest regression and partial least squares regression were implemented to estimate predictive ability (PA). Phenomics-driven models markedly outperformed GP across most traits, achieving PA up to 0.61 for SPI, 0.56 for FE, 0.71 for grains/m2 (GN), and 0.66 for grain yield (GY). Hyperspectral data provided higher accuracy than vegetation indices, and multi-omic integration slightly improved prediction (PA up to 0.73 for GN). These results demonstrate that UAV-based hyperspectral phenotyping can effectively capture canopy-level physiological signals associated with biomass partitioning, offering a scalable and data-driven approach for in-season selections. This can help wheat breeding programs to optimize biomass partitioning in modern wheat cultivars for long-term yield resilience and genetic gain.
Integrating genomic and environmental information holds the potential for enhancing the predictive power of genomic prediction models when accounting for the genotype-by-environment interactions. Hence, incorporating environmental covariates (EC) into these models can significantly influence their predictive accuracy. In this study, we utilized 1379 genotypes from the SoyNAM dataset, evaluated across four environments and genotyped with 4611 single-nucleotide polymorphism markers, to compare models incorporating genotype-by-environment and genotype-by-environmental covariate interactions using different covariance matrices. We evaluated four approaches: summarizing EC by averaging (AVG), filtering ECs based on a coefficient of determination criterion (FILT), segmenting ECs by crop phenology (STG), and a naïve approach that utilized all available information (ALL). Predictive ability was assessed as the Pearson's correlation between the genomic estimated breeding values and the adjusted phenotypes considering 10 replicates of three cross-validation scenarios (CV2: predicting tested genotypes in observed environments; CV1: untested genotypes in observed environments; CV0: tested genotypes in novel environments). Incorporating EC information into the models increased average predictive ability from 0.42 to 0.56 for CV1 and CV2. In these cases, the predictive ability was lower when EC information was averaged to compute the environmental kinship matrix, with slight differences observed with respect to the other approaches. Regarding the CV0 scheme, the model incorporating only genotype-by-environment information performed better (0.33). The naïve method, which utilized all available EC information (ALL), proved to be a promising approach, as it effectively improved the results in these scenarios while eliminating the need for additional steps in selecting variables.
We investigated the potential of incorporating grid-cell-based environmental covariates (ECs) in the genetic evaluation of total sperm count (TSC), sperm motility (MOT), and sperm morphology (MOR) for Duroc boars. A total of 188,665 records derived from 3,684 genotyped boars, born between December 2018 and October 2024 and raised in three stud farms located in different U.S. states, were analyzed using multi-trait linear-threshold repeatability models. To account for genotype by environment interactions (GE), we constructed an interaction matrix as the Hadamard product of the genomic relationship matrix and an environmental (co)variance matrix. The environmental groups were defined in three ways: farm, farm-season, and farm-year-season. The (co)variance matrix was constructed based on daily ECs obtained from the NASA POWER database for each environmental group. Of all available ECs, those significantly associated with TSC, MOT and MOR (temperature, relative humidity, atmospheric pressure, and wind speed and direction) were retained. We evaluated five models with different GE structures: M1 represented the baseline without accounting for GE, in M2 the GE included farm as environmental groups, in M3 the GE included farm-season as environmental groups, in M4 the GE included farm-year-season as environmental groups, and M5 involved M3 with an additional random effect of the farm-season. Estimates of heritability for TSC, MOT, and MOR ranged from 0.03 to 0.04, 0.05 to 0.08, and 0.04 to 0.08, respectively. Corresponding repeatability ranged from 0.15 to 0.23, 0.28 to 0.49, and 0.28 to 0.49. The proportion of phenotypic variance attributed to GE variance ranged from 0.00 to 0.32, 0.00 to 0.44, and 0.00 to 0.44. Lastly, estimates of genetic correlation, TSC-MOT, TSC-MOR, and MOT-MOR ranged from 0.27 to 0.31, 0.24 to 0.31, and 0.98 to 0.99, respectively, with minor differences across models. We assessed the predictive ability of models using the linear regression validation. Across traits and models, bias ranged from -0.05 to 0.02 standard deviations, slope varied from 0.88 to 0.99, the correlation ranged from 0.75 to 0.84, and accuracy from 0.41 to 0.53. Overall, building the GE matrix considering grid-cell-based ECs helped to account for GE, thereby reducing the proportion of phenotypic variance attributed to genetic components; however, it did not improve the validation metrics. Additional on-farm records for ECs may improve the model performance.
The main approach for improving multiple traits simultaneously is the selection index. The most widely used selection indices are those based on factor analysis, which overcome statistical limitations such as multicollinearity and the reliance on arbitrary weights of the classical Smith–Hazel approach and support multi-environment trials. Nevertheless, the efficiency indices are affected by factors such as genotype number, environment and trait correlation, and heritability. In this study, we simulated different scenarios varying the mentioned factors to evaluate the performance of the Factor-Analysis and Ideotype-Design-Based Index (FAI-BLUP), Multi-trait Genotype–Ideotype Distance Index (MGIDI), and Multi-Trait Stability Index (MTSI). All correlations were positive and constant within each scenario, while the ideotype sought genetic gains for traits in opposite directions. Simulations were conducted using AlphaSimR and FieldSimR, and indices were implemented via the metan package. Results showed that index efficiency was higher in scenarios with larger numbers of genotypes, low-to-moderate trait correlations, and moderate-to-high inter-environment correlations. However, strong correlations among traits, particularly when combined with high heritability, compromise selection index efficiency in scenarios with antagonistic trait objectives. Despite that, the MGIDI consistently outperformed the other indices across most scenarios. Therefore, we emphasize accounting for trait genetic architectures, genotype–trait correlations, and target environment correlations.
Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in conventional perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse bio-based products. Increasing biomass yield will increase profitability and environmental benefits, so it is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed implementing sparse testing designs. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; G×E interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for G×E interaction consistently presented the highest PA and the lowest MSE: for CNN (PA: ~0.77, MSE: ~0.5) and YDY (PA: ~0.70, MSE: ~1.3) while for TCM and AIL these ranged from ~0.28 to 0.41 and ~1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.
Genomic prediction (GP)-based sparse testing allows evaluation of more genotypes within a fixed budget in multi-environment trials (METs), thereby reducing phenotyping costs. It assesses untested genotype-environment combinations using varying training sets and allocation schemes. This study implements sparse testing designs in tetraploid potato (Solanum tuberosum L.) breeding trials employing varying compositions of overlapping and nonoverlapping genotypes in different allocation schemes, and utilizing three calibration set sizes (19, 16, and 13). A total of 114 unique genotypes were tested for dry matter content (YDY), tuber length (TL), and tuber count (TC), using three prediction models: E + L (environment + line), E + L + G (environment + line + markers), and E + L + G + GE (environment + line + markers + genotype-by-environment interaction [G & times;E]). Considering the largest training set and the complete nonoverlapping and zero-overlapping genotypes strategy, no differences were observed between models, yielding predictive ability (PA) values of 0.83, 0.70, and 0.50 for YDY, TL, and TC, respectively. As more overlapping genotypes were included in the designs, PA declined across schemes, with the E + L model showing a more rapid decline than models with genomic data. The similar performance of the E + L + G and E + L + G + GE models indicated minimal G & times;E interaction in the dataset. Additionally, reducing the training set size affected PA in designs with fewer nonoverlapping genotypes. There was no restoration of PA after adding more overlapping genotypes to the designs. Our findings suggest that evaluating more nonoverlapping and a few overlapping genotypes could reduce phenotyping costs by fivefold and increase testing capacity for tetraploid potato cultivars in METs for maximizing overall genetic gain.
Abstract Lima bean ( Phaseolus lunatus L.) is an economically and agronomically important grain legume. Lima beans (or limas) show a range of climatic adaptations with independent domestications in the Andes (large-seeded) and Mesoamerica (small- or medium-seeded). We generated and integrated genotypic and comprehensive field- and laboratory-based phenotypic information for the available accessions in the USDA National Plant Germplasm System collection across multiple environments to inform germplasm utilization in breeding. A total of 810 accessions were genotyped using short-read, low-coverage sequencing. Accession geographic origin and domestication explained population structure. A partially overlapping subset of the panel ( n =141-308) was field-evaluated across two years in each of Davis, CA, Central Ferry, WA, and Coachella Valley, CA (the latter was fall-planted for evaluation of photoperiod-sensitive accessions) to assess trait performance in contrasting environments. Agronomic traits such as determinacy and flowering time, and seed traits such as seed coat color and hundred-seed weight, were scored. Macronutrient traits (protein, starch, fat, and ash content) were measured on dry (mature) harvested grain via near-infrared spectroscopy. Genome-wide association analyses identified loci significantly associated with descriptive, agronomic, and seed traits, including orthologs of known genes in common bean and novel candidate regions. Genomic predictive abilities were moderate to high for key traits. Finally, we established a conditional core collection that was constrained to include 211 extensively phenotyped accessions and for which 91 supplemental accessions were selected to maximize genetic diversity from among the genotyped accessions. Overall, these resources provide a foundation to support genomics-assisted breeding of limas.
Abstract Genotype-by-environment interaction (GEI) has been studied to identify environment-stable/favorable genotypes. The GEI simulation could help refine the inference by incorporating tangible factors such as genomic and environmental information. The Bayesian additive main effect and multiplicative interaction (Bayesian AMMI) model captures the genotype-specific responses across environments, reflecting directional relationships between genotypes and environments. Thus, we propose a Bayesian AMMI-based GEI simulation framework that utilizes high-throughput environmental covariance matrices to generate GEI effects with interpretable directional structure. To demonstrate the proposed approach, two simulated phenotypes were assessed under four levels of GEI variance. In the first simulation (Sim1), GEI effects were sampled from a multivariate normal distribution defined by the GEI matrix. In the second simulation (Sim2), GEI effects were generated by extending Sim1 with the Bayesian AMMI model. In both simulations, increasing GEI variance resulted in lower correlations of phenotypes across environments and stronger genotype-specific sensitivity to environmental variation. Across five cross-validation designs, models accounting for GEI consistently outperformed one that did not, with prediction accuracy generally decreasing as GEI variance increased. Clear distinctions between the two simulated phenotypes were evident from biplot analyses: Sim2 successfully captured environmental relatedness and genotype-specific responses, whereas such structure was absent in Sim1. These results demonstrate that the proposed Bayesian AMMI-based GEI simulation framework enables interpretable visualization of GEI and supports genomic selection strategies under complex environmental conditions.
In genomic prediction, it remains unclear whether increasingly complex or ensemble models improve prediction over established linear approaches, and why prediction accuracy varies among traits. Here, we evaluated a comprehensive suite of genomic prediction models, including linear mixed models, Bayesian variable selection, kernel methods, machine learning algorithms, graph attention networks, and stacked ensembles, in mango (Mangifera indica L.). Across 5 traits, prediction accuracy converged across linear, Bayesian, kernel, and ensemble models, with only marginal gains derived from stacking and no systematic advantage of machine learning approaches. Ensemble ablation and weight analyses revealed that predictive signal was dominated by additive and smooth kernel components, while more complex learners contributed little or negatively upon performance. To explain these trait-dependent patterns in predictability, we quantified the phylogenetic signal using genome-wide marker-based trees. All traits showed a significant phylogenetic signal, with the magnitude varying widely and strongly associated with prediction accuracy (r ≈ 0.71). Traits with strong phylogenetic structure achieved the highest prediction accuracies, whereas traits with a weaker signal were consistently harder to predict, regardless of model choice. Together, these results confirm that, in mango, genomic prediction accuracy is determined more by evolutionary structure and trait architecture rather than increasing model complexity. Aligning prediction strategies with the evolutionary basis of trait variation may therefore be more effective than adopting increasingly complex models.
Florida and California produce 98% of U.S. strawberries, with Florida growers' profitability depending on high yields early in the season (November-January), when prices are the highest in the U.S. market. This study aims to model the cumulative Marketable Yield and Delta Yield curves (difference in cumulative Marketable Yield between consecutive harvest time points) of strawberry genotypes to facilitate selection for greater early season productivity. The dataset comprised thirteen seasons (2013-14 to 2025-26) of advanced selection trial data. Marketable Yield trajectories were modeled using Legendre polynomial smoothing, with optimal degree selection balancing flexibility and noise reduction. The resulting coefficients served as surrogate phenotypes for genomic prediction. Forward prediction cross-validation was implemented for five seasons (2021-22, 2022-23, 2023-24, 2024-25, and 2025-26), with each season predicted using the information from all preceding seasons. For cumulative yield, across all five seasons, reconstructed curves from the Legendre models presented a clear temporal trend, with predictive ability increasing from low early-season values to peaks around Trait Dates (weeks) 7-8. In the 2021-22 and 2022-23 seasons, Legendre models showed higher predictive ability than single time-point predictions but were comparable to single-time point predictions for the other seasons. Legendre polynomial models utilizing Delta Yield achieved moderate predictive ability across five validation seasons, with consistent advantages over single time point models particularly in earlier seasons, indicating that genetic control extends beyond total yield to the trajectory of yield accumulation. Overall, Legendre modeling effectively captured the temporal dynamics of yield development while describing the trajectory with only a few parameters.
Meloidogyne enterolobii is a virulent root-knot nematode (RKN) species posing a significant threat to watermelon production across the United States. The USDA, ARS, Plant Introduction (PI) collection of Citrullus amarus, a wild relative of cultivated watermelon (Citrullus lanatus), contains RKN-resistance. However, incorporating RKN resistance into watermelon cultivars is challenging. Genomic selection could be a useful strategy for improving RKN resistance in watermelon as it is a polygenic trait primarily controlled through additive gene action, as indicated by the broad-sense heritability being close or equal to the narrow-sense heritability: H2 = 0.69, h2 = 0.23 for galling and H2 = 0.47, h2 = 0.46 for eggs per gram of root. Here, a genomic selection approach was used to evaluate 97 C. amarus PIs (S3 inbred lines) for RKN resistance based on two criteria: (1) root galling percentage and (2) eggs per gram of root. Variants from whole-genome resequencing (2.1 million SNPs) of PIs were used here. Genomic prediction (GP) models, including genomic best linear unbiased prediction (GBLUP), sparse GBLUP, and the reproducing kernel Hilbert space (RKHS) regression, were tested. The RKHS model based on 2.1 million SNPs had the highest prediction ability for galling (0.42) and for eggs per gram of root (0.50). The most resistant PIs (galling: PI_596692, PI_532664, PI_288316, PI_271767, and PI_542118; eggs per gram root: PI_271775, PI_485584, and PI_299379) based on both best linear unbiased estimates and genome estimated breeding values under 10% selection intensity were selected and are being used in our breeding program to enhance RKN resistance in watermelon cultivars.
Genomic selection (GS) is a promising strategy for accelerating genetic gains of complex traits in breeding programs. Despite the recent advancements in high-throughput genotyping technologies, the selection of the type of marker systems needed for GS remains challenging in breeding programs. In this study, we explored 3K array single nucleotide polymorphisms (SNPs) and genotyping by sequencing (GBS) SNP markers for genomic prediction of oat biomass yield using different statistical and machine learning approaches. An oat panel consisting of 420 lines was phenotyped for biomass-related traits for 3 years and genotyped using two different marker platforms (3K array and GBS). Our results showed similar performance of both the 3K array and GBS-based SNPs in terms of training population optimization, forward prediction, and univariate and multivariate genomic prediction of forage yield. The genomic best linear unbiased prediction (GBLUP), Bayes-B, and random forest models gave similar predictive ability for dry matter yield (DMY) in different harvest-year combinations and for both marker platforms. The multivariate models involving various combinations of secondary traits (simple breeders' field notes and data) resulted in more than twofold increases in predictive abilities compared to the univariate models. Comparison of the 25% top-performing observed and predicted genotypes showed a higher overlap percentage (30.10%-66.99%) for multivariate GBLUP models compared to the univariate models (27.18%-51.46%). This further elucidates the great potential of multivariate GS models incorporating the more robust and easily reproducible 3K array SNP markers for improving the genetic gains of DMY in breeding programs.
Abstract The processing of phenotypic information prior to training genomic selection (GS) models is a key factor that is frequently overlooked. Several approaches have been proposed to isolate the genetic signal from the field variability. However, in most cases, the estimated genetic signal still carries the field variability print. In addition, the statistical metrics are not conclusive about the model that isolates the signal the best since the breeding values are unknown. In this study, we evaluate the effects of different spatial models for separating the genetic from the field variability components, and their repercussions implementing GS models. A real soybean (Glycine max L. Merr.) data and a simulation study under controlled conditions were analyzed. Three standard models were implemented accounting for different field variability components (M1: block, M2: block + row + column, and M3: block + row + column + row × column). Results derived from the real dataset showed that accounting for field variability reduces predictive ability of the isolated genetic signals. In the simulated data, however, it was found that field variability corrections improved the predictive ability of breeding values. We conclude that training GS models with isolated genetic signals improves the predictability of breeding values and that the current benchmarks, relying on the correlation between predicted and observed values, can be misleading due to the lack of comparability between phenotypes and breeding values.
Improving grain yield remains the central objective of soybean breeding programs. During early-stage yield trials, breeders often evaluate thousands of genotypes; however, limited seed availability constrains the number of tested environments and replications, reducing selection accuracy. Genomic prediction offers a promising approach to identify high-yielding and stable genotypes earlier in the breeding pipeline. The objective of this study was to develop a classification-based genomic prediction framework that directly targets advancement decisions by assigning genotypes to yield performance classes while estimating the probability of class membership to prioritize genotypes with higher confidence. A total of 1,789 soybean genotypes, ranging from maturity groups III to V, were evaluated for grain yield across 10 environments (year × location combinations) in Arkansas and Missouri during the 2023 and 2024 growing seasons. Genomic Best Linear Unbiased Predictors (GBLUPs) were obtained for each genotype in each environment, and a selection index (MSI) was calculated as the average yield deviation from the mean of the checks across the tested environments, centered at zero. This metric captures both yield and consistency across environments using a simple, check-referenced scale that is directly interpretable in breeding decisions. Genotypes were then classified as high-yielding (MSI ≥ –5), moderate (–5 > MSI ≥ –15), or low-yielding (MSI < –15). Two classification-based genomic prediction models, Generalized Linear Model via Elastic Net Regularization (GLMNet) and Random Forest (RF), were trained using the SoySNP3K BeadChip markers as predictors and the MSI-based yield classes as response categories. The MSI ranged from -32.4 to 7.2, with a small proportion of genotypes in the high-yielding class. GLMNet and RF achieved macro-averaged balanced accuracies of 0.84 and 0.83, respectively, with high specificity (0.89 for both) and sensitivity (0.78 and 0.76), and minimal extreme misclassification between low- and high-yielding classes. Compared to regression-based genomic prediction, this classification framework aligns with advancement decisions, is less sensitive to early-stage noise, and retains greater genetic diversity than GBLUP-based ranking, enabling more efficient resource allocation and more targeted advancement of promising genotypes.
Abstract The development of improved cultivars requires establishing multi‐environment trials (METs) to evaluate their performance under a wide range of environmental conditions. However, the high phenotyping costs often limit the capacity to evaluate genotypes in all the target environments. Our main objective was to explore the potential of implementing sparse testing in cassava breeding programs to reduce the cost of phenotyping in METs. The population used in this study consisted of 435 cassava genotypes evaluated in five environments in Nigeria for dry matter (dm) and fresh root yield (fyld). Sparse testing designs were developed based on non‐overlapping (NOL), completely overlapping (OL), and intermediates between NOL and OL genotypes. Three prediction models were assessed (one based on phenotypes only, while two had genomic data). All the three models had a higher predictive ability and a lower mean square error (MSE) when a large training set was used. Predictive ability increased and MSE reduced when genotype‐by‐environment interaction (G × E) was modeled for the same training set sizes and allocations. Predictive ability decreased while MSE increased with the increasing OL genotypes across the environments, suggesting that only a few OL genotypes may be required to set up METs for model training. Sparse testing using a model incorporating G × E could be implemented to reduce cost of phenotyping in cassava METs. If data were available, integrating crop growth models (CGMs) with genomic prediction holds the potential to improve predictive ability. The training population used for sparse testing could be optimized to determine the optimal size and distribution of genotypes to increase the predictive ability and reduce cost under a fixed budget.