Genomic selection holds the potential to serve as a strategic tool to enhance the genetic gain of complex traits in Miscanthus breeding programs. The development of improved cultivars requires their assessment for various traits across diverse environments to ensure suitable overall performance. Hence, the multi-trait multi-environment (MTME) genomic prediction (GP) models offer an opportunity to improve selection accuracy. This study aims to evaluate the potential of five GP models: (1) three MTME models including genotype-by-trait-by-environment interaction (G×E×T) and (2) two single-trait multi-environment (STME) models (with and without G×E interaction). A Miscanthus sacchariflorus population comprising 336 genotypes evaluated in three environments and scored for four traits (biomass yield YDY, total culm number TCM, average internode length AIL, and culm node number CNN) was analyzed. The predictive ability of the models was evaluated considering three cross-validation schemes resembling realistic scenarios (CV1: predicting new genotypes, CVP: predicting missing traits in a given environment, and CV2: predicting partially observed genotypes). On average, in all cross-validation schemes compared to the STME the predictive ability of the MTME models was 10% to 70% higher for TCM and AIL. On the other hand, for YDY and CNN, both STME models performed similarly or slightly better (between 5 to 64%) than the MTME models in most environments. While the MTME models were not successful for all traits when compared to their STME counterparts, MTME models improved the prediction of the performance of genotypes that were untested across environments or lacked trait information in a specific environment. Overall, our study suggests that MTME GP models can be implemented in Miscanthus breeding programs to improve the predictive ability of the complex traits, shorten breeding cycles, and accelerate selection decisions.
The presence of non-informative markers in Genome Wide Selection needs to be evaluated so that the genomic prediction is more efficient in a breeding program. This study proposes to evaluate the efficiency of genomic prediction after reducing the dimensionality of SNP's markers in the presence of different levels of dominance, heritability, and epistatic interactions to demonstrate that the results obtained with reduced information improve prediction and preserve the same biological conclusions when using a larger data set. Ten F2 populations of a diploid species with an effective size of 1000 individuals were simulated, involving the random combination of 2000 gametes generated from contrasting homozygous parents. Ten linkage groups with a size of 100 cM each and comprised 2010 bi-allelic SNP´s distributed equally and equidistant form. Nine traits were simulated, formed by different degrees of dominance, heritability, and epistatic interactions. The dimensionality reduction was performed randomly in the simulated population and then the efficiency of genomic prediction was tested in two different studies. The parameters square of correlation ( ), root mean squares error (RMSE), and the Akaike Information Criterion (AIC) was used to evaluate the efficiency of the model used in the RR-BLUP. The results obtained from the reduced information predicted by the RR-BLUP were able to improve the prediction and preserve the same biological conclusions when using a larger data set. Non-informational or small effect markers can be removed from the original data set. The inclusion of dominance effects was an efficient strategy to improve predictive capacity.
Genomic selection (GS) estimates the GEBV from genome-wide markers to reduce generation intervals and optimize germplasm selection, which is particularly advantageous for high-cost or late-expressed traits. While models like GBLUP are popular, they assume a polygenic architecture. In contrast, the Bayesian alphabet and machine learning (ML) can accommodate other types of genetic architectures. Given that no single model is universally optimal, stacking ensembles, which train a meta-model using predictions from diverse base learners, emerge as a compelling solution. However, the application of stacking in GS often overlooks non-additive effects. This study evaluated different stacking configurations for genomic prediction across 10 simulated traits, covering additive, dominance, and epistatic genetic architectures. A 5-fold cross-validation scheme was used to assess predictive ability and other evaluation metrics. The stacking approach demonstrated superior predictive ability in all scenarios. Gains were especially pronounced in complex architectures (100 QTLs, h2 = 0.3), reaching an 83% increment over the best individual model (BayesA with dominance), and also in oligogenic scenarios with epistasis (10 QTLs, h2 = 0.6), with a 27.59% gain. The success of stacking was attributed to two key strategies: base learner selection and the use of robust meta-learners (such as principal component or penalized regression) that effectively handled multicollinearity.
Genome-wide association studies (GWAS) are essential for identifying genomic regions associated with agronomic traits, but Linear Mixed Model (LMM)-based GWAS face challenges in capturing complex gene interactions. This study explores the potential of machine learning (ML) methodologies to enhance marker identification and association modeling in plant breeding. Unlike LMM-based GWAS, ML approaches do not require prior assumptions about marker–phenotype relationships, enabling the detection of epistatic effects and non-linear interactions. The research sought to assess and contrast approaches utilizing ML (Decision Tree—DT; Bagging—BA; Random Forest—RF; Boosting—BO; and Multivariate Adaptive Regression Splines—MARS) and LMM-based GWAS. A simulated F2 population comprising 1000 individuals was analyzed using 4010 SNP markers and ten traits modeled with epistatic interactions. The simulation included quantitative trait loci (QTL) counts varying between 8 and 240, with heritability levels set at 0.5 and 0.8. These characteristics simulate traits of candidate crops that represent a diverse range of agronomic species, including major cereal crops (e.g., maize and wheat) as well as leguminous crops (e.g., soybean), such as yield, with moderate heritability and a high number of QTLs, and plant height, with high heritability and an average number of QTLs, among others. To validate the simulation findings, the methodologies were further applied to a real Coffea arabica population (n = 195) to identify genomic regions associated with yield, a complex polygenic trait. Results demonstrated a fundamental trade-off between sensitivity and precision. Specifically, for the most complex trait evaluated (240 QTLs under epistatic control), Ensemble methods (Bagging and Random Forest) maintained a Detection Power (DP) exceeding 90%, significantly outperforming state-of-the-art GWAS methods (FarmCPU), which dropped to approximately 30%, and traditional Linear Mixed Models, which failed to detect signals (0%). However, this sensitivity resulted in lower precision for ensembles. In contrast, MARS (Degree 1) and BLINK achieved exceptional Specificity (>99%) and Precision (>90%), effectively minimizing false positives. The real data analysis corroborated these trends: while standard GWAS models failed to detect significant associations, the ML framework successfully prioritized consensus genomic regions harboring functional candidates, such as SWEET sugar transporters and NAC transcription factors. In conclusion, ML Ensembles are recommended for broad exploratory screening to recover missing heritability, while MARS and BLINK are the most effective methods for precise candidate gene validation.
Integrating genomic and environmental information holds the potential for enhancing the predictive power of genomic prediction models when accounting for the genotype-by-environment interactions. Hence, incorporating environmental covariates (EC) into these models can significantly influence their predictive accuracy. In this study, we utilized 1379 genotypes from the SoyNAM dataset, evaluated across four environments and genotyped with 4611 single-nucleotide polymorphism markers, to compare models incorporating genotype-by-environment and genotype-by-environmental covariate interactions using different covariance matrices. We evaluated four approaches: summarizing EC by averaging (AVG), filtering ECs based on a coefficient of determination criterion (FILT), segmenting ECs by crop phenology (STG), and a naïve approach that utilized all available information (ALL). Predictive ability was assessed as the Pearson's correlation between the genomic estimated breeding values and the adjusted phenotypes considering 10 replicates of three cross-validation scenarios (CV2: predicting tested genotypes in observed environments; CV1: untested genotypes in observed environments; CV0: tested genotypes in novel environments). Incorporating EC information into the models increased average predictive ability from 0.42 to 0.56 for CV1 and CV2. In these cases, the predictive ability was lower when EC information was averaged to compute the environmental kinship matrix, with slight differences observed with respect to the other approaches. Regarding the CV0 scheme, the model incorporating only genotype-by-environment information performed better (0.33). The naïve method, which utilized all available EC information (ALL), proved to be a promising approach, as it effectively improved the results in these scenarios while eliminating the need for additional steps in selecting variables.
The main approach for improving multiple traits simultaneously is the selection index. The most widely used selection indices are those based on factor analysis, which overcome statistical limitations such as multicollinearity and the reliance on arbitrary weights of the classical Smith–Hazel approach and support multi-environment trials. Nevertheless, the efficiency indices are affected by factors such as genotype number, environment and trait correlation, and heritability. In this study, we simulated different scenarios varying the mentioned factors to evaluate the performance of the Factor-Analysis and Ideotype-Design-Based Index (FAI-BLUP), Multi-trait Genotype–Ideotype Distance Index (MGIDI), and Multi-Trait Stability Index (MTSI). All correlations were positive and constant within each scenario, while the ideotype sought genetic gains for traits in opposite directions. Simulations were conducted using AlphaSimR and FieldSimR, and indices were implemented via the metan package. Results showed that index efficiency was higher in scenarios with larger numbers of genotypes, low-to-moderate trait correlations, and moderate-to-high inter-environment correlations. However, strong correlations among traits, particularly when combined with high heritability, compromise selection index efficiency in scenarios with antagonistic trait objectives. Despite that, the MGIDI consistently outperformed the other indices across most scenarios. Therefore, we emphasize accounting for trait genetic architectures, genotype–trait correlations, and target environment correlations.
Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in conventional perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse bio-based products. Increasing biomass yield will increase profitability and environmental benefits, so it is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed implementing sparse testing designs. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; G×E interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for G×E interaction consistently presented the highest PA and the lowest MSE: for CNN (PA: ~0.77, MSE: ~0.5) and YDY (PA: ~0.70, MSE: ~1.3) while for TCM and AIL these ranged from ~0.28 to 0.41 and ~1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.
The objective of this work was to investigate the use of Extreme Learning Machines (ELM) for the genomic prediction of rust resistance in Coffea arabica. With the objective of identifying an effective predictive model for the selection of resistant genotypes, ELM was compared to Artificial Neural Networks (ANN) and Bayesian Generalized Linear Regression (GBLR) in terms of accuracy measures and computational time. To this end, an F2 population of 245 C. arabica plants genotyped with 137 markers was used to evaluate the application of ELM for the genomic prediction of coffee rust resistance. The results indicate that ELM and ANN show a higher accuracy-on average 15% greater than that of GBLR-in predicting rust resistance. Additionally, ELM proves to be computationally more efficient, with a processing speed 5.5 and 19.45 times slower than that of ANN and BGLR, respectively, making it promising for large-scale analyses.
ABSTRACT: Wheat is a crop of significant global and Brazilian economic importance. The success of cultivation depends on several factors. Studying genotype by environment (G x E) interaction is fundamental for developing more productive and resilient varieties. Due to the climatic diversity, it is important to understand how weather covariates (WC) interact with the environments under study. To understand the complex relationships between WC and environments, it is possible to perform a graphical analysis based on the singular value decomposition (SVD) of the matrix composed by the values of the WC evaluated in the environments under study. This study proposed graphical analysis to understand the interaction between WC and evaluated environments. We used data from a randomized block design (four replications) assessed 42 wheat genotypes across 10 environments (2 years and 5 Rio Grande do Sul locations). Temperature (°C), relative humidity (%), and precipitation (mm) data were obtained from the NASA POWER project, using the experiment’s years and coordinates. The visualization of the interaction between weather covariates and environments revealed the WC that most influenced the performance of the genotypes. The environment E6 and E8 stood out for their high productivity, mainly associated with precipitation and relative humidity. This tool allows breeders to visualize the interconnection between different cultivation scenarios, obtaining valuable information for more accurate and effective decisions.
Plant genetic improvement has prioritized morphological traits capable of increasing plant productivity, adaptation, and resistance to biotic and abiotic stresses. Among these traits in soybean, phenotypic descriptors evaluated at juvenile stages stand out for enabling rapid, efficient, and low-cost analyses. In this context, the present study aimed to estimate the repeatability coefficients of morphological traits in soybean plants and to determine the minimum number of measurements required to predict genotypic values with high reliability. Three experiments were conducted under greenhouse conditions. A total of 34 soybean genotypes were evaluated in a randomized complete block design with four replications, in which each experimental unit consisted of the mean of two plants grown in a 3 dm³ pot. The variables epicotyl length, internode length between the unifoliolate node and the first trifoliolate leaf node, petiole length, rachis length of the first trifoliolate leaf, and plant height were assessed at the V2 and V3 stages. Repeatability coefficients were estimated, and the minimum number of measurements required to achieve predetermined levels of determination were established. The results indicated that the epicotyl length, internode length, and plant height showed higher repeatability coefficients compared with petiole length and rachis length of the first trifoliolate leaf. And, with four measurements, it was possible to predict genotype values for epicotyl length and plant height with 90% reliability in all experiments and at both evaluated developmental stages (V2 and V3), and with an increase to seven measurements, this level of reliability was achieved for internode length. Thus, the identification of the minimum number of plants that should be evaluated to reliably represent each variable, providing important support for experimental planning in soybean breeding programs.
Machine learning was applied to predict Arabica and Robusta coffee prices 2-6 months ahead using climatic, production, and economic data from 2009 to 2025. SHAP analysis revealed that global and Brazilian stock levels, Vietnamese drought, Colombian rainfall, and U.S. dollar exchange rate were the most influential drivers of price variation, showing that even features with weak simple correlations can have high predictive power. We compared Multilayer Perceptron (MLP), Extreme Learning Machine (ELM), Random Forest (RF), XGBoost, and two stacking configurations (RF/XGBoost/ELM) and (RF/XGBoost), both combined through a meta-learner. For Robusta, the three-model stacking configuration (RF/XGBoost/ELM) achieved the best overall accuracy with Pearson's correlation coefficient (PA) of 0.9264, Hit Rate (HR) of 92.8000%, Root-Mean-Square Error (RMSE) of 0.0793, and Mean Absolute Percentage Error (MAPE) of 9.3051%. Arabica proved more difficult to forecast, and the ELM delivered the highest independent test performance with PA of 0.9064 and HR of 94.1167%. To assess model robustness under future warming, we also tested two extreme climate scenarios derived from historical data and literature for major Arabica and Robusta producing regions (Brazil, Colombia, and Vietnam). These scenarios combined higher temperatures, reduced production, lower stocks, expanded production area, increased global consumption, and concurrent logistic and climatic shocks. Despite these stringent assumptions, the models did not project strong price surges and showed greater difficulty when extrapolating Arabica prices, while Robusta forecasts remained more consistent with realistic market values. These results underscore the ability of machine-learning models and SHAP interpretation to reveal the complex economic, climatic and productive factors governing global coffee price dynamics and highlight their conservative extrapolation when confronted with unprecedented climatic conditions.
Our aim was to assess the reduction of computational demand compared to traditional multivariate analyzes, as well as identify latent variables with biological relevance (pseudo-phenotypes) through a factor analysis for morphological and gait traits in MM horses, estimating their genetic parameters within a Bayesian framework. A total of 15 morphological and performance traits were collected from 50,172 horses: withers height, croup height, thoracic perimeter, cannon bone perimeter, head length, neck length, back and loin length, croup length, shoulder length, body length, head width, croup width, morphofunctional points, gait and total points. The dataset was provided by the Brazilian Association of Mangalarga Marchador Horse Breeders (ABCCMM). A factor analysis was performed and six latent factors that explained 73 % of the total variance of the data were extracted. From them, four factors showed biological relevance, namely: "height" (F1), "gait ability" (F2), "robustness" (F3), and "head size" (F4), and were explored in the genetic analysis. The heritability estimates for F1 and F4 were of moderate to high magnitudes (0.47 f 0.10 and 0.39 f 0.10 for F1 and F4, respectively). The genetic correlations ranged from zero to high magnitudes, with notable values observed between F1 and F4 (0.80 f 0.00), F1 and F3 (0.44 f 0.01), and F3 and F4 (0.33 f 0.01). The factor analysis proved to be an adequate alternative for reducing the dimensionality of variables and, consequently, the computational demand in the genetic evaluation of MM horses. Moderate heritabilities indicate the possibility of effective selection of horses based on these traits. In general, the correlations suggest that the selection of factors may result in genetic gains ranging from negative to high for other factors. Therefore, two factors are recommended to be used as selection criteria in MM horse breeding programs, which could contribute to the genetic improvement of the breed.
The phytonematode Meloidogyne paranaensis is one of the main threats to coffee production. The development of Coffea arabica cultivars resistant to this pathogen is an urgent demand for coffee growers. Progenies derived from the wild germplasm Amphillo are considered potential sources of resistance to M. paranaensis, however the mechanisms involved in this resistance have not yet been elucidated. In the present work, the resistance of different progenies derived from Amphillo was studied and molecular markers associated with resistance were identified. Through Genomic-Wide Association, SNP markers associated with genes potentially involved in resistance control were identified. A total of 158 genotypes belonging to four progenies derived from crosses between Amphillo and Catuaí Vermelho were analyzed. These coffee plants were phenotyped for five traits related to resistance. A total of 7116 SNP markers were genotyped and, after quality filtering, 931 SNPs were selected to conduct the genome-wide association study. The mixed linear model identified 12 SNPs with significant associations with at least one of the evaluated variables and eighteen genes were mapped. The results obtained support the development of markers for assisted selection, studies on genetic inheritance, and elucidating molecular mechanisms involved in the resistance of C. arabica to M. paranaensis.
This study provided a comprehensive overview of the behavior of alfalfa genotypes in response to environmental variations. We utilized established methods from literature and examined the unique aspects of each approach to collectively create a criterion for recommending cultivars. To this end, seventy-seven genotypes were cultivated with 24 consecutive cuts (months), during two years. Adaptability and stability analyses were conducted using multiple information estimates. The results indicated no significant effect of the genotypes, but there were significant effects from the environment and the genotype 21 was identified as the most promising due to its superior dry matter yield, predictable performance, and responsiveness to environmental changes across various cuts. Combining data and thoroughly describing the behavior of alfalfa genotypes has proven to be an effective method for studying their adaptability and stability.
Developing new cultivars, particularly in perennial species like Coffea arabica, can be a time-consuming process. Employing molecular markers in genome-wide selection (GWS) for predicting genetic values offers an alternative to accelerate this process. However, implementing GWS typically involves genotyping many markers for both training and candidate individuals, which can increase the total genotyping cost for the breeding program. Therefore, this study aimed to assess the feasibility of using low-density marker panels to predict the genetic merit of C. arabica for a range of desirable agronomic traits. For this purpose, GWS analyses were performed using the G-BLUP method with panels of varying marker densities, selected based on marker effect magnitude. The results indicate that employing lower-density panels might be advantageous for this species' improvement. Models based on these panels yielded accurate predictions for various traits and demonstrated high agreement in terms of selected individuals compared to more complex models.
Plant breeders utilize the additive main effects and multiplicative interaction (AMMI) model for analyzing yield data from multi-environment trials (METs) to visualize interaction patterns between genotypes and environments. AMMI-based selection indexes, such as the weighted average of absolute scores (WAAS) and the weighted average of absolute scores combining yield (WAASY), guide breeders in identifying superior varieties within METs. Despite being powerful, the frequentist approach of AMMI model and its derived indices presents challenges for identifying genotypes and environments, causing significant genotype-by-environment (G x E) interactions. This study built upon the Bayesian AMMI framework to allow to perform inferences on AMMI-based selection indexes. The Bayesian versions of WAAS and WAASY (Bayesian weighted average of absolute scores and Bayesian weighted average of absolute scores combining yield) were compared with the frequentist approach. A novel stability measure (SM), using Mahalanobis distance, was also proposed and integrated with yield performance into a graphical tool called the stability Mahalanobis trait (SMT) plot. Nine maize genotypes evaluated for grain yield across 20 environments were analyzed. The B-WAAS, B-WAASY, and SM indexes provided informative statistical inference through posterior distribution and credible intervals (highest posterior density [HPD]). HPD intervals allowed grouping similar genotypes based on stability and performance, offering reliable information for selection and recommendation. The SMT plot allows a direct comparison to an ideal scenario of high stability, facilitating the identification of genotypes aligned with breeding goals by analyzing the four quadrants. Genotypes in quadrant IV, exhibiting both high yield and high stability, are particularly valuable for breeding programs.
Soybean drought tolerance relies on root traits. Genomic prediction (GP) offers a non-destructive alternative to laborious phenotyping. This study explores a multi-kernel GP approach for predicting soybean root traits by also incorporating easily measurable non-destructive aerial traits as secondary covariates. The main idea is to leverage the correlation between the aerial (visible) and root traits (not visible). In addition, we contrasted the predictive ability (PA) shown by the multi-kernel approach to those obtained from single-trait and multi-trait genomic prediction models. Data comprising 100 cultivars evaluated in two years and genotyped for 5,403 single-nucleotide polymorphism markers was analyzed. To comprehensively assess model performance, two cross-validation schemes were considered (CV1 and CV0). CV1 used a five-fold approach, and CV0 used a time-lagged cross-validation (i.e., data from years 1 and 2 were used for training and testing, respectively). Aerial traits added as covariates enhanced the GP predictive ability for all the traits and cross-validation (CV) schemes, outperforming single- and multi-trait models without this information. The inclusion of the interaction term between markers and secondary traits did not improved PA compared to the main effects models.
The use of Euterpe edulis for the sustainable management of its fruits requires the development of superior genotypes. As a native species, the commercial plantations are challenged by limited information on seedling production, planting, cultural practices and production process. This study aimed to generate initial databases for conducting E. edulis breeding focused on fruit characteristics. A total of 499 plants from a commercial population (a cultivated group of plants grown specifically for commercial production purposes) were evaluated for nine characteristics of fruit. High diversity was detected, through the restricted maximum likelihood method and the prediction of genotypic values using the best linear unbiased prediction. Six groups were detected. The repeatability estimates ranged from medium to high, indicating accuracy in predicting the true value. The average difference between the base population (the original group of plants used for selection) and the selected population reached 97.40
ABSTRACT Regional heritability mapping (RHM) employs mixed-model and single-genomic region approaches for Genome-Wide Association Studies (GWAS). Although it utilizes marker groups and possesses a high detection power for identifying markers associated with target phenotypes, the significant linkage disequilibrium (LD) among markers limits the RHM's capacity to estimate the effects of one genomic region simultaneously. This constraint may hinder its capacity to detect complex associations or relationships among various genomic regions. Within a Bayesian framework, RHM can operate with multiple structured covariance matrices and can incorporate several genomic regions into a single model, effectively utilizing the LD present in these regions. In this study, our objectives were: (i) to propose the simultaneous estimation of multiple genomic region effects using a Bayesian model; (ii) to compare the efficiency of this simultaneous estimation with that of the single-region estimation using simulated data, focusing on the detection of significant genomic regions for phenotypes characterized by diverse genetic architectures; and (iii) to demonstrate the applicability of these models in breeding programs, particularly applying them to rice data. The results indicated that the simultaneous estimation of genomic region effects using the Bayesian approach offered greater detection power for more complex traits in simulated data. In the rice dataset, the simultaneous estimation method identified more regions than those previously reported in the literature, as well as newly uncovered genomic regions that deserve further investigation in post-GWAS analyses. This methodology holds promise for exploring and applying new genomic regions associated with target traits.
This study focused on incorporating dimensionality reduction based on marker significance to better harness the potential of machine learning for genomic prediction in different trait-genomic structures. The aim was to show that outcomes achieved with reduced data would improve predictive accuracy ( ) and precision (root-mean-square error: RMSE) while reducing computational time. Distinct subsets of markers, in simulated data, were chosen by prioritizing importance via the Bagging technique. Predictive modelling was subsequently conducted using both Bagging and the diverse architectures of a Multilayer Perceptron (MLP) neural network. This study was carried out with six traits of an F2 simulated population (derived from contrasting homozygotes) with 1,000 individuals. Three traits had three different heritabilities (0.4, 0.6, and 0.8) and were controlled by a set of 40 quantitative trait loci (QTLs). Additionally, four QTLs with more pronounced heritability effects (set at unity) were introduced in three other traits while preserving the same genetic control structure as the earlier traits. In our investigation, as the number of markers increased, both techniques gradually increased training time; however, the time needed for computation notably extended beyond the threshold of 100 markers for Bagging. In comparison to the MLP model, the Bagging model generally obtained better accuracy (higher ) and precision (lower RMSE) values regardless of heritability and added QTLs. Most importantly, results highlight that for traits subject to robust genetic control of additional QTLs, MLP networks experienced a decline in prediction performance from a few markers (~10). In contrast, Bagging kept constant or subtly improved predication performance. Finally, the dimensionality reduction procedure effectively improves genomic prediction, and Bagging captures complex genetic control structures for prediction better than MLP networks.