The increasing availability of multi-omics data is promising in enhancing genomic prediction in breeding and human genetics. However, integrating multi-omics data into genomic prediction models remains challenging due to complex relationships between omics layers and phenotypic outcomes. We propose Fusion Similarity Best Linear Unbiased Prediction (FSBLUP), a novel strategy that integrates genomic and intermediate omics data using a unified similarity matrix approach. FSBLUP systematically estimates how different omics layers contribute to phenotypic variation via machine-learning-optimized parameters that capture underlying genetic architecture of complex traits. FSBLUP demonstrates greater predictive accuracy than existing methods, as validated through theoretical and practical evaluations.
BACKGROUND:Multibreed genomic prediction (MBGP) is crucial for improving prediction accuracy for breeds with small populations, for which limited data are often available. Recent studies have demonstrated that partitioning the genome into nonoverlapping blocks to model heterogeneous genetic (co)variance in multitrait models can achieve higher joint prediction accuracy. However, the block partitioning method, a key factor influencing model performance, has not been extensively explored. RESULTS:We introduce mbBayesABLD, a novel Bayesian MBGP model that partitions each chromosome into nonoverlapping blocks on the basis of linkage disequilibrium (LD) patterns. In this model, marker effects within each block are assumed to follow normal distributions with block-specific parameters. We employ simulated data as well as empirical datasets from pigs and beans to assess genomic prediction accuracy across different models using cross-validation. The results demonstrate that mbBayesABLD significantly outperforms conventional MBGP models, such as GBLUP and BayesR. For the meat marbling score trait in pigs, compared with GBLUP, which does not account for heterogeneous genetic (co)variance, mbBayesABLD improves the prediction accuracy for the small-population breed Landrace by 15.6%. Furthermore, our findings indicate that a moderate level of similarity in LD patterns between breeds (with an average correlation of 0.6) is sufficient to improve the prediction accuracy of the target breed. CONCLUSIONS:This study presents a novel LD block-based approach for multibreed genomic prediction. Our work provides a practical tool for livestock breeding programs and offers new insights into leveraging genetic diversity across breeds for improved genomic prediction.
Nonparametric models have recently been receiving increased attention due to their effectiveness in genomic prediction for complex traits. However, regular nonparametric models cannot effectively differentiate the relative importance of various SNPs, which significantly impedes the further application of these methods for genomic prediction. To enhance the fitting ability of nonparametric models and improve genomic prediction accuracy, a weighted kernel ridge regression model (WKRR) was proposed in this study. For this new method, different weights were assigned to different SNPs according to the p-values from GWAS, and then a KRR model based on these weighted SNPs was constructed for genomic prediction. Cross-validation was further adopted to choose appropriate hyper-parameters during the weighting and prediction process for generalization. We compared the predictive accuracy of WKRR with the genomic best linear unbiased prediction (GBLUP), BayesR, and unweighted KRR using both simulated and real datasets. The results showed that WKRR outperformed unweighted KRR in all simulated scenarios. Additionally, WKRR achieved an average improvement of 1.70% in accuracies across all traits in a mice dataset and 2.17% for three lactation-related traits in a cattle dataset compared to GBLUP, and yielded competitive results compared to BayesR. These findings demonstrated the great potential of weighted nonparametric models for genomic prediction.
Regulatory interactions across biological layers are highly complex. Intermediate omics (e.g., transcriptomes) can reveal how DNA variation influences traits, but raw omics measurements are noisy and commonly mix genetic and non-genetic signals. We developed the Genetically Regulated Additive and Dominance model (GRAD), a single-step framework. GRAD first decomposes each omics feature into genetically regulated additive and dominance components, and then jointly models additive and dominance effects at both the genomic and intermediate-omics levels to predict complex traits. We evaluated GRAD using simulations with dominance contributions ranging from 0 to 50% of genetic variance, and using a Landrace pig dataset (1 494 genotyped individuals, of which 105 had whole-blood transcriptome sequencing). In simulations, GRAD increased up to 21.98% (4.73% on average) accuracy gain relative to standard genomic best linear unbiased prediction, and it also outperformed models that used raw omics data or only additive omics signals. In the pig data, GRAD improved prediction for age to 100 kg and total number of piglets born by 7.4 and 14.9%, respectively. Our findings demonstrate that explicitly extracting genetically regulated additive and dominance signals from omics data increases the genetic information available for prediction. However, practical benefit depends strongly on the proportion of animals profiled and on sampling tissues that are mechanistically relevant to the target trait.
Pigs are one of the most essential sources of high-quality proteins in human diets.Structural variants(SVs)are a major source of genetic variants associated with diverse traits and evolutionary events.However,the current linear reference genome of pigs restricts the accurate presentation of position information for SVs.In this study,we generated a pangenome of pigs and a genome variation map of 599 deeply sequenced genomes across Eurasia.Additionally,we established a section-wide gene repertoire,revealing that core genes are more evolutionarily conserved than vari-able genes.Furthermore,we identified 546,137 SVs,their enrichment regions,and relationships with genomic features and found significant diver-gence across Eurasian pigs.More importantly,the pangenome-detected SVs could complement heritability estimates and genome-wide associa-tion studies based only on single nucleotide polymorphisms.Among the SVs shaped by selection,we identified an insertion in the promoter region of the TBX19 gene,which may be related to the development,growth,and timidity traits of Asian pigs and may affect the gene expression.The constructed pig pangenome and the identified SVs in this study provide rich resources for future functional genomic research on pigs.
As the scale of deep whole-genome sequencing (WGS) data has grown exponentially, hundreds of millions of single nucleotide polymorphisms (SNPs) have been identified in livestock. Utilizing these massive SNP data in population stratification analysis, ancestry prediction, and breed diversity assessments leads to overfitting issues in computational models and creates computational bottlenecks. Therefore, selecting genetic variants that express high amounts of information for use in population diversity studies and ancestry inference becomes critically important. Here, we develop a method, HITSNP, that combines feature selection and machine learning algorithms to select high-representative SNPs that can effectively estimate breed diversity and infer ancestry. HITSNP outperforms existing feature selection methods in estimating accuracy and computational stability. Furthermore, HITSNP offers a new algorithm to predict the number and composition of ancestral populations using a small number of SNPs, and avoiding calculating the number of clusters. Taken together, HITSNP facilitates the research of population structure, animal breeding, and animal resource protection.
Pigs play a central role in human livelihoods in China, but a lack of systematic large-scale whole-genome sequencing of Chinese domestic pigs has hindered genetic studies. Here, we present the 1000 Chinese Indigenous Pig Genomes Project sequencing dataset, comprising 1011 indigenous individuals from 50 pig populations covering approximately two-thirds of China's administrative divisions. Based on the deep sequencing (similar to 25.95x) of these pigs, we identify 63.62 million genomic variants, and provide a population-specific reference panel to improve the imputation performance of Chinese domestic pig populations. Using a combination of methods, we detect an ancient admixture event related to a human immigration climax in the 13(th) century, which may have contributed to the formation of southeast-central Chinese pig populations. Analyzing the haplotypes of the Y chromosome shows that the indigenous populations residing around the Taihu Lake Basin exhibit a unique haplotype. Furthermore, we find a 13 kb region in the THSD7A gene that may relate to high-altitude adaptation, and a 0.47 Mb region on chromosome 7 that is significantly associated with body size traits. These results highlight the value of our genomic resource in facilitating genomic architecture and complex traits studies in pigs.
The design of breeding programs is crucial for maximizing economic gains. Simulation provides the most efficient measures to test these programs, as real-world trials are often costly and time-consuming. We developed GOplan, a comprehensive and user-friendly R package designed to develop animal breeding programs considering pure-bred populations and crossbreeding systems. Compared with other traditional simulators, it has mainstream crossbreeding frameworks that streamline modeling and use Gene Flow and Bayesian optimization methods to enhance breeding program efficiency. GOplan includes 3 key functions: runCore() to evaluate the effects of nucleus breeding programs, runWhole() to predict economic outcomes and the production performance of crossbreeding systems, and runOpt() to optimize crossbreeding structures for greater profitability. These functions support breeders in better planning and accelerating breeding goals. Additionally, the application of Bayesian optimization algorithms in this study provides valuable insights for developing new optimization algorithms in the future. The software is available at https://github.com/CAU-TeamLiuJF/GOplan.
There is an increasing understanding that a reference genome representing an individual cannot capture all the gene repertoire of a species. Here, we conduct a population-scale missing sequences detection of Chinese domestic pigs using whole-genome sequencing data from 534 individuals. We identify 132.41 Mb of sequences absent in the reference assembly, including eight novel genes. In particular, the breeds spread in Chinese highaltitude regions perform significantly different frequencies of new sequences in promoters than other breeds. Furthermore, we dissect the role of non-coding variants and identify a novel sequence inserted in the 3'UTR of the FMO3 gene, which may be associated with the intramuscular fat phenotype. This novel sequence could be a candidate marker for meat quality. Our study provides a comprehensive overview of the missing sequences in Chinese domestic pigs and indicates that this dataset is a valuable resource for understanding the diversity and biology of pigs.
MOTIVATION:Utilizing both purebred and crossbred data in animal genetics is widely recognized as an optimal strategy for enhancing the predictive accuracy of breeding values. Practically, the different genetic background among several purebred populations and their crossbred offspring populations limits the application of traditional prediction methods. Several studies endeavor to predict the crossbred performance via the partial relationship, which divides the data into distinct sub-populations based on the common genetic background, such as one single purebred population and its corresponding crossbred descendant. However, this strategy makes prediction inaccurate due to ignoring half of the parental information of crossbreed animals. Furthermore, dominance effects, although playing a significant role in crossbreeding systems, cannot be modeled under such a prediction model. RESULTS:To overcome this weakness, we developed a novel multi-breed single-step model using metafounders to assess ancestral relationships across diverse breeds under a unified framework. We proposed to use multi-breed dominance combined relationship matrices to model additive and dominance effects simultaneously. Our method provides a straightforward way to evaluate the heterosis of crossbreeds and the breeding values of purebred parents efficiently and accurately. We performed simulation and real data analyses to verify the potential of our proposed method. Our proposed model improved prediction accuracy under all scenarios considered compared to commonly used methods. AVAILABILITY AND IMPLEMENTATION:The software for implementing our method is available at https://github.com/CAU-TeamLiuJF/MAGE.
Excreta traits comprise a very important characteristic in breeding that have been neglected for a long time. With the growth of intensive pig farming, plenty of environment problems have been raised, and people have begun to pay attention to pig excreta behaviors from genetics and breeding perspectives. However, the genetic architecture of excreta traits remains unclear. To investigate the genetic architecture of excreta traits in pigs, eight excreta traits and feed conversion ratio (FCR) were analyzed in this study. We performed genome-wide association studies (GWASs) on 213 Yorkshire pigs and estimated genetic parameters for a total number of 290 pigs, comprising 213 Yorkshire, 52 Landrace and 25 Duroc. After analysis, eight and 22 genome-wide significant SNPs were detected for FCR and the eight excreta traits in single-trait GWASs separately, and 18 were detected in a multi-trait meta-analysis for excreta traits, six of which were detected in both the single-trait and the multi-trait GWAS. Eighty, 182 and 133 genes were detected within 1 Mb of the genome-wide significant SNPs for FCR, excreta traits and multi-trait meta-analysis, respectively. Five candidate genes (BCKDC, DBT, ANKRD7, SHPRH and HCRT) with biochemical and physiological effects relevant to feed efficiency and excreta traits might be interesting markers for future breeding. Meanwhile, functional enrichment analysis indicates that most of the significant pathways are associated with the glutathione catabolic process, DNA topological change and replication fork protection complex. This study reveals the architecture of excreta traits in commercial pigs and offers an opportunity for decreasing the pollution from excreta using genomic selection in pigs.
我国是世界第一养猪大国,但种猪的生产性能与国际先进水平仍有较大差距.为了提升生猪种业的核心竞争力,农业农村部先后颁布了 2 轮《全国生猪遗传改良计划》.国家生猪核心育种场遗传评估报告是了解我国生猪育种工作的重要途径,是推动遗传改良计划有序进行的重要措施.《国家生猪核心育种场年度遗传评估报告(2022 年度)》基于全国 92 家瘦肉型品种国家核心场、5 家地方品种核心场和 6 家国家核心种公猪站的相关数据,对 2022 年度国家核心场的育种进展等情况进行了详细全面地分析和阐述.本文对该报告中核心场数据量统计信息、重要经济性状指标、遗传评估、育种措施、场间关联等指标做了全面分析,以期为我国从事猪育种科研、产业相关人士及行业管理人员提供准确信息解读,完善相关育种工作,提升我国猪遗传改良效率.
本研究旨在探讨系谱错误对猪基因组选择的影响.模拟数据研究表明,随着系谱错误率增加,基因组选择估计育种值的准确性、无偏性和秩相关系数均逐步减小,20%系谱错误率相较 0%时准确性、无偏性和秩相关系数分别由 0.4233、0.1749、0.4099 降为 0.3582、0.1031、0.3469;当基因型检测个体数目增多,0%与 20%系谱错误率下基因组选择估计育种值准确性的差值缩小.研究表明,本研究所分析育种场群中系谱错误率约为 3.2%;应用基因组选择一步法对有表型及系谱记录的大白猪进行育种值估计的准确性为0.4471,利用基因型数据矫正部分系谱错误后,准确性提高 0.42%.以上结果表明,系谱错误的存在会降低基因组选择的育种值估计准确性,使育种值的估计无偏性变小;通过增加基因型检测数目可以减少系谱错误对基因组选择造成的负面影响.
With the broad application of genomic information, SNP-based measures of estimating inbreeding have been widely used in animal breeding, especially based on runs of homozygosity. Inbreeding depression is better estimated by SNP-based inbreeding coefficients than pedigree-based inbreeding in general. However, there are few comprehensive comparisons of multiple methods in pigs so far, to some extent limiting their application. In this study, to explore an appropriate strategy for estimating inbreeding depression on both growth traits and reproductive traits in a Large White pig population, we compared multiple methods for the inbreeding coefficient estimation based on both pedigree and genomic information. This pig population for analyzing the influence of inbreeding was from a pig breeding farm in the Inner Mongolia of China. There were 26,204 pigs with records of age at 100 kg (AGE) and back-fat thickness at 100 kg (BF), and 6,656 sows with reproductive records of the total number of piglets at birth (TNB), and the number of alive piglets at birth (NBA), and litter weight at birth. Inbreeding depression affected growth and reproductive traits. The results indicated that pedigree-based and SNP-based inbreeding coefficients had significant effects on AGE, TNB, and NBA, except for BF. However, only SNP-based inbreeding coefficients revealed a strong association with inbreeding depression on litter weight at birth. Runs of homozygosity-based methods showed a slight advantage over other methods in the correlation analysis of inbreeding coefficients and estimation of inbreeding depression. Furthermore, our results demonstrated that the model-based approach (RZooRoH) could avoid miscalculations of inbreeding and inbreeding depression caused by inappropriate parameters, which had a good performance on both AGE and reproductive traits. These findings might improve the extensive application of runs of homozygosity analysis in pig breeding and breed conservation.
Growth rate plays a critical role in the pig industry and is related to quantitative traits controlled by many genes. Here, we aimed to identify causative mutations and candidate genes responsible for pig growth traits. In this study, 2360 Duroc pigs were used to detect significant additive, dominance, and epistatic effects associated with growth traits. As a result, a total number of 32 significant SNPs for additive or dominance effects were found to be associated with various factors, including adjusted age at a specified weight (AGE), average daily gain (ADG), backfat thickness (BF), and loin muscle depth (LMD). In addition, the detected additive significant SNPs explained 2.49%, 3.02%, 3.18%, and 1.96% of the deregressed estimated breeding value (DEBV) variance for AGE, ADG, BF, and LMD, respectively, while significant dominance SNPs could explain 2.24%, 13.26%, and 4.08% of AGE, BF, and LMD, respectively. Meanwhile, a total of 805 significant epistatic effects SNPs were associated with one of ADG, AGE, and LMD, from which 11 sub-networks were constructed. In total, 46 potential genes involved in muscle development, fat deposition, and regulation of cell growth were considered as candidates for growth traits, including CD55 and NRIP1 for AGE and ADG, TRIP11 and MIS2 for BF, and VRTN and ZEB2 for LMD, respectively. Generally, in this study, we detected both new and reported variants and potential candidate genes for growth traits of Duroc pigs, which might to be taken into account in future molecular breeding programs to improve the growth performance of pigs.