In brief: Correlated regions of systemic interindividual epigenetic variation (CoRSIVs) are genomic regions with CpG-methylation patterns that differ between individuals, yet are consistent between tissues, within the same individual. Analyzing two groups of Holstein bull methylomes-nine with a high sire-conception rate (SCR) and nine with a low SCR-we found that a common type of CoRSIVs was significantly associated with reduced SCR and is thus suggested as a biomarker for SCR because it was highly methylated in sperm, but failed to retain hypermethylation in the gametes of males with low SCR. Abstract: Correlated regions of systemic interindividual epigenetic variation (CoRSIVs) are genomic regions with CpG-methylation patterns that differ between individuals, yet are consistent between tissues, within the same individual; therefore, their methylation can be profiled in bodily fluids that are easily obtained, such as blood and semen. Bearing in mind the simple epigenetic profiling of CoRSIVs, we tested whether this type of differentially methylated region (DMR) is associated with bovine fertility. Sequence Read Archive (SRA) meth BLAST was used to estimate CoRSIVs methylation status in 18 healthy, representative, and age-matched Holstein bulls, among which nine had high (H) sire-conception rate (SCR), and the other nine had low (L) SCR (group averages of SCR: 3.3 ± 0.6 and -3.8 ± 1.8, respectively). This method was also applied to morula and trophoblast SRA methylomes. Analysis with meth BLAST was effective for most (80%) CoRSIVs and showed that CoRSIVs are reprogrammed during blastocyst formation, although this method was incapable of specifically determining the methylation level in CoRSIVs with retrotransposons. In sperm, the effect of global methylation was evident in a common (25%) type of CoRSIVs that is highly (94.5% ± 4.3%) methylated in sperm. Specifically, a failure to retain hypermethylation in the sperm plus strand was significantly (p < 0.00025) indicative of low SCR. Comparing global DNA methylation using the latter type of CoRSIVs between sperm and blood can be used as a better biomarker for fertility than using other differentially methylated regions with more complex epigenetics.
We propose a nonparametric approach to testing conditional independence and estimating conditional association, generalizing the Cochran-Mantel-Haenszel (CMH) test and odds-ratio estimator to continuous sample spaces. It leverages a multiscale scanning approach to decompose the sample space into a cascade of 2× 2 × T tables. Following the CMH test, we condition on the marginal order statistics, which are "almost ancillary" regarding conditional dependency. This strategy helps overcome a key challenge faced by other methods that discretize the sample space: we achieve consistency without requiring stratum sample sizes to grow to infinity, a constraint often difficult to satisfy in practice. Our method produces easy-to-compute test statistics with a known asymptotic null distribution under the conditional sampling model, scaling almost linearly with the sample size. Our simulation results demonstrate reliable Type I error control, even with small samples and high-dimensional conditioning, and competitive power compared to state-of-the-art tests. Finally, a case study on Uber ride-share data highlights the method's unique dual capability, inherited from the CMH, to both test and identify the nature of the inferred conditional association. By providing summary statistics that capture the strength and direction of local associations, our method offers practitioners a useful tool for learning conditional dependencies.
Abstract While pangenome studies emphasize structural variations (SVs), their functional significance beyond single‐nucleotide polymorphisms (SNPs) in cattle remains unexplored. We established a comprehensive variation map for 1379 European cattle from nine breeds using a pangenome graph. Although SV and SNP distributions were similar genome‐wide, 19.3% of SVs exhibited low linkage disequilibrium (r² < 0.2) with nearby SNPs. These non‐SNP‐tagged SVs are preferentially localized to chromosome ends and promoters, and play crucial roles in artificial selection for production traits. Integrated analyses of large‐scale chromosomal recombination data and transposon‐mediated SV events revealed that terminal chromosomal recombination and strong selection on recurrent young SVs are likely to prevent SNP tagging. For SVs tagged by SNPs, they were found to play a dominant role in haplotype diversity and inter‐population differences. Genome‐wide association studies of over 50,000 dairy bulls using imputed SV + SNP data identified functional SVs missed by standard analysis, including an SV (chr6:32,772,14–32,772,409) showing strong milk yield association (p < 1 × 10−10) and divergent selection between dairy and beef breeds, demonstrating SVs' unique contribution to complex trait architecture. Overall, our study illustrates the independent or dominant roles of SV in the breeding processes of the European cattle, facilitating the use of pangenome SVs to recover missing heritability from SNP‐based research.
Pangenomes of several species have been assembled recently, facilitating the detection and genotyping of structural variants. As part of the FarmGTEx Project, we previously constructed a Holstein pangenome (H20D) based on 40 phased haploid assemblies. Here, we use this breed specific pangenome to genotype 93,059 structural variants from whole-genome sequences of 1,571 cattle. We then develop a Holstein pangenome variation imputation reference panel we name HolPIP. Leveraging HolPIP, we impute 86.65% (68,354/78,886) of structural variants for 50,299 bulls with Beagle R² ≥ 0.8. Using these imputed structural variants and phenotypes for 43 complex traits, we conduct GWAS, identifying 1,225 structural variant-trait associations. We next use fine-mapping to prioritize 32 high-confidence candidate structural variants, including a 75-bp deletion in ANKRD11 linked to dairy form, rump width, and stature, as well as an insertion in DHX32 associated with RNA metabolism. Compared to SNPs across various functional annotations, structural variants show a stronger genome-wide enrichment across most complex traits in cattle, suggesting that structural variants may have an important contribution to the genetic basis of dairy traits.
Microbiome compositional data are often high-dimensional, sparse, and exhibit pervasive cross-sample heterogeneity. We introduce the "logistic-tree normal" (LTN) model, a generative model that allows flexible covariance among the microbiome taxa, enables scalable computation, and effectively captures other key characteristics of microbiome compositional data such as the abundance of zeros. LTN incorporates a tree-based decomposition for effective aggregation over sparse taxa counts and models the relative abundance at the tree splits jointly using a (multivariate) logistic-normal distribution. The latent Gaussian structure allows a wide range of multivariate analysis and modeling tools for high-dimensional data-such as those enforcing sparsity or low-rank assumptions on the covariance structure-to be readily incorporated. As a general-purpose, fully generative model, LTN can be applied in a wide range of contexts, while at the same time, efficient computational recipes for Bayesian inference under LTN are available through conjugate blocked Gibbs sampling enabled by pólya-gamma augmentation. We demonstrate the use of LTN in a compositional mixed-effects model for differential abundance analysis through both numerical experiments and a reanalysis of the infant cohort in the DIABIMMUNE study. We explain and showcase through numerical experiments and the case study how LTN, through adequately accounting for the cross-sample heterogeneity, is capable of generating the appropriate proportion of zeros without incurring an explicit zero-inflation component. This confirms a recent viewpoint that "zero-inflation" in count-based sequencing data are often results of unaccounted cross-sample variation.
We develop unbalanced Haar (UH) wavelet tree ensembles for regression on triangulable manifolds. Given data sampled on a triangulated manifold, we construct UH wavelet trees whose atoms are supported on geodesic triangles and form an orthonormal system in L^2(μ_n), where μ_n is the empirical measure on the sample, which allows us to use UH trees as weak learners in additive ensembles. Our construction extends classical UH wavelet trees from regular Euclidean grids to generic triangulable manifolds while preserving three key properties: (i) orthogonality and exact reconstruction at the sampled locations, (ii) recursive, data-driven partitions adapted to the geometry of the manifold via geodesic triangulations, and (iii) compatibility with optimization-based and Bayesian ensemble building. In Euclidean settings, the framework reduces to standard UH wavelet tree regression and provides a baseline for comparison. We illustrate the method on synthetic regression on the sphere and on climate anomaly fields on a spherical mesh, where UH ensembles on triangulated manifolds substantially outperform classical tree ensembles and non-adaptive mesh-based wavelets. For completeness, we also report results on image denoising on regular grids. A Bayesian variant (RUHWT) provides posterior uncertainty quantification for function estimates on manifolds. Our implementation is available at http://www.github.com/hrluo/WaveletTrees.
Structural variants are an underexplored source of genetic diversity. As part of the FarmGTEx Project, here we report a Holstein breed-specific pangenome graph (H20D) using Minigraph-Cactus and 40 phased haploid assemblies from 20 cows. H20D outperforms both assembly- and read-based long-read callers, and far exceeds short-read approaches, identifying over 10,000 additional structural variants per sample. It also significantly improves structural variant detection and genotyping relative to graphs built across breeds or from fewer/unphased assemblies, with particular advantages in complex regions. Using H20D, we genotype variants in 173 cattle and performed a GWAS, where a larger fraction of structural variants than SNPs reach genome-wide significance, implicating them as potential causal variants. Together, these results demonstrate the power of phased, within-breed pangenome graphs for accurate SV genotyping and trait mapping in dairy cattle.
Cattle are integral to global food security, yet the molecular architecture of their complex traits remains poorly understood. Here, we present the Cattle Genotype–Tissue Expression (CattleG-TEx) Phase 1 resource (https://cattlegtex.farmgtex.org/), a substantial expansion of the pilot study. By leveraging 12,422 RNA-seq profiles across 43 tissues and 82 breeds, we characterized 433,972 primary and 161,428 non-primary regulatory effects spanning seven molecular phenotypes. This high-resolution atlas resolves 75% of GWAS signals for 44 complex traits, significantly addressing the "missing regulation" in livestock. We propose a genetic regulatory model demonstrating how variants across multiple biological layers interact with specific biological contexts to shape pheno-typic variation. Furthermore, CattleGTEx elucidates mechanisms underlying adaptive evolution between Bos taurus and Bos indicus, as well as artificial selection in dairy and beef breeds. Finally, by mapping evolutionary constraints on these regulatory effects, we demonstrate the translational value of this resource for prioritizing causal variants in human complex diseases. Together, Phase 1 of CattleGTEx provides a transformative framework for functional genomics, precision breeding, and comparative genetics.
Many modern statistical applications involve a two-level sampling scheme that first samples subjects from a population and then samples observations on each subject. These schemes often are designed to learn both the population-level functional structures shared by the subjects and the functional characteristics specific to individual subjects. Common wisdom suggests that learning population-level structures benefits from sampling more subjects whereas learning subject-specific structures benefits from deeper sampling within each subject. Oftentimes these two objectives compete for limited sampling resources, which raises the question of how to optimally sample at the two levels. We quantify such sampling-depth trade-offs by establishing the L-2 minimax risk rates for learning the population-level and subject-specific structures under a hierarchical Gaussian process model framework where we consider a Bayesian and a frequentist perspective on the unknown population-level structure. These rates provide general lessons for designing two-level sampling schemes given a fixed sampling budget. Interestingly, they show that subject-specific learning occasionally benefits more by sampling more subjects than by deeper within-subject sampling. We show that the corresponding minimax rates can be readily achieved in practice through simple adaptive estimators without assuming prior knowledge on the underlying variability at the two sampling levels. We validate our theory and illustrate the sampling trade-off in practice through both simulation experiments and two real datasets. While we carry out all the theoretical analysis in the context of Gaussian process models for analytical tractability, the results provide insights on effective two-level sampling designs more broadly.
The bovine liver is a highly compartmentalized organ that plays essential roles in continuous gluconeogenesis and nitrogen recycling; however, its spatial molecular architecture has remained largely uncharacterized due to the limitations of traditional bulk and single-cell approaches. To address this gap, Spatial Enhanced Resolution Omics-sequencing (Stereo-seq) was utilized to generate a subcellular-resolution (500 nm) transcriptomic map of an adult Holstein cattle liver, and a refined reference-guided workflow was implemented to overcome standard annotation limitations in livestock. Raw sequencing data were processed using the Stereo-seq Analysis Workflow and analyzed with Stereopy, Seurat, SingleR, and reference-guided workflows. Spatial aggregation was evaluated at Bin20, Bin50, Bin100, Bin150, and Bin200. Increasing bin size increased molecular identifier counts and detected-gene complexity while progressively reducing spatial granularity. Bin50, corresponding to 50 × 50 DNA nanoballs and an approximate nominal footprint of 25 × 25 µm, was therefore selected as a practical intermediate aggregation level for the primary analyses. Quality-control assessment, Leiden clustering, UMAP visualization, reference-based cell-type annotation, cluster-marker analysis, and spatial mapping of canonical hepatic genes demonstrated preservation of biologically interpretable liver transcriptional organization. Raw sequencing data processed spatial matrices, annotated objects, and analysis code are publicly available to support reanalysis and computational benchmarking. In summary, we present a Stereo-seq spatial transcriptomic resource generated from liver tissue of an adult Holstein cow. This initial resource provides a valuable foundation for future studies of bovine liver biology, comparative genomics, and the spatial basis of livestock health and production traits.
Understanding the genetic and molecular architecture of complex traits and artificial selection is crucial for advancing sustainable precision breeding in cattle and other livestock. Yet, how genetic variation affects cellular gene expression remains elusive in cattle. Here, by integrating 8,866 bulk RNA-seq samples and 999,192 single cells of 81 cell types in 22 bovine tissues, we presented a comprehensive atlas of regulatory variants at the cell type resolution in cattle. By colocalizing with bulk-tissue expression quantitative trait loci (beQTL), we detected 57,043 novel cell-type stratified eQTL and cell-type/state interaction eQTL in 18,153 genes, which also exhibited a stronger tissue/cell-type specificity than beQTL. By examining genome-wide associations (GWAS) of 44 complex traits, these cell-resolved eQTL were colocalized with 505 (24%) additional GWAS loci compared to beQTL. Through integrating this resource with selection signatures between dairy and beef cattle, we provided tissue/cell-specific regulatory insights into cattle breeding. Overall, the current atlas of cell-type-specific regulatory variants will serve as an invaluable resource for cattle genomics and selective breeding.
Flow matching (FM) is a family of training algorithms for fitting continuous normalizing flows (CNFs). Conditional flow matching (CFM) exploits the fact that the marginal vector field of a CNF can be learned by fitting least-squares regression to the conditional vector field specified given one or both ends of the flow path. In this paper, we extend the CFM algorithm by defining conditional probability paths along “streams”, instances of latent stochastic paths that connect data pairs of source and target, which are modeled with Gaussian process (GP) distributions. The unique distributional properties of GPs help preserve the “simulation-free" nature of CFM training. We show that this generalization of the CFM can effectively reduce the variance in the estimated marginal vector field at a moderate computational cost, thereby improving the quality of the generated samples under common metrics. Additionally, adopting the GP on the streams allows for flexibly linking multiple correlated training data points (e.g., time series). We empirically validate our claim through both simulations and applications to image and neural time series data.
Deciphering the regulatory syntax of the genome is essential to understand the genetic and molecular architecture of complex traits, as most trait-associated variants lie in non-coding regions. Yet, functional annotation of the bovine genome remains limited, hindering our ability to unravel the mechanisms underpinning complex traits of economic and ecological importance in cattle. Here, we present a comprehensive epigenetic atlas comprising 1,138 genome-wide epigenetic profiles, including chromatin accessibility, six histone modifications, CCCTC-binding factor (CTCF) transcription factor binding, DNA methylation, chromatin conformation, and transcriptomes across 53 adult tissues, five fetal tissues, and seven primary cell types. This atlas-level data enables us to annotate around 45% of the genome as putative regulatory elements exhibiting tissue- or cell-specific regulatory activity. Leveraging sequence-to-function deep learning models, we discovered 301 sequence motifs and predicted the functional impact of genetic variants through in silico mutagenesis, thereby facilitating the decoding of the regulatory syntax of the cattle genome and fine-mapping of GWAS loci for 22 complex traits. Cross-species analysis further revealed evolutionarily conserved features of regulatory architecture and provided evolutionary insights into complex traits and diseases in humans. Together, this atlas offers a foundational resource for advancing cattle functional genomics, sustainable breeding, and studies of regulatory evolution.
DDX3X neurodevelopmental disorder (DDX3X-NDD) represents a recently identified genetic syndrome characterized by intellectual disability (ID) and developmental delays, primarily caused by pathogenic variants in the DDX3X gene. The physiological ramifications of these mutations remain largely unexplored. In this study, we reported 21 DDX3X variants from 22 Chinese patients with DDX3X-NDD by whole exome sequencing. We selected five variants for further functional analyses, including two previously reported by our group. Three frameshift variants (c.280_281dup p.R95Efs*127, c.669_670del p.A224Pfs*70, and c.1579del p.H527Ifs*9) resulted in either the loss of DDX3X protein or the production of truncated proteins. Additionally, two missense variants (c.1051C > G p.R351G and c.1501G > A p.A501T) significantly reduced DDX3X protein expression. Notably, variants DDX3X-R95Efs*127 and DDX3X-A224Pfs*70 triggered marked apoptosis induction and failed to form stress granules in HEK293T cells compared to wild-type DDX3X. This defect may stem from their inability to interact with the stress particle marker PABPC1, as evidenced by co-immunoprecipitation assays. Moreover, DDX3X-H527Ifs*9 and DDX3X-R351G variants were found to disrupt the cell cycle, extending the S phase relative to the wild type. Collectively, our findings provide mechanistic insights into the pathogenic consequences of DDX3X-NDD associated mutations, suggesting that the loss-of-function variants of DDX3X lack a context-dependent survival advantage, potentially contributing to the pathology of this syndrome.
A genome-wide association study (GWAS) of daughter pregnancy rate (DPR) was conducted using 75,133 SNPs and 40,203 first lactation crossbred dairy cows mostly from Jersey–Holstein crosses. The GWAS analysis detected 6528 additive effects, 65 dominance effects, 1638 additive × additive (A × A) effects, 3 additive × dominance effects, and 18 intra-chromosome dominance × dominance (D × D) effects. Of the 1638 A × A effects, 1634 were intra-chromosome and four were inter-chromosome A × A effects. The distance between two SNPs with intra-chromosome epistasis effects was in the range of 3.61 Kb to 2.68 Mb, and many interacting SNP pairs were within the same genes. The additive and A × A effects were distributed on all chromosomes showing genome-wide involvement in DPR heterosis. The dominance and D × D effects all had homozygous advantages and heterozygous disadvantages. The GWAS results identified four genetic mechanisms underlying DPR heterosis in crossbred dairy cows: complementary additive effects from different breeds and new additive effects due to cross breeding, two-locus allelic interactions between loci and between breeds, within-locus allelic interactions between breeds, and genotype × genotype interactions enabled by allelic interactions between breeds. Results in this study provided a novel understanding about the genetic factors and mechanisms underlying DPR heterosis in crossbred dairy cows.
Modeling multiple sampling densities within a hierarchical framework enables borrowing of information across samples. These density random effects can act as kernels in latent variable models to represent exchangeable subgroups or clusters. A key feature of these kernels is the (functional) covariance they induce, which determines how densities are grouped in mixture models. Our motivating problem is clustering chromatin accessibility profiles from high-throughput DNase-seq experiments to detect transcription factor (TF) binding. TF binding typically produces footprint profiles with spatial patterns, creating long-range dependency across genomic locations. Existing nonparametric hierarchical models impose restrictive covariance assumptions and cannot accommodate such dependencies, often leading to biologically uninformative clusters. We propose a nonparametric density kernel flexible enough to capture diverse covariance structures and adaptive to various spatial patterns of TF footprints. The kernel specifies dyadic tree splitting probabilities via a multivariate logit-normal model with a sparse precision matrix. Bayesian inference for latent variable models using this kernel is implemented through Gibbs sampling with Polya-Gamma augmentation. Extensive simulations show that our kernel substantially improves clustering accuracy. We apply the proposed mixture model to DNase-seq data from the ENCODE project, which results in biologically meaningful clusters corresponding to binding events of two common TFs.
We propose a functional evaluation metric for generative models based on the relative density ratio (RDR) designed to characterize distributional differences between real and generated samples. We show that the RDR as a functional summary of the goodness-of-fit for the generative model, possesses several desirable theoretical properties. It preserves ϕ-divergence between two distributions, enables sample-level evaluation that facilitates downstream investigations of feature-specific distributional differences, and has a bounded range that affords clear interpretability and numerical stability. Functional estimation of the RDR is achieved efficiently through convex optimization on the variational form of ϕ-divergence. We provide theoretical convergence rate guarantees for general estimators based on M-estimator theory, as well as the convergence rates of neural network-based estimators when the true ratio is in the anisotropic Besov space. We demonstrate the power of the proposed RDR-based evaluation through numerical experiments on MNIST, CelebA64, and the American Gut project microbiome data. We show that the estimated RDR not only allows for an effective comparison of the overall performance of competing generative models, but it can also offer a convenient means of revealing the nature of the underlying goodness-of-fit. This enables one to assess support overlap, coverage, and fidelity while pinpointing regions of the sample space where generators concentrate and revealing the features that drive the most salient distributional differences.
A unique line of Holstein cattle has been maintained without selection in Minnesota since 1964. After many generations, unselected cattle produce less milk, but have better reproductive performance and health traits when compared with contemporary cows. Comparisons between this line of unselected Holstein and those under selection provide useful insights that connect selection and complex traits in cattle. Utilizing these unique resources and sequence data, we sought to identify genome changes due to selection. We sequenced 30 unselected and 54 selected Holstein cattle and compared their sequence variants to identify selection signatures. After many years, the two populations showed completely different patterns in their genome-level population structures and linkage disequilibrium. By integrating signals from five different detection methods, we detected consensus selection signatures from at least four methods covering 14,533 SNPs and 155 protein-coding genes. An integrated analysis of selection signatures with gene annotation, pathways, and the cattle QTL database demonstrated that the genomic regions under selection are related to milk productivity, health, and reproductive efficiency. The polygenic nature of these complex traits is evident from hundreds of selection signatures and candidate genes, suggesting that long-term artificial selection has acted on the whole genome rather than a few major genes. In summary, our study identified candidate selection signatures underlying phenotypic differences between unselected and selected Holstein cows and revealed insights into the genetic basis of complex traits in cattle.
Identifying causal genetic variants underlying economically important traits in dairy cattle is essential for understanding their genetic basis and optimizing breeding programs. The growing availability of sequenced reference genomes and individuals with both phenotypic and genotypic data notably enhances our ability to detect genetic associations and further pinpoint causal effects. This comprehensive GWAS of dairy cattle used deregressed breeding values as phenotypes and analyzed 11,292,243 quality-controlled, imputed sequence variants from 50,309 Holstein bulls. The number of bulls with available phenotypes ranged from 23,121 to 50,309 across 30 complex traits categorized into production and yield, type, and longevity and health. We performed GWAS using our SLEMM-GWA approach, which accounts for the varying reliability of deregressed breeding values across individuals and demonstrates computational efficiency for large sample sizes and sequence data. This analysis identified 381 significant association peaks, of which 126 are novel findings. Subsequent Bayesian fine-mapping provided statistical prioritization by assigning posterior conditional inclusion probabilities to individual variants and genes, yielding a list of credible candidate genes-an advancement over conventional GWAS reporting of all proximal genes. This prioritization offered direct statistical support for previously reported genes and, more importantly, identified credible candidate genes within the 126 newly discovered peaks for specific traits, including AOPEP, GC, E2F6, MGST1, VPS13B, ZNF652, ASPH, SFMBT1, and MAPRE2. These findings enhance the understanding of the genetic architecture of these complex dairy traits and provide valuable insights for the refinement of genomic selection strategies and breeding programs in Holstein cattle.