Abstract The pecan weevil, Curculio caryae (Horn), is an obligate feeder of pecan and native hickory trees (genus Carya) throughout North America. Subsequently it is a significant agricultural pest in pecan orchards. In this study, we present a reference quality genome using deep-coverage, ~40x PacBio HiFi genome sequence reads, and chromatin confirmation, Hi-C, scaffolding. The final genome assembly is approximately 2.2 Gb, which was confirmed by flow cytometry. The primary genome scaffolds have an N50 of 132 Mb and a BUSCO completeness of 95.4% [S:94.3%, D:1.1%]. Furthermore, we employed PacBio long-read RNA, Iso-seq, for de novo gene annotation, in conjunction with InterProscan to identify approximately 19,000 protein coding genes. Repeat content is extensive, contributing at least >80% of the total genome. This data set provides a valuable resource for comparative genomics and evolutionary studies of an economically impactful group of insect pests that currently lack extensive genomic resources.
Annotation of regulatory elements is essential for understanding mechanisms underlying gene regulation, particularly tissue-specific regulation in human and animals. Here, we characterize 274,682 enhancers and 25,975 promoters across 24 tissues from an adult female sheep using ChIP-seq, ATAC-seq, CAGE-seq, RRBS, WGBS, and RNA-seq. We identify seven neural development-related genes with over 10 enhancers in brain tissues, highlighting the role of tissue-specific regulation. Cis-regulatory enhancer-promoter combinations provide insights into tissue-specific enhancers, such as the cerebellum-specific enhancer (chr15: 57390520-57390685) regulating BDNF, which is expressed in both the cerebellum and cerebral cortex. Comparative analysis of enhancer-promoter combinations in human, mouse, pig, cattle, and sheep reveals ruminant-specific pathways, including pentose catabolism and long-chain fatty acid import regulation. A milk fat yield quantitative trait locus (QTL) identified within an enhancer interacts with the fat metabolism-related gene COMMD1, and a birth weight-associated QTL detected within a cerebellum-specific enhancer regulates XKR4. This study provides a robust framework for exploring cis-regulatory mechanisms and tissue-specific regulation, advancing the functional annotation of the sheep reference genome.
Pangenomes of several species have been assembled recently, facilitating the detection and genotyping of structural variants. As part of the FarmGTEx Project, we previously constructed a Holstein pangenome (H20D) based on 40 phased haploid assemblies. Here, we use this breed specific pangenome to genotype 93,059 structural variants from whole-genome sequences of 1,571 cattle. We then develop a Holstein pangenome variation imputation reference panel we name HolPIP. Leveraging HolPIP, we impute 86.65% (68,354/78,886) of structural variants for 50,299 bulls with Beagle R² ≥ 0.8. Using these imputed structural variants and phenotypes for 43 complex traits, we conduct GWAS, identifying 1,225 structural variant-trait associations. We next use fine-mapping to prioritize 32 high-confidence candidate structural variants, including a 75-bp deletion in ANKRD11 linked to dairy form, rump width, and stature, as well as an insertion in DHX32 associated with RNA metabolism. Compared to SNPs across various functional annotations, structural variants show a stronger genome-wide enrichment across most complex traits in cattle, suggesting that structural variants may have an important contribution to the genetic basis of dairy traits.
Structural variants are an underexplored source of genetic diversity. As part of the FarmGTEx Project, here we report a Holstein breed-specific pangenome graph (H20D) using Minigraph-Cactus and 40 phased haploid assemblies from 20 cows. H20D outperforms both assembly- and read-based long-read callers, and far exceeds short-read approaches, identifying over 10,000 additional structural variants per sample. It also significantly improves structural variant detection and genotyping relative to graphs built across breeds or from fewer/unphased assemblies, with particular advantages in complex regions. Using H20D, we genotype variants in 173 cattle and performed a GWAS, where a larger fraction of structural variants than SNPs reach genome-wide significance, implicating them as potential causal variants. Together, these results demonstrate the power of phased, within-breed pangenome graphs for accurate SV genotyping and trait mapping in dairy cattle.
The domestication of the banteng in Southeast Asia is one of the world’s least known livestock domestications yet a vital component of the agricultural system in Indonesia and surrounding countries. Here, we generated the first reference genome of the banteng and used it to analyze a set of 78 resequenced wild and domesticated bantengs, including 19 newly generated whole-genome-sequenced samples, of which three are historical samples. We found low heterozygosity and significant differentiation, the latter primarily driven by recent genetic drift and inbreeding in two populations and clearly attributable to anthropogenically driven founder events or ex situ breeding. Population structure, when excluding these two populations, was limited, and we found that the evolutionary divergence between wild and domestic banteng was moderate (FST = 0.14), relatively young (10,356 years), and accompanied by post-divergence gene flow. We found only weak signals of a domestication bottleneck between ∼6,100 and 2,900 years ago, and genetic diversity was, on average, higher in domestic than in wild banteng. Despite the soft domestication history, we found 36 candidate genes potentially under selection during domestication, with the leptin receptor gene (LEPR) of particular interest due to the robust selection signal across methods and its known association with metabolism, obesity, and energy homeostasis. Finally, genetic load estimation revealed that Bali cattle in Australia have high realized load, whereas Bali cattle from Bali have high masked load. These findings provide the first genomic insights into an understudied bovine that is critically endangered in its wild form and agriculturally important in its domesticated form.
High-quality reference genomes are critical for studying the biology of the genome, but current methods often leave gaps and errors, especially in repetitive regions. These issues arise from challenges in genome graph curation and are not fully resolvable by standard polishing approaches. To address this, we developed Verkko-Fillet, a Python-based interactive framework for inspecting, editing, and refining genome assembly graphs. It integrates multiple data sources and provides tools for visualization, gap filling, and structural correction. Applied to a giraffe and the benchmark human genome, Verkko-Fillet improves a draft assembly (Q61.5) to a complete telomere-to-telomere genome (Q73.6), increasing both contiguity and accuracy. This work highlights the importance of graph-based curation for producing a finished, gapless genome assembly suitable for downstream analyses.
The cattle genome is crucial for understanding ruminant biology, but it remains incomplete. Here we present a telomere-to-telomere haplotype-resolved X chromosome and four autosomes of cattle in a near-complete assembly that is 431 Mb (16%) longer than the current reference genome. Using this assembly (UOA_Wagyu_1) we identify 738 new protein-coding genes and support the characterization of centromeric repeats, identification of transposable elements, and enabled the detection of 2397 more structural variants from 20 Wagyu animals than using ARS-UCD2.0. We find that the cattle X centromere is a natural neocentromere with highly identical inverted repeats, no bovine satellite repeats, low CENP-A signal, low methylation, and low CpG content, in contrast to the autosomal centromeres that are comprised of typical bovine satellite repeats and epigenetic features. Our results suggest it likely formed from transposable element expansion and CpG deamination, suggesting dynamic evolution. We find eighteen X-pseudoautosomal region genes have conserved testes expression between cattle and apes. We also find all cattle X neocentromere protein-coding genes are expressed in testes, which suggests they potentially play a role in reproduction.
Deciphering the regulatory syntax of the genome is essential to understand the genetic and molecular architecture of complex traits, as most trait-associated variants lie in non-coding regions. Yet, functional annotation of the bovine genome remains limited, hindering our ability to unravel the mechanisms underpinning complex traits of economic and ecological importance in cattle. Here, we present a comprehensive epigenetic atlas comprising 1,138 genome-wide epigenetic profiles, including chromatin accessibility, six histone modifications, CCCTC-binding factor (CTCF) transcription factor binding, DNA methylation, chromatin conformation, and transcriptomes across 53 adult tissues, five fetal tissues, and seven primary cell types. This atlas-level data enables us to annotate around 45% of the genome as putative regulatory elements exhibiting tissue- or cell-specific regulatory activity. Leveraging sequence-to-function deep learning models, we discovered 301 sequence motifs and predicted the functional impact of genetic variants through in silico mutagenesis, thereby facilitating the decoding of the regulatory syntax of the cattle genome and fine-mapping of GWAS loci for 22 complex traits. Cross-species analysis further revealed evolutionarily conserved features of regulatory architecture and provided evolutionary insights into complex traits and diseases in humans. Together, this atlas offers a foundational resource for advancing cattle functional genomics, sustainable breeding, and studies of regulatory evolution.
Mannheimia haemolytica is a key agent in bovine respiratory disease (BRD), driving antibiotic use in feedyard cattle. As a facultative anaerobe commonly found in the upper respiratory tracts of cattle, its role under anaerobic conditions in BRD has not been extensively studied despite the known importance of anaerobiosis in human respiratory infections. Utilizing a combined omics approach, we refuted the null hypothesis that the M. haemolytica genome sequence is independent of the environment. This finding provides the rationale for research exploring how anaerobiosis-driven genome plasticity might contribute to a transition from a commensal to a virulent state. Genome sequencing of colony morphology variants from aerobic and anaerobic cultures revealed phase variation via homologous recombination between ribosomal RNA (rRNA) operons and slipped strand mispairing at simple sequence repeats (SSRs), frameshifting genes "on" and "off." Homologous recombination was exclusive to anaerobic conditions. SSR length variation in a DNA methyltransferase gene, targeting 5'-GACAT, correlated with methylation status. We observed statistically significant differences in the fraction of methylated motifs in isolates derived from anaerobic and aerobic cultures, as well as in the transcript abundances of genes in a nitrate reduction pathway and in a ribosomal protein class among anaerobic culture-derived colonial variants. We propose that instead of representing a M. haemolytica isolate's genome as a single static sequence, dynamic genome models should be developed. These models should account for the stochastic changes in genome sequence induced by phase variation, represented as a probability-weighted mixture of homologous recombinants and SSR switches.IMPORTANCEThe anaerobic growth of Mannheimia haemolytica generates phase variants through homologous recombination and slipped strand mispairing at simple sequence repeats. Homologous recombination between ribosomal RNA operons occurred exclusively in anaerobic conditions, likely similar to those present in infected lung tissue. Phase variation may yield variants better adapted for persistence in pulmonary tissues, potentially impacting antimicrobial persistence, as observed in persisters of other species. The genomic diversity resulting from phase variation complicates analyses based on static genome sequences, highlighting the need for dynamic genome models that reflect the stochastic nature of phase variation to understand M. haemolytica's role in host disease biology.
The bighorn sheep (Ovis canadensis), despite its close relation to domestic sheep, suffer higher morbidity and mortality from respiratory disease complexes, likely due to genetic differences in immune responses. Unraveling highly repetitive regions such as immune loci and genetic differences was problematic until now. We generated a bighorn sheep telomere-to-telomere assembly, adding 14.28% of novel sequence compared to the previous reference. This enabled the first complete immune loci annotation revealing the IGL and TR loci are significantly short in bighorn sheep. Importantly, a critical immune gene GBP5 and ZNF501, involved in Golgi-mediated immune response, are lacking in bighorn but present in domestic sheep. Re-analysis of a Mycoplasma ovipneumoniae carriage study, using this assembly, identified the immune gene CAPN2 as a key genetic marker for disease carriage, not observable in the original study. This work provides a critical resource for identifying phenotype-linked genetic variation and exploring evolutionary adaptations of bighorn sheep.
Genetic diversity is a crucial resource in livestock, determining their traits and ability to respond to selection. Indonesian cattle are unique due to their history of admixture involving both zebu (Bos indicus) and banteng (B. javanicus), and may therefore contain novel cattle genetic resources. We generated whole genome sequences from 126 Indonesian cattle, 51 domesticated banteng and three captive banteng. We show that Indonesian cattle have very high genetic diversity, especially the Madura breed due to introgression from banteng and possibly other Bos species, contributing up to 36.6% of the Madura's genome. We find that Indonesian zebu ancestry can be traced to at least three distinct ancestral populations, two of which were introduced more than 1345 years ago from mainland Southeast or eastern Asia. Peaks and valleys in banteng ancestry across the genome in admixed breeds suggest that both negative and positive selection act on introgressed haplotypes. Despite adaptive introgression being mainly breed-specific, we found evidence that some phenotypes, such as coat color, have experienced convergent adaptive introgression. Overall, our results provide insights into the historical movement of cattle in Asia, and showcase the potential for genetic improvement of cattle by identifying ~3.5 million novel SNPs introgressed into Indonesian cattle.
Ruminant species are vital for agriculture, ecosystems, and conservation and remain vulnerable to infectious and zoonotic diseases. Advances in genome sequencing and genomics now enable high-resolution analysis of immunoglobulin (IG) loci and antibody repertoires uncovering extensive germline diversity, structural variation, and lineage-specific adaptations, such as ultralong cysteine-rich Abs in cattle. This review summarizes current knowledge of ruminant IG locus organization and repertoire generation and discusses the evolutionary origins of ultralong Abs. It also examines the challenges highly repetitive IG loci pose for assembly, annotation, and nomenclature and highlights emerging solutions. Finally, it describes genomic approaches for linking immune genotypes to phenotypes that create promise for improving ruminant health.
Introduction Most SV studies in livestock rely on short-read sequencing, posing challenges in accurately characterizing large genomic variants due to their limited read length. Objectives Our goal is to reveal structural variation and novel sequences specific to Holstein and Jersey cattle breeds using long-read and pan-genome analyses. Methods We sequenced 20 Holsteins and 8 Jersey cattle using PacBio HiFi to 20×, and integrated five read-based and one assembly-based SV caller to determine SVs. Results We assembled the 28 genomes averaging 3.25 Gb with a contig N50 of 69.36 Mb and using the ARS-UCD1.2 reference, we acquired Holstein/Jersey SV catalogs with 74,068/54,689 events spanning 202/135 Mb (7.43 %/4.97 % of the genome). SVs were enriched in less conserved, non-coding, and non-regulatory regions. Comparing Holsteins with differing feed efficiency (FE), SVs unique to high FE were linked to energy metabolism and olfactory receptors, while those specific to low FE were associated with material transport. We constructed Holstein/Jersey pangenome graphs with 148,598/105,875 nodes and 208,891/147,990 edges, representing 47,028/37,137 biallelic and multi-allelic events, and 63.75/42.34 Mb of novel sequence. We observed SV count saturation with 20 Holsteins, while adding Jerseys significantly increased the SV count, highlighting breed-specific SV events. Conclusion Our long-read data and SV catalogs are valuable resources, revealing that the cattle genome is more complex than previously thought.
Lack of freezing tolerance is a major constraint for the production of agronomically important Brassica species, particularly in the Northern Great Plains (NGP) of the United States and Canada. However, within the Brassicaceae family, winter germplasm of camelina have shown excellent freezing tolerance and overwinter potential in the NGPs. Differences in freezing tolerance between a winter variety (Joelle) and a spring variety (C046) of camelina appear to be controlled by a small number of dominant or co-dominant genes. To unravel the genetic mechanisms for the differences in freezing tolerance, 254 Recombinant Inbred Lines (RILs) were developed using reciprocal crosses between these two camelina varieties. The RIL population was phenotyped at the F7 stage for freezing tolerance under controlled conditions and genotyped by whole-genome skim sequencing. A one-way ANOVA test revealed a significant (P < 0.001) difference exists among the RILs for freezing tolerance. A significant and strong correlation (r = 0.60, P < 0.001) was also observed between freezing tolerance and flowering time, indicating that regulation of flowering time might also influence freezing tolerance in camelina. A de novo linkage map was constructed using 4507 SNP markers covering a total of 1208.5 cM map distance with an average of 0.3 cM distances between the markers, which formed 20 linkage groups representing the 20 chromosomes (Chr) of C. sativa. The QTL analyses using three different programs revealed significant loci on Chr 8, 11, 13, 16 and 18 with LOD threshold value of over 3.5 for freezing tolerance. The QTL peaks with the greatest LOD values of 20.7 and 26.8 were observed at Chr 8 and Chr 13 and accounted for 18.3 % and 25.3 % of the phenotypic variation respectively. A total of 3369 annotated camelina genes were identified within +/- 50 Kb from the consensus QTL intervals generated from the output of the three mapping programs. Among them, 125 were transcription factors including twelve MIKC_MADS on Chr 8, 11, 13, 16 and 18 and two that annotate as the floral regulators FLOWERING LOCUS C (FLC) on Chr 8 and 13, an orthologue of MADS AFFECTING FLOWERING 3 and 4 (MAF4 and MAF3) of arabidopsis on Chr18, and an orthologue of SHORT VEGETATIVE PHASE (SVP) on Chr 16. Although many of the candidate genes identified near the freezing tolerance QTLs have previously been associated with flowering time, further studies are needed to help unravel how these genes impact freezing tolerance mechanisms and improve freezing tolerance in camelina and other Brassica crop species.
Reference genomes of cattle and sheep have lacked contiguous assemblies of the sex-determining Y chromosome. Here, we assemble complete and gapless telomere to telomere (T2T) Y chromosomes for these species. We find that the pseudo-autosomal regions are similar in length, but the total chromosome size is substantially different, with the cattle Y more than twice the length of the sheep Y. The length disparity is accounted for by expanded ampliconic region in cattle. The genic amplification in cattle contrasts with pseudogenization in sheep suggesting opposite evolutionary mechanisms since their divergence 19MYA. The centromeres also differ dramatically despite the close relationship between these species at the overall genome sequence level. These Y chromosomes have been added to the current reference assemblies in GenBank opening new opportunities for the study of evolution and variation while supporting efforts to improve sustainability in these important livestock species that generally use sire-driven genetic improvement strategies.
Camelina sativa is an emerging oilseed crop with potential for multi-cropping with more traditional cash crops. Until very recently, the only reference genome for camelina was produced using short read technologies and was considerably fragmented. To facilitate the mapping of traits associated with vernalization requirement and freezing tolerance in an F7 recombinant inbred line (RIL) population developed by crossing a spring- and winter-biotype of camelina, we performed long read sequencing using PacBio HiFi technology. Here we report on the genome assembly and gene annotation obtained from the long read sequencing of the two parental lines. Both assemblies formed 20 chromosomal units with a genome size of 667,897,934 and 714,129,675 bases for C046 and Joelle, respectively. Assessment of completeness by BUSCO analysis indicated the C046 and Joelle genomes were both 99.5 complete with 98.1 % of the conserved genes being duplicated. Skim-sequencing of the F6:7 RILs resulted in 823,142 markers that mapped to variants identified by comparisons of the parental genome sequences. These sequences provide a valuable resource for breeders seeking to improve the food and industrial attributes of camelina.