Genotype-to-phenotype (G2P) research in farmed animals has entered a phase in which deep genome annotation, multi-omics and AI-enabled prediction can nominate variants, regulatory mechanisms and cellular pathways at scale, but causal validation remains a major bottleneck. In vitro cellular systems, from tractable primary cells to organoids and other advanced models, provide experimentally controlled contexts in which massively parallel reporter assays and CRISPR-based perturbations can test variant effects and define cellular phenotypes. Here, we examine how in vitro approaches can be deployed across host–pathogen interactions, genotype-by-environment responses, nutrition, adaptation and One Health research, and outline priorities for integrating in vivo, in vitro and in silico data into mechanistic G2P workflows.
Oral swabs of dairy cows have been suggested as a proxy for direct ruminal sampling, and this approach can identify the presence of up to 70% of the rumen microbial community. Here, we further extend the utility of this approach by correlating the bacterial community of swabs collected from 226 dairy cows on a research farm in Wisconsin, USA, with average milk yield and days in milk, two phenotypes previously associated with differences in the ruminal microbiome. We then obtained milk production efficiency data for a subset of these animals (gross feed efficiency [GFE] and residual feed intake [RFI]) and correlated these metrics against their associated microbial data. We found that when using the oral swabs, we could identify correlations between bacterial genera and days in milk (P < 0.05). We further show that the ruminal microbiota was associated with average milk yield and days in milk for animals in their first lactation. Differential abundance testing identified amplicon sequence variants (ASVs) associated with these metrics (P < 0.05). Our comparison of bacterial communities between high and low efficiency groups, as determined by GFE and RFI, identified a significant difference in Shannon's diversity in second lactation cows (P < 0.05). We also found that RFI was significantly correlated with the bacterial community in second lactation animals (P < 0.05). Differential abundance analysis identified multiple oral- and rumen-associated ASVs correlated with GFE and RFI (P < 0.05). This study further establishes the utility of oral swabs as a ruminal proxy.IMPORTANCEImproving milk production efficiency is a key goal in the dairy industry and is traditionally pursued through genetic selection, diet optimization, and herd management practices. The ruminal microbiome, essential for digesting feed, has been linked to milk production efficiency, suggesting that microbiome modulation could improve efficiency. However, the integration of rumen microbiology into current management practices is hampered by the difficulty of large-scale rumen sampling, as proxies like fecal samples do not accurately reflect the ruminal microbiota. Traditional methods, like cannulation and stomach tubing, are labor-intensive and impractical for extensive sampling. Our research demonstrates the potential use of oral swabs as a scalable, effective method for characterizing the microbiome and its associations with milk production metrics, recapitulating established associations obtained through traditional ruminal sampling methods.
Using oral swabs to collect the remnants of stomach content regurgitation during rumination in dairy cows can replicate up to 70% of the ruminal bacterial community, offering potential for broad-scale population-based studies on the rumen microbiome. The swabs collected from dairy cows often vary widely with respect to sample quality, likely due to several factors such as time of sample collection and cow rumination behavior, which may limit the ability of a given swab to accurately represent the ruminal microbiome. One such factor is the color of the swab, which can vary significantly across different cows. Here, we hypothesize that darker-colored swabs contain more rumen contents, thereby better representing the ruminal bacterial community than lighter-colored swabs. To address this, we collected oral swabs from 402 dairy cows and rumen samples from 13 cannulated cows on a research farm in Wisconsin, United States and subjected them to 16S rRNA sequencing. In addition, given that little is known about the ability of oral swabs to recapitulate the ruminal fungal community, we also conducted ITS sequencing of these samples. To correlate swab color to the microbiota we developed and utilized a novel imaging approach to colorimetrically quantify each swab from a range of light to dark. We found that swabs with increasing darkness scores were significantly associated with increased bacterial alpha diversity (p < 0.05). Lighter swabs exhibited greater variation in their community structure, with many identified amplicon sequence variants (ASVs) categorized as belonging to known bovine oral and environmental taxa. Our analysis of the fungal microbiome found that swabs with increasing darkness scores were associated with decreased alpha diversity (p < 0.05) and were also significantly associated with the ruminal solids fungal community, but not with the ruminal liquid community. Our study refines the utility of oral swabs as a useful proxy for capturing the ruminal microbiome and demonstrates that swab color is an important factor to consider when using this approach for documenting both the bacterial and fungal communities.
Vicia villosa is an incompletely domesticated annual legume of the Fabaceae family native to Europe and Western Asia. V. villosa is widely used as a cover crop and forage due to its ability to withstand harsh winters. Here, we generated a reference-quality genome assembly (Vvill1.0) from low error-rate long-sequence reads to improve the genetic-based trait selection of this species. Our Vvill1.0 assembly includes seven scaffolds corresponding to the seven estimated linkage groups and comprising approximately 68% of the total genome size of 2.03 Gbp. This assembly is expected to be a useful resource for genetically improving this emerging cover crop species and provide useful insights into legume genomics and plant genome evolution.
The Bovine Pangenome Consortium (BPC) is an international collaboration dedicated to the assembly of cattle genomes to develop a more complete representation of cattle genomic diversity. The goal of the BPC is to provide genome assemblies and a community-agreed pangenome representation to replace breed-specific reference assemblies for cattle genomics. The BPC invites partners sharing our vision to participate in the production of these assemblies and the development of a common, community-approved, pangenome reference as a public resource for the research community ( https://bovinepangenome.github.io/ ). This community-driven resource will provide the context for comparison between studies and the future foundation for cattle genomic selection.
Background The gaur (Bos gaurus) is the largest extant wild bovine species, native to South and Southeast Asia, with unique traits, and is listed as vulnerable by the International Union for Conservation of Nature (IUCN). Results We report the first gaur reference genome and identify three biological pathways including lysozyme activity, proton transmembrane transporter activity, and oxygen transport with significant changes in gene copy number in gaur compared to other mammals. These may reflect adaptation to challenges related to climate and nutrition. Comparative analyses with domesticated indicine (Bos indicus) and taurine (Bos taurus) cattle revealed genomic signatures of artificial selection, including the expansion of sperm odorant receptor genes in domesticated cattle, which may have important implications for understanding selection for male fertility. Conclusions Apart from aiding dissection of economically important traits, the gaur genome will also provide the foundation to conserve the species.
Advantages of pangenomes over linear reference assemblies for genome research have recently been established. However, potential effects of sequence platform and assembly approach, or of combining assemblies created by different approaches, on pangenome construction have not been investigated. Here we generate haplotype-resolved assemblies from the offspring of three bovine trios representing increasing levels of heterozygosity that each demonstrate a substantial improvement in contiguity, completeness, and accuracy over the current Bos taurus reference genome. Diploid coverage as low as 20x for HiFi or 60x for ONT is sufficient to produce two haplotype-resolved assemblies meeting standards set by the Vertebrate Genomes Project. Structural variant-based pangenomes created from the haplotype-resolved assemblies demonstrate significant consensus regardless of sequence platform, assembler algorithm, or coverage. Inspecting pangenome topologies identifies 90 thousand structural variants including 931 overlapping with coding sequences; this approach reveals variants affecting QRICH2 , PRDM9 , HSPA1A , TAS2R46 , and GC that have potential to affect phenotype.
Recessive alleles represent genetic risk in populations that have undergone bottleneck events. We present a comprehensive framework for identification and validation of these genetic defects, including haplotype-based detection, variant selection from sequence data, and validation using knockout embryos. Holstein haplotype 2 (HH2), which causes embryonic death, was used to demonstrate the approach. Holstein haplotype 2 was identified using a deficiency-of-homozygotes approach and confirmed to negatively affect conception rate and stillbirths. Five carriers were present in a group of 183 sequenced Holstein bulls selected to maximize the coverage of unique haplotypes. Three variants concordant with haplotype calls were found in HH2: a high-priority frameshift mutation resulting, and 2 low-priority variants (1 synonymous variant, 1 premature stop codon). The frameshift in intraflagellar 80 (IFT80) was confirmed in a separate group of Holsteins from the 1000 Bull Genomes Project that shared no animals with the discovery set. IFT80-null embryos were generated by truncating the IFT80 transcript at exon 2 or 11 using a CRISPR-Cas9 system. Abattoir-derived oocytes were fertilized in vitro, and zygotes were injected at the one-cell stage either with a guide RNA and CAS9 mRNA complex (n = 100) or Cas9 mRNA (control, n = 100) before return to culture, and replicated 3 times. IFT80 is activated at the 8-cell stage, and IFT80-null embryos arrested at this stage of development, which is consistent with data from mouse hypomorphs and HH2 carrier-to-carrier matings. This frameshift in IFT80 on chromosome 1 at 107,172,615 bp (p.Leu381fs) disrupts WNT and hedgehog signaling, and is responsible for the death of homozygous embryos.
A cattle pangenome representation was created based on the genome sequences of 898 cattle representing 57 breeds. The pangenome identified 83 Mb of sequence not found in the cattle reference genome, representing 3.1% novel sequence compared with the 2.71-Gb reference. A catalog of structural variants developed from this cattle population identified 3.3 million deletions, 0.12 million inversions, and 0.18 million duplications. Estimates of breed ancestry and hybridization between cattle breeds using insertion/deletions as markers were similar to those produced by single nucleotide polymorphism–based analysis. Hundreds of deletions were observed to have stratification based on subspecies and breed. For example, an insertion of a Bov-tA1 repeat element was identified in the first intron of the APPL2 gene and correlated with cattle breed geographic distribution. This insertion falls within a segment overlapping predicted enhancer and promoter regions of the gene, and could affect important traits such as immune response, olfactory functions, cell proliferation, and glucose metabolism in muscle. The results indicate that pangenomes are a valuable resource for studying diversity and evolutionary history, and help to delineate how domestication, trait-based breeding, and adaptive introgression have shaped the cattle genome.
Relative to other crops, red clover (Trifolium pratense L.) has various favorable traits making it an ideal forage crop. Conventional breeding has improved varieties, but modern genomic methods could accelerate progress and facilitate gene discovery. Existing short-read-based genome assemblies of the ∼420 megabase pair (Mbp) genome are fragmented into >135,000 contigs, with numerous order and orientation errors within scaffolds, probably associated with the plant’s biology, which displays gametophytic self-incompatibility resulting in inherent high heterozygosity. Here, we present a high-quality long-read-based assembly of red clover with a more than 500-fold reduction in contigs, improved per-base quality, and increased contig N50 by three orders of magnitude. The 413.5 Mbp assembly is nearly 20% longer than the 350 Mbp short-read assembly, closer to the predicted genome size. We also present quality measures and full-length isoform RNA transcript sequences for assessing accuracy and future genome annotation. The assembly accurately represents the seven main linkage groups in an allogamous (outcrossing), highly heterozygous plant genome.
Background The domestic sheep (Ovis aries) is an important agricultural species raised for meat, wool, and milk across the world. A high-quality reference genome for this species enhances the ability to discover genetic mechanisms influencing biological traits. Furthermore, a high-quality reference genome allows for precise functional annotation of gene regulatory elements. The rapid advances in genome assembly algorithms and emergence of sequencing technologies with increasingly long reads provide the opportunity for an improved de novo assembly of the sheep reference genome. Findings Short-read Illumina (55x coverage), long-read Pacific Biosciences (75x coverage), and Hi-C data from this ewe retrieved from public databases were combined with an additional 50x coverage of Oxford Nanopore data and assembled with canu v1.9. The assembled contigs were scaffolded using Hi-C data with Salsa v2.2, gaps filled with PBsuitev15.8.24, and polished with Nanopolish v0.12.5. After duplicate contig removal with PurgeDups v1.0.1, chromosomes were oriented and polished with 2 rounds of a pipeline that consisted of freebayes v1.3.1 to call variants, Merfin to validate them, and BCFtools to generate the consensus fasta. The ARS-UI_Ramb_v2.0 assembly is 2.63 Gb in length and has improved continuity (contig NG50 of 43.18 Mb), with a 19- and 38-fold decrease in the number of scaffolds compared with Oar_rambouillet_v1.0 and Oar_v4.0. ARS-UI_Ramb_v2.0 has greater per-base accuracy and fewer insertions and deletions identified from mapped RNA sequence than previous assemblies. Conclusions The ARS-UI_Ramb_v2.0 assembly is a substantial improvement in contiguity that will optimize the functional annotation of the sheep genome and facilitate improved mapping accuracy of genetic variant and expression data for traits in sheep.
Workshop cluster 1 (WC1) molecules are part of the scavenger receptor cysteine-rich (SRCR) superfamily and act as hybrid co-receptors for the γδ T cell receptor and as pattern recognition receptors for binding pathogens. These members of the CD163 gene family are expressed on γδ T cells in the blood of ruminants. While the presence of WC1+ γδ T cells in the blood of goats has been demonstrated using monoclonal antibodies, there was no information available about the goat WC1 gene family. The caprine WC1 multigenic array was characterized here for number, structure and expression of genes, and similarity to WC1 genes of cattle and among goat breeds. We found sequence for 17 complete WC1 genes and evidence for up to 30 SRCR a1 or d1 domains which represent distinct signature domains for individual genes. This suggests substantially more WC1 genes than in cattle. Moreover, goats had seven different WC1 gene structures of which 4 are unique to goats. Caprine WC1 genes also had multiple transcript splice variants of their intracytoplasmic domains that eliminated tyrosines shown previously to be important for signal transduction. The most distal WC1 SRCR a1 domains were highly conserved among goat breeds, but fewer were conserved between goats and cattle. Since goats have a greater number of WC1 genes and unique WC1 gene structures relative to cattle, goat WC1 molecules may have expanded functions. This finding may impact research on next-generation vaccines designed to stimulate γδ T cells.
Red clover (Trifolium pratense L.) is an important forage crop and serves as a major contributor of nitrogen input in pasture settings because of its ability to fix atmospheric nitrogen. During the legume-rhizobial symbiosis, the host plant undergoes a large number of gene expression changes, leading to development of root nodules that house the rhizobium bacteria as they are converted into nitrogen-fixing bacteroids. Many of the genes involved in symbiosis are conserved across legume species, while others are species-specific with little or no homology across species and likely regulate the specific plant genotype/symbiont strain interactions. Red clover has not been widely used for studying symbiotic nitrogen fixation, primarily due to its outcrossing nature, making genetic analysis rather complicated. With the addition of recent annotated genomic resources and use of RNA-seq tools, we annotated and characterized a number of genes that are expressed only in nodule forming roots. These genes include those encoding nodule-specific cysteine rich peptides (NCRs) and nodule-specific Polycystin-1, Lipoxygenase, Alpha toxic (PLAT) domain proteins (NPDs). Our results show that red clover encodes one of the highest number of NCRs and ATS3-like/NPDs, which are postulated to increase nitrogen fixation efficiency, in the Inverted-Repeat Lacking Clade (IRLC) of legumes. Knowledge of the variation and expression of these genes in red clover will provide more insights into the function of these genes in regulating legume-rhizobial symbiosis and aid in breeding of red clover genotypes with increased nitrogen fixation efficiency.
Microbial communities might include distinct lineages of closely related organisms that complicate metagenomic assembly and prevent the generation of complete metagenome-assembled genomes (MAGs). Here we show that deep sequencing using long (HiFi) reads combined with Hi-C binning can address this challenge even for complex microbial communities. Using existing methods, we sequenced the sheep fecal metagenome and identified 428 MAGs with more than 90% completeness, including 44 MAGs in single circular contigs. To resolve closely related strains (lineages), we developed MAGPhase, which separates lineages of related organisms by discriminating variant haplotypes across hundreds of kilobases of genomic sequence. MAGPhase identified 220 lineage-resolved MAGs in our dataset. The ability to resolve closely related microbes in complex microbial communities improves the identification of biosynthetic gene clusters and the precision of assigning mobile genetic elements to host genomes. We identified 1,400 complete and 350 partial biosynthetic gene clusters, most of which are novel, as well as 424 (298) potential host–viral (host–plasmid) associations using Hi-C data. Metagenome sequencing can now distinguish closely related microbes using long reads and haplotype phasing.
The increasing demand for natural products is currently transforming the meat industry, making grass-fed and finished beef a valuable option for improving profits. However, the transformation of conventional operations to grass-fed systems comprises many modifications, such as logistical, technological, and financial that could be very complex and expensive, involving economic risk. Therefore, in this study, we analyzed the growth curve, critical economic traits, and carcass quality and finished characteristics over several consecutive years in closely related grass-fed and finished Angus steers, to reduce the genetic effect on the results. We found that grass-fed steers require around 188 additional days to reach the market weight (approx. 470 kg) and had approximately 70% less average daily gain compared to the grain-fed and finished steers. Regression analysis demonstrated an interaction between feed and age (P < 0.01); thus, individual regressions were fitted for each regimen style, obtaining almost perfect linear curves for both treatments, which could be straightforwardly used in practical situations due to its simplicity. Six of eight carcass traits were different between grain-fed and grass-fed and finished steers. Hot-carcass weight, dressing, back fat, and quality grade were superior in grain-fed individuals, contrarily to yield grade and ribeye area/carcass ratio, which were better in grass-fed and finished steers (P < 0.05). Interestingly, the meat tenderness was certainly low and similar in both treatments (P = 0.25), indicating the feasibility of producing tender meat with animals under a grass-fed diet. Nevertheless, according to the quality grade analysis, grain-fed carcasses were greater ranked compared to grass-fed bodies (P < 0.01), regardless of their same tenderness. The results will provide valuable information for better understanding beef cattle in grass-feeding finishing systems, especially from weaning to harvest. Additionally, the study will expand the knowledge about the quality of meat obtained from animals that received grass exclusively, becoming relevant information for economic evaluation and management decisions for grass-based cattle operations.
Kimberly M. Davenport1, Alisha T. Massa2, Michelle R. Mousel3,4, Maria K. Herndon2, Stephen N. White2,3,5, Mazdak Salavati6, Emily Clark6, Alan Archibald6, Suraj Bhattarai7, Stephanie D. McKay7, Kim C. Worley8, Brian Dalrymple9, James Kijas10, Alex Caulton11, Shannon Clarke11, Rudiger Brauning11, Tracy Hadfield12, Noelle E. Cockett12, Timothy P.L. Smith13, and Brenda M. Murdoch1,5 on behalf of The Ovine FAANG Project Consortium
Viruses play crucial roles in the ecology of microbial communities, yet they remain relatively understudied in their native environments. Despite many advancements in high-throughput whole-genome sequencing (WGS), sequence assembly, and annotation of viruses, the reconstruction of full-length viral genomes directly from metagenomic sequencing is possible only for the most abundant phages and requires long-read sequencing technologies. Additionally, the prediction of their cellular hosts remains difficult from conventional metagenomic sequencing alone. To address these gaps in the field and to accelerate the study of viruses directly in their native microbiomes, we developed an end-to-end bioinformatics platform for viral genome reconstruction and host attribution from metagenomic data using proximity-ligation sequencing (i.e., Hi-C). We demonstrate the capabilities of the platform by recovering and characterizing the metavirome of a variety of metagenomes, including a fecal microbiome that has also been sequenced with accurate long reads, allowing for the assessment and benchmarking of the new methods. The platform can accurately extract numerous near-complete viral genomes even from highly fragmented short-read assemblies and can reliably predict their cellular hosts with minimal false positives. To our knowledge, this is the first software for performing these tasks. Being significantly cheaper than long-read sequencing of comparable depth, the incorporation of proximity-ligation sequencing in microbiome research shows promise to greatly accelerate future advancements in the field. ### Competing Interest Statement GU, MP, SE, AW, JG, BA, SS, and IL are past or present employees of Phase Genomics. MP is an employee of Inscripta. All other authors have no competing interests.
Sheep are an important agricultural species used for both food and fiber in the United States and globally. A high-quality reference genome enhances the ability to discover genetic and biological mechanisms influencing important traits, such as meat and wool quality. The rapid advances in genome assembly algorithms and emergence of increasingly long sequence read length provide the opportunity for an improved de novo assembly of the sheep reference genome. Tissue was collected postmortem from an adult Rambouillet ewe selected by USDA-ARS for the Ovine Functional Annotation of Animal Genomes project. Short-read (55x coverage), long-read PacBio (75x coverage), and Hi-C data from this ewe were retrieved from public databases. We generated an additional 50x coverage of Oxford Nanopore data and assembled the combined long-read data with canu v1.9. The assembled contigs were polished with Nanopolish v0.12.5 and scaffolded using Hi-C data with Salsa v2.2. Gaps were filled with PBsuite v15.8.24 and polished with Nanopolish v0.12.5 followed by removal of duplicate contigs with PurgeDups v1.0.1. Chromosomes were oriented by identifying centromeres and telomeres with RepeatMasker v4.1.1, indicating a need to reverse the orientation of chromosome 11 relative to Oar_rambouillet_v1.0. Final polishing was performed with two rounds of a pipeline which consisted of freebayes v1.3.1 to call variants, Merfin to validate them, and BCFtools to generate the consensus fasta. The ARS-UI_Ramb_v2.0 assembly has improved continuity (contig N50 of 43.19 Mb) with a 19-fold and 38-fold decrease in the number of scaffolds compared with Oar_rambouillet_v1.0 and Oar_v4.0. ARS-UI_Ramb_v2.0 has greater per-base accuracy and fewer insertions and deletions identified from mapped RNA sequence than previous assemblies. This significantly improved reference assembly, public at NCBI GenBank under accession number GCA_016772045, will optimize the functional annotation of the sheep genome and facilitate improved mapping accuracy of genetic variant and expression data for traits relevant the sheep industry.
Goats and cattle diverged 30 million years ago but retain similarities in immune system genes. Here, the caprine T cell receptor (TCR) gene loci and transcription of its genes were examined and compared to cattle. We annotated the TCR loci using an improved genome assembly (ARS1) of a highly homozygous San Clemente goat. This assembly has already proven useful for describing other immune system genes including antibody and leucocyte receptors. Both the TCRγ (TRG) and TCRδ (TRD) loci were similarly organized in goats as in cattle and the gene sequences were highly conserved. However, the number of genes varied slightly as a result of duplications and differences occurred in mutations resulting in pseudogenes. WC1+ γδ T cells in cattle have been shown to use TCRγ genes from only one of the six available cassettes. The structure of that Cγ gene product is unique and may be necessary to interact with WC1 for signal transduction following antigen ligation. Using RT-PCR and PacBio sequencing, we observed the same restriction for goat WC1+ γδ T cells. In contrast, caprine WC1+ and WC1− γδ T cell populations had a diverse TCRδ gene usage although the propensity for particular gene usage differed between the two cell populations. Noncanonical recombination signal sequences (RSS) largely correlated with restricted expression of TCRγ and δ genes. Finally, caprine γδ T cells were found to incorporate multiple TRD diversity gene sequences in a single transcript, an unusual feature among mammals but also previously observed in cattle.