Pigeons and doves (family Columbidae) are one of the most diverse extant avian lineages, and many species have served as key models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Building genomic resources for columbids is essential to further many of these studies. Here, we present high-quality genome assemblies and annotations for 2 columbid species, Columba livia and Columba guinea. We simultaneously assembled C. livia and C. guinea genomes from long-read sequencing of a single F1 hybrid individual. The new C. livia genome assembly (Cliv_3) shows improved completeness and contiguity relative to Cliv_2.1, with an annotation incorporating long-read IsoSeq data for more accurate gene models. Intensive selective breeding of C. livia has given rise to hundreds of breeds with diverse morphological and behavioral characteristics, and Cliv_3 offers improved tools for mapping the genomic architecture of interesting traits. The C. guinea genome assembly is the first for this species and is a new resource for avian comparative genomics. Together, these assemblies and annotations provide improved resources for functional studies of columbids and avian comparative genomics in general.
Pigeons and doves (family Columbidae) are one of the most diverse extant avian lineages, and many species have served as key models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Building genomic resources for colubids is essential to further many of these studies. Here, we present high-quality genome assemblies and annotations for two columbid species, Columba livia and C. guinea. We simultaneously assembled C. livia and C. guinea genomes from long-read sequencing of a single F1 hybrid individual. The new C. livia genome assembly (Cliv_3) shows improved completeness and contiguity relative to Cliv_2.1, with an annotation incorporating long-read IsoSeq data for more accurate gene models. Intensive selective breeding of C. livia has given rise to hundreds of breeds with diverse morphological and behavioral characteristics, and Cliv_3 offers improved tools for mapping the genomic architecture of interesting traits. The C. guinea genome assembly is the first for this species and is a new resource for avian comparative genomics. Together, these assemblies and annotations provide improved resources for functional studies of columbids and avian comparative genomics in general.
ABSTRACT Pigeons and doves (family Columbidae) are one of the most diverse extant avian lineages, and many species have served as key models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Building genomic resources for colubids is essential to further many of these studies. Here, we present high-quality genome assemblies and annotations for two columbid species, Columba livia and C. guinea . We simultaneously assembled C. livia and C. guinea genomes from long-read sequencing of a single F 1 hybrid individual. The new C. livia genome assembly (Cliv_3) shows improved completeness and contiguity relative to Cliv_2.1, with an annotation incorporating long-read IsoSeq data for more accurate gene models. Intensive selective breeding of C. livia has given rise to hundreds of breeds with diverse morphological and behavioral characteristics, and Cliv_3 offers improved tools for mapping the genomic architecture of interesting traits. The C. guinea genome assembly is the first for this species and is a new resource for avian comparative genomics. Together, these assemblies and annotations provide improved resources for functional studies of columbids and avian comparative genomics in general. ARTICLE SUMMARY Pigeons and doves are important models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Here, we present high-quality reference genome assemblies and annotations for two pigeon species, the domestic rock pigeon ( Columba livia ) and the African speckled pigeon ( C. guinea ). These assemblies and annotations provide improved resources for both comparative genomics and functional studies.
Programmed DNA loss is a gene silencing mechanism that is employed by several vertebrate and nonverte-brate lineages, including all living jawless vertebrates and songbirds. Reconstructing the evolution of somat-ically eliminated (germline-specific) sequences in these species has proven challenging due to a high content of repeats and gene duplications in eliminated sequences and a corresponding lack of highly accurate and contiguous assemblies for these regions. Here, we present an improved assembly of the sea lamprey (Petromyzon marinus) genome that was generated using recently standardized methods that increase the contiguity and accuracy of vertebrate genome assemblies. This assembly resolves highly contiguous, somat-ically retained chromosomes and at least one germline-specific chromosome, permitting new analyses that reconstruct the timing, mode, and repercussions of recruitment of genes to the germline-specific fraction. These analyses reveal major roles of interchromosomal segmental duplication, intrachromosomal duplication, and positive selection for germline functions in the long-term evolution of germline-specific chromosomes.
Congenital myasthenic syndrome (CMS) is a group of 32 disorders involving genetic dysfunction at the neuromuscular junction resulting in skeletal muscle weakness that worsens with physical activity. Precise diagnosis and molecular subtype identification are critical for treatment as medication for one subtype may exacerbate disease in another (Engel et al., Lancet Neurol 14: 420 [2015]; Finsterer, Orphanet J Rare Dis 14: 57 [2019]; Prior and Ghosh, J Child Neurol 36: 610 [2021]). The SNAP25- related CMS subtype (congenital myasthenic syndrome 18, CMS18; MIM #616330) is a rare disorder characterized by muscle fatigability, delayed psychomotor development, and ataxia. Herein, we performed rapid whole-genome sequencing (rWGS) on a critically ill newborn leading to the discovery of an unreported pathogenic de novo SNAP25 c.529C > T; p.Gln177Ter variant. In this report, we present a novel case of CMS18 with complex neonatal consequence. This discovery offers unique insight into the extent of phenotypic severity in CMS18, expands the reported SNAP25 variant phenotype, and paves a foundation for personalized management for CMS18.
BACKGROUND:Genetic disorders contribute to significant morbidity and mortality in critically ill newborns. Despite advances in genome sequencing technologies, a majority of neonatal cases remain unsolved. Complex structural variants (SVs) often elude conventional genome sequencing variant calling pipelines and will explain a portion of these unsolved cases. METHODS:As part of the Utah NeoSeq project, we used a research-based, rapid whole-genome sequencing (WGS) protocol to investigate the genomic etiology for a newborn with a left-sided congenital diaphragmatic hernia (CDH) and cardiac malformations, whose mother also had a history of CDH and atrial septal defect. RESULTS:Using both a novel, alignment-free and traditional alignment-based variant callers, we identified a maternally inherited complex SV on chromosome 8, consisting of an inversion flanked by deletions. This complex inversion, further confirmed using orthogonal molecular techniques, disrupts the ZFPM2 gene, which is associated with both CDH and various congenital heart defects. CONCLUSIONS:Our results demonstrate that complex structural events, which often are unidentifiable or not reported by clinically validated testing procedures, can be discovered and accurately characterized with conventional, short-read sequencing and underscore the utility of WGS as a first-line diagnostic tool.
Vertebrate craniofacial morphogenesis is a highly orchestrated process that is directed by evolutionarily conserved developmental pathways.1,2 Within species, canalized development typically produces modest morphological variation. However, as a result of millennia of artificial selection, the domestic pigeon displays radical craniofacial variation within a single species. One of the most striking cases of pigeon craniofacial variation is the short-beak phenotype, which has been selected in numerous breeds. Classical genetic experiments suggest that pigeon beak length is regulated by a small number of genetic factors, one of which is sex linked (Ku2 locus).3-5 However, the genetic underpinnings of pigeon craniofacial variation remain unknown. Using geometric morphometrics and quantitative trait locus (QTL) mapping on an F2 intercross between a short-beaked Old German Owl (OGO) and a medium-beaked Racing Homer (RH), we identified a single Z chromosome locus that explains a majority of the variation in beak morphology in the F2 population. Complementary comparative genomic analyses revealed that the same locus is strongly differentiated between breeds with short and medium beaks. Within the Ku2 locus, we identified an amino acid substitution in the non-canonical Wnt receptor ROR2 as a putative regulator of pigeon beak length. The non-canonical Wnt pathway serves critical roles in vertebrate neural crest cell migration and craniofacial morphogenesis.6,7 In humans, ROR2 mutations cause Robinow syndrome, a congenital disorder characterized by skeletal abnormalities, including a widened and shortened facial skeleton.8,9 Our results illustrate how the extraordinary craniofacial variation among pigeons can reveal genetic regulators of vertebrate craniofacial diversity.
The iris of the eye shows striking color variation across vertebrate species, and may play important roles in crypsis and communication. The domestic pigeon (Columba livia) has three common iris colors, orange, pearl (white), and bull (dark brown), segregating in a single species, thereby providing a unique opportunity to identify the genetic basis of iris coloration. We used comparative genomics and genetic mapping in laboratory crosses to identify two candidate genes that control variation in iris color in domestic pigeons. We identified a nonsense mutation in the solute carrier SLC2A11B that is shared among all pigeons with pearl eye color, and a locus associated with bull eye color that includes EDNRB2, a gene involved in neural crest migration and pigment development. However, bull eye is likely controlled by a heterogeneous collection of alleles across pigeon breeds. We also found that the EDNRB2 region is associated with regionalized plumage depigmentation (piebalding). Our study identifies two candidate genes for eye colors variation, and establishes a genetic link between iris and plumage color, two traits that vary widely in the evolution of birds and other vertebrates.
Small noncoding microRNAs (miRNAs) play essential roles in post-transcriptional gene regulation in development and disease, predominantly through interactions with mRNA untranslated regions (UTRs). However, the dynamic and tissue specific expression of miRNAs and target UTRs throughout development is poorly characterized. In order to understand the RNA regulatory events that drive cardiac morphogenesis, we captured a simultaneous small RNA and total RNA-seq time course in zebrafish hearts. We extracted de novo transcriptome annotation of dynamically changing UTRs, revealing over ten thousand significantly altered transcripts, with novel 5’ and 3’UTRs for approximately half of all protein coding genes, and predicted regulatory miRNA/UTR interactions during development. In addition to miRNA and mRNA transcriptome resources for studies of RNA regulatory events in cardiac development and disease, this study provides a robust bioinformatic pipeline that should facilitate discovery of posttranscriptional regulatory networks in other developmental systems.
When published, this article did not initially appear open access. This error has been corrected, and the open access status of the paper is noted in all versions of the paper. Additionally, affiliation 16 denoting equal contribution was missing from author Robb Krumlauf in the PDF originally published. This error has also been corrected.
The domestic rock pigeon (Columba livia) is among the most widely distributed and phenotypically diverse avian species. C. livia is broadly studied in ecology, genetics, physiology, behavior, and evolutionary biology, and has recently emerged as a model for understanding the molecular basis of anatomical diversity, the magnetic sense, and other key aspects of avian biology. Here we report an update to the C. livia genome reference assembly and gene annotation dataset. Greatly increased scaffold lengths in the updated reference assembly, along with an updated annotation set, provide improved tools for evolutionary and functional genetic studies of the pigeon, and for comparative avian genomics in general.
Motivation:Genetic variation that disrupts gene function by altering gene splicing between individuals can substantially influence traits and disease. In those cases, accurately predicting the effects of genetic variation on splicing can be highly valuable for investigating the mechanisms underlying those traits and diseases. While methods have been developed to generate high quality computational predictions of gene structures in reference genomes, the same methods perform poorly when used to predict the potentially deleterious effects of genetic changes that alter gene splicing between individuals. Underlying that discrepancy in predictive ability are the common assumptions by reference gene finding algorithms that genes are conserved, well-formed and produce functional proteins. Results:We describe a probabilistic approach for predicting recent changes to gene structure that may or may not conserve function. The model is applicable to both coding and non-coding genes, and can be trained on existing gene annotations without requiring curated examples of aberrant splicing. We apply this model to the problem of predicting altered splicing patterns in the genomes of individual humans, and we demonstrate that performing gene-structure prediction without relying on conserved coding features is feasible. The model predicts an unexpected abundance of variants that create de novo splice sites, an observation supported by both simulations and empirical data from RNA-seq experiments. While these de novo splice variants are commonly misinterpreted by other tools as coding or non-coding variants of little or no effect, we find that in some cases they can have large effects on splicing activity and protein products and we propose that they may commonly act as cryptic factors in disease. Availability and implementation:The software is available from geneprediction.org/SGRF. Supplementary information:Supplementary information is available at Bioinformatics online.
A reference genome sequence for Pseudotsuga menziesii var. menziesii (Mirb.) Franco (Coastal Douglas-fir) is reported, thus providing a reference sequence for a third genus of the family Pinaceae. The contiguity and quality of the genome assembly far exceeds that of other conifer reference genome sequences (contig N50 = 44,136 bp and scaffold N50 = 340,704 bp). Incremental improvements in sequencing and assembly technologies are in part responsible for the higher quality reference genome, but it may also be due to a slightly lower exact repeat content in Douglas-fir vs. pine and spruce. Comparative genome annotation with angiosperm species reveals gene-family expansion and contraction in Douglas-fir and other conifers which may account for some of the major morphological and physiological differences between the two major plant groups. Notable differences in the size of the NDH-complex gene family and genes underlying the functional basis of shade tolerance/intolerance were observed. This reference genome sequence not only provides an important resource for Douglas-fir breeders and geneticists but also sheds additional light on the evolutionary processes that have led to the divergence of modern angiosperms from the more ancient gymnosperms.
Motivation: The accurate interpretation of genetic variants is critical for characterizing genotype‐phenotype associations. Because the effects of genetic variants can depend strongly on their local genomic context, accurate genome annotations are essential. Furthermore, as some variants have the potential to disrupt or alter gene structure, variant interpretation efforts stand to gain from the use of individualized annotations that account for differences in gene structure between individuals or strains. Results : We describe a suite of software tools for identifying possible functional changes in gene structure that may result from sequence variants. ACE (‘Assessing Changes to Exons’) converts phased genotype calls to a collection of explicit haplotype sequences, maps transcript annotations onto them, detects gene‐structure changes and their possible repercussions, and identifies several classes of possible loss of function. Novel transcripts predicted by ACE are commonly supported by spliced RNA‐seq reads, and can be used to improve read alignment and transcript quantification when an individual‐specific genome sequence is available. Using publicly available RNA‐seq data, we show that ACE predictions confirm earlier results regarding the quantitative effects of nonsense‐mediated decay, and we show that predicted loss‐of‐function events are highly concordant with patterns of intolerance to mutations across the human population. ACE can be readily applied to diverse species including animals and plants, making it a broadly useful tool for use in eukaryotic population‐based resequencing projects, particularly for assessing the joint impact of all variants at a locus. Availability and Implementation: ACE is written in open‐source C ++ and Perl and is available from geneprediction.org/ACE Contact: myandell@genetics.utah.edu or tim.reddy@duke.edu Supplementary information: Supplementary information is available at Bioinformatics online.
High-throughput sequencing data are increasingly being made available to the research community for secondary analyses, providing new opportunities for large-scale association studies. However, heterogeneity in target capture and sequencing technologies often introduce strong technological stratification biases that overwhelm subtle signals of association in studies of complex traits. Here, we introduce the Cross-Platform Association Toolkit, XPAT, which provides a suite of tools designed to support and conduct large-scale association studies with heterogeneous sequencing datasets. XPAT includes tools to support cross-platform aware variant calling, quality control filtering, gene-based association testing and rare variant effect size estimation. To evaluate the performance of XPAT, we conducted case-control association studies for three diseases, including 783 breast cancer cases, 272 ovarian cancer cases, 205 Crohn disease cases and 3507 shared controls (including 1722 females) using sequencing data from multiple sources. XPAT greatly reduced Type I error inflation in the case-control analyses, while replicating many previously identified disease-gene associations. We also show that association tests conducted with XPAT using cross-platform data have comparable performance to tests using matched platform data. XPAT enables new association studies that combine existing sequencing datasets to identify genetic loci associated with common diseases and other complex traits.
Abstract Pancreatic adenocarcinoma (PDAC) is the 4th most common cause of cancer deaths in North America, for both men and women with a 5-year survival rate of less than 5%. The poor prognosis rate is attributed to late presentation of the disease and the lack of effective treatment options. Large-scale genome sequencing efforts on PDAC tumors show evidence of high mutational burden and revealed a number of mutated genes affecting multiple oncogenic pathways. While there are significant endeavors in developing specific targeted agents against “driver” mutations, tumor diversity within and across patient population remains a key factor affecting therapeutic efficacy. In this context, the availability of large cohorts of genomically characterized patient-derived xenograft (PDX) tumor models may help to accelerate the development of novel therapies against this lethal cancer. PDX models provide a renewable resource to maintain a patient's tumor ex vivo for pre-clinical or co-clinical studies. As part of The International Cancer Genome Consortium (ICGC), our laboratory has established 93 PDX models in non-obese diabetic and severe combined immune-deficient (NOD-SCID) mice from Whipple resection specimens. These tumors represent a heterogeneous group of neoplasms arising from the head, body and tail of pancreas, bile duct and Ampulla of Vater. All implantations including in the subcutaneous pocket at the flank or at the orthotopic pancreas site, were performed using 4-8 weeks old NOD-SCID mice. Successful growth and serial transplant to multiple mouse generations were observed in in 74 PDX models of the 93 implanted PDAC specimens, achieving an 80% engraftment rate, one of the highest reported in any type of cancer. Histology fidelity was preserved in the PDX models compared to corresponding patient tumors. Failed implants were due to specimens characterized by borderline malignancy and absence of tumor cells. Whole exome sequencing and copy number aberration profiling was completed for 61 PDXs and blood from the matched patients. Cancer-specific single nucleotide variation (SNV) load varied widely from 38 to 305 in PDXs. The most recurrent activating mutation was observed in KRAS with 77% of PDX models showing alterations at codon G12 (65%), G13 (8%) and Q61 (4%); in addition, 26% PDXs had a copy number gain in KRAS. Molecular comparisons of the 21 PDX models and their matched patient tumors showed that alternate allele frequency of KRAS mutation from exome sequencing of primary tumor is a strong indicator of the tumor cellularity; a higher tumor cellularity results in a larger overlap of cancer specific alterations between xenografts and corresponding patient tumors. We have demonstrated a successful establishment of PDX models that represent genomic architecture of major subclonal populations of patient PDAC primary tumors. Citation Format: Nikolina Radulovich, Emin Ibrahimov, Carson Holt, Vibha Raghavan, Tracy Zhao, Rob Denroch, Nhu-An Pham, Steve Gallinger, Melania Pintilie, Lincoln Stein, John McPherson, Lakshmi Muthuswamy, Ming Sound Tsao. Establishment and molecular characterization of patient-derived tumor xenografts from resected tumors or ascites fluids of patients with pancreatic/ampullary/bile duct carcinomas. [abstract]. In: Proceedings of the AACR Special Conference: Patient-Derived Cancer Models: Present and Future Applications from Basic Science to the Clinic; Feb 11-14, 2016; New Orleans, LA. Philadelphia (PA): AACR; Clin Cancer Res 2016;22(16_Suppl):Abstract nr B31.