Pigeons and doves (family Columbidae) are one of the most diverse extant avian lineages, and many species have served as key models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Building genomic resources for columbids is essential to further many of these studies. Here, we present high-quality genome assemblies and annotations for 2 columbid species, Columba livia and Columba guinea. We simultaneously assembled C. livia and C. guinea genomes from long-read sequencing of a single F1 hybrid individual. The new C. livia genome assembly (Cliv_3) shows improved completeness and contiguity relative to Cliv_2.1, with an annotation incorporating long-read IsoSeq data for more accurate gene models. Intensive selective breeding of C. livia has given rise to hundreds of breeds with diverse morphological and behavioral characteristics, and Cliv_3 offers improved tools for mapping the genomic architecture of interesting traits. The C. guinea genome assembly is the first for this species and is a new resource for avian comparative genomics. Together, these assemblies and annotations provide improved resources for functional studies of columbids and avian comparative genomics in general.
Pigeons and doves (family Columbidae) are one of the most diverse extant avian lineages, and many species have served as key models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Building genomic resources for colubids is essential to further many of these studies. Here, we present high-quality genome assemblies and annotations for two columbid species, Columba livia and C. guinea. We simultaneously assembled C. livia and C. guinea genomes from long-read sequencing of a single F1 hybrid individual. The new C. livia genome assembly (Cliv_3) shows improved completeness and contiguity relative to Cliv_2.1, with an annotation incorporating long-read IsoSeq data for more accurate gene models. Intensive selective breeding of C. livia has given rise to hundreds of breeds with diverse morphological and behavioral characteristics, and Cliv_3 offers improved tools for mapping the genomic architecture of interesting traits. The C. guinea genome assembly is the first for this species and is a new resource for avian comparative genomics. Together, these assemblies and annotations provide improved resources for functional studies of columbids and avian comparative genomics in general.
ABSTRACT Pigeons and doves (family Columbidae) are one of the most diverse extant avian lineages, and many species have served as key models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Building genomic resources for colubids is essential to further many of these studies. Here, we present high-quality genome assemblies and annotations for two columbid species, Columba livia and C. guinea . We simultaneously assembled C. livia and C. guinea genomes from long-read sequencing of a single F 1 hybrid individual. The new C. livia genome assembly (Cliv_3) shows improved completeness and contiguity relative to Cliv_2.1, with an annotation incorporating long-read IsoSeq data for more accurate gene models. Intensive selective breeding of C. livia has given rise to hundreds of breeds with diverse morphological and behavioral characteristics, and Cliv_3 offers improved tools for mapping the genomic architecture of interesting traits. The C. guinea genome assembly is the first for this species and is a new resource for avian comparative genomics. Together, these assemblies and annotations provide improved resources for functional studies of columbids and avian comparative genomics in general. ARTICLE SUMMARY Pigeons and doves are important models for evolutionary genomics, developmental biology, physiology, and behavioral studies. Here, we present high-quality reference genome assemblies and annotations for two pigeon species, the domestic rock pigeon ( Columba livia ) and the African speckled pigeon ( C. guinea ). These assemblies and annotations provide improved resources for both comparative genomics and functional studies.
Haplotype-resolved genome assemblies are important for understanding how combinations of variants impact phenotypes. To date, these assemblies have been best created with complex protocols, such as cultured cells that contain a single-haplotype (haploid) genome, single cells where haplotypes are separated, or co-sequencing of parental genomes in a trio-based approach. These approaches are impractical in most situations. To address this issue, we present FALCON-Phase, a phasing tool that uses ultra-long-range Hi-C chromatin interaction data to extend phase blocks of partially-phased diploid assembles to chromosome or scaffold scale. FALCON-Phase uses the inherent phasing information in Hi-C reads, skipping variant calling, and reduces the computational complexity of phasing. Our method is validated on three benchmark datasets generated as part of the Vertebrate Genomes Project (VGP), including human, cow, and zebra finch, for which high-quality, fully haplotype-resolved assemblies are available using the trio-based approach. FALCON-Phase is accurate without having parental data and performance is better in samples with higher heterozygosity. For cow and zebra finch the accuracy is 97% compared to 80–91% for human. FALCON-Phase is applicable to any draft assembly that contains long primary contigs and phased associate contigs.
Hop (Humulus lupulus L. var Lupulus) is a diploid, dioecious plant with a history of cultivation spanning more than one thousand years. Hop cones are valued for their use in brewing, and around the world, hop has been used in traditional medicine to treat a variety of ailments. Efforts to determine how biochemical pathways responsible for desirable traits are regulated have been challenged by the large, repetitive, and heterozygous genome of hop. We present the first report of a haplotype-phased assembly of a large plant genome. Our assembly and annotation of the Cascade cultivar genome is the most extensive to date. PacBio long-read sequences from hop were assembled with FALCON and phased with FALCON-Unzip. Using the diploid assembly to assess haplotype variation, we discovered genes under positive selection enriched for stress-response, growth, and flowering functions. Comparative analysis of haplotypes provides insight into large-scale structural variation and the selective pressures that have driven hop evolution. Previous studies estimated repeat content at around 60%. With improved resolution of long terminal retrotransposons (LTRs) due to long-read sequencing, we found that hop is nearly 78% repetitive. Our quantification of repeat content provides context for the size of the hop genome, and supports the hypothesis of whole genome duplication (WGD), rather than expansion due to LTRs. With our more complete assembly, we have identified a homolog of cannabidiolic acid synthase (CBDAS) that is expressed in multiple tissues. The approaches we developed to analyze a phased, diploid assembly serve to deepen our understanding of the genomic landscape of hop and may have broader applicability to the study of other large, complex genomes.
Haplotype-resolved de novo assembly is the ultimate solution to the study of sequence variations in a genome. However, existing algorithms either collapse heterozygous alleles into one consensus copy or fail to cleanly separate the haplotypes to produce high-quality phased assemblies. Here we describe hifiasm, a de novo assembler that takes advantage of long high-fidelity sequence reads to faithfully represent the haplotype information in a phased assembly graph. Unlike other graph-based assemblers that only aim to maintain the contiguity of one haplotype, hifiasm strives to preserve the contiguity of all haplotypes. This feature enables the development of a graph trio binning algorithm that greatly advances over standard trio binning. On three human and five nonhuman datasets, including California redwood with a ~30-Gb hexaploid genome, we show that hifiasm frequently delivers better assemblies than existing tools and consistently outperforms others on haplotype-resolved assembly.
In recent years, improved sequencing technology and computational tools have made de novo genome assembly more accessible. Many approaches, however, generate either an unphased or only partially resolved representation of a diploid genome, in which polymorphisms are detected but not assigned to one or the other of the homologous chromosomes. Yet chromosomal phase information is invaluable for the understanding of phenotypic trait inheritance in the cases of compound heterozygosity, allele-specific expression or cis-acting variants. Here we use a combination of tools and sequencing technologies to generate a de novo diploid assembly of the human primary cell line WI-38. First, data from PacBio single molecule sequencing and Bionano Genomics optical mapping were combined to generate an unphased assembly. Next, 10x Genomics linked reads were combined with the hybrid assembly to generate a partially phased assembly. Lastly, we developed and optimized methods to use short-read (Illumina) sequencing of flow cytometry-sorted metaphase chromosomes to provide phase information. The final genome assembly was almost fully (94%) phased with the addition of approximately 2.5-fold coverage of Illumina data from the sequenced metaphase chromosomes. The diploid nature of the final de novo genome assembly improved the resolution of structural variants between the WI-38 genome and the human reference genome. The phased WI-38 sequence data are available for browsing and download at wi38.research.calicolabs.com. Our work shows that assembling a completely phased diploid genome de novo from the DNA of a single individual is now readily achievable.
Cannabis is a diverse and polymorphic species. To better understand cannabinoid synthesis inheritance and its impact on pathogen resistance, we shotgun sequenced and assembled a trio (sibling pair and their offspring) utilizing long read single molecule sequencing. This resulted in the most contiguous assemblies to date. These reference assemblies were further annotated with full-length male and female mRNA sequencing (Iso-Seq) to help inform isoform complexity, gene model predictions and identification of the Y chromosome. To further annotate the genetic diversity in the species, 40 male, female, and monoecious cannabis and hemp varietals were evaluated for copy number variation (CNV) and RNA expression. This identified multiple CNVs governing cannabinoid expression and 82 genes associated with resistance to , the causal agent of powdery mildew in cannabis. Results indicated that breeding for plants with low tetrahydrocannabinolic acid (THCA) concentrations may result in deletion of pathogen resistance genes. Low THCA cultivars also have a polymorphism every 51 bases while dispensary grade high THCA cannabis exhibited a variant every 73 bases. A refined genetic map of the variation in cannabis can guide more stable and directed breeding efforts for desired chemotypes and pathogen-resistant cultivars.
Non-human assemblies evaluated in Cheng et al (2021). Human assemblies are available via doi:10.5281/zenodo.4393631.
The sequence and assembly of human genomes using long-read sequencing technologies has revolutionized our understanding of structural variation and genome organization. We compared the accuracy, continuity, and gene annotation of genome assemblies generated from either high-fidelity (HiFi) or continuous long-read (CLR) datasets from the same complete hydatidiform mole human genome. We find that the HiFi sequence data assemble an additional 10% of duplicated regions and more accurately represent the structure of tandem repeats, as validated with orthogonal analyses. As a result, an additional 5 Mbp of pericentromeric sequences are recovered in the HiFi assembly, resulting in a 2.5-fold increase in the NG50 within 1 Mbp of the centromere (HiFi 480.6 kbp, CLR 191.5 kbp). Additionally, the HiFi genome assembly was generated in significantly less time with fewer computational resources than the CLR assembly. Although the HiFi assembly has significantly improved continuity and accuracy in many complex regions of the genome, it still falls short of the assembly of centromeric DNA and the largest regions of segmental duplication using existing assemblers. Despite these shortcomings, our results suggest that HiFi may be the most effective stand-alone technology for de novo assembly of human genomes.
The DNA sequencing technologies in use today produce either highly accurate short reads or less-accurate long reads. We report the optimization of circular consensus sequencing (CCS) to improve the accuracy of single-molecule real-time (SMRT) sequencing (PacBio) and generate highly accurate (99.8%) long high-fidelity (HiFi) reads with an average length of 13.5 kilobases (kb). We applied our approach to sequence the well-characterized human HG002/NA24385 genome and obtained precision and recall rates of at least 99.91% for single-nucleotide variants (SNVs), 95.98% for insertions and deletions <50 bp (indels) and 95.99% for structural variants. Our CCS method matches or exceeds the ability of short-read sequencing to detect small variants and structural variants. We estimate that 2,434 discordances are correctable mistakes in the 'genome in a bottle' (GIAB) benchmark set. Nearly all (99.64%) variants can be phased into haplotypes, further improving variant detection. De novo genome assembly using CCS reads alone produced a contiguous and accurate genome with a contig N50 of >15 megabases (Mb) and concordance of 99.997%, substantially outperforming assembly with less-accurate long reads.
ABSTRACT Haplotype-resolved genome assemblies are important for understanding how combinations of variants impact phenotypes. These assemblies can be created in various ways, such as use of tissues that contain single-haplotype (haploid) genomes, or by co-sequencing of parental genomes, but these approaches can be impractical in many situations. We present FALCON-Phase, which integrates long-read sequencing data and ultra-long-range Hi-C chromatin interaction data of a diploid individual to create high-quality, phased diploid genome assemblies. The method was evaluated by application to three datasets, including human, cattle, and zebra finch, for which high-quality, fully haplotype resolved assemblies were available for benchmarking. Phasing algorithm accuracy was affected by heterozygosity of the individual sequenced, with higher accuracy for cattle and zebra finch (>97%) compared to human (82%). In addition, scaffolding with the same Hi-C chromatin contact data resulted in phased chromosome-scale scaffolds.
While genome assembly projects have been successful in many haploid and inbred species, the assembly of noninbred or rearranged heterozygous genomes remains a major challenge. To address this challenge, we introduce the open-source FALCON and FALCON-Unzip algorithms (https://github.com/PacificBiosciences/FALCON/) to assemble long-read sequencing data into highly accurate, contiguous, and correctly phased diploid genomes. We generate new reference sequences for heterozygous samples including an F1 hybrid of Arabidopsis thaliana, the widely cultivated Vitis vinifera cv. Cabernet Sauvignon, and the coral fungus Clavicorona pyxidata, samples that have challenged short-read assembly approaches. The FALCON-based assemblies are substantially more contiguous and complete than alternate short-or long-read approaches. The phased diploid assembly enabled the study of haplotype structure and heterozygosities between homologous chromosomes, including the identification of widespread heterozygous structural variation within coding sequences.
In Hawaiʻi, Acropora cytherea (Dana, 1846) is essentially restricted to the central portion of the archipelago in the Northwestern Hawaiian Islands (NWHI), although it was present in the Main Hawaiian Islands (MHI) in the fossil record (Grigg 1981).Recently however, two relatively young (<5 years) colonies were first documented (Kenyon 2007) from Kauaʻi in the MHI (Fig. 1).Grigg (1981) proposed Johnson Atoll was the origin of the recent Hawaiian Acropora fauna, and computer simulations predict larval transport corridors linking Johnston Atoll with both French Frigate Shoals and Kauaʻi (Kobayashi 2006).Alternatively, these newly discovered colonies could be a southerly range expansion from the large NWHI population, or originate from another source not previously considered.To test these alternative hypotheses regarding the source population, samples were collected from one of these juvenile A. cytherea colonies on Kauaʻi (23.12N, 159.61W), 51 colonies at French Frigate Shoals (23.75N, 166.15W), 57 at Johnston Atoll (16.73N, 169.53W) and 25 at Kingman Reef (6.38N, 162.42W) in the Northern Line Islands; sites separated from Kauaʻi by roughly 700 km, 1200 km, and 1750 km, respectively.These sites represent the closest emergent coral reef habitat to the MHI on which A. cf cytherea corals are known to exist.Samples were identified by Jean Kenyon and Jim Maragos, and genotyped for each of seven microsatellite loci as outlined in Concepcion et al. (2010).Pairwise FST values among the possible source sites ranged from a low of 0.016 (FFS to JO) to a high of 0.06 (FFS to KI).Comparison of these sites with Kwajalein Atoll (8.72N, 167.73E) gives pairwise FST values on the order of 0.15, suggesting that the more proximate sites are the more likely sources of these colonists.Pairwise FST values among these sites are certainly on the low end of the range recommended for confident assignment to source populations (reviewed by Manel et al. 2005), but equivalent to those used in some successful applications of this approach (e.g., Guillemaud et al. 2010).An assignment test was then performed and probabilities of assignment to a reference population were computed with 100,000 iterations of a Monte Carlo resampling algorithm as implemented in GENECLASS2 (Piry et al. 2004).Surprisingly, the population with the highest probability of assignment (99.9%) was the Kingman Reef in the Line Islands, followed distantly by French Frigate Shoals (78.5%) and Johnston Atoll (75.2%).None of the populations could be positively excluded as possible sources, which is consistent with the relatively low pairwise FST values among these populations, but it is striking that the highest likelihood of assignment is to the most distant population, which has not
Montipora capitata Dana, 1846 is one of the most successful reef-building corals in the Hawaiian Archipelago, both in terms of geographic distribution and relative abundance. Here, we examine population genetic structure using eight microsatellite loci to make inferences about exchange among geographical regions throughout Hawaiian waters to inform management and conservation efforts. We collected biopsy samples (n = 560) from colonies at each of 11 islands/atolls along the archipelago in addition to Johnston Atoll, about 1328 km to the southwest. We found very few potential clones (<2%) in our sampling (551 of 560 colonies had unique multi-locus genotypes), indicating that reproduction is predominantly sexual. Likewise, significant genetic structuring among most locations (pairwise F-ST' = 0.05 to 0.49, only two <0.10; P < 0.01) indicates that gene flow between islands is highly limited. Overall, we found four main regional genetic groupings of M. capitata within state waters, one comprised of the Main Hawaiian Islands, one off the three northwestern-most Hawaiian Islands, and two groupings encompassing the middle of the northwestern chain and Johnston Atoll. Despite the potential for extended pelagic larval development periods (>200 d), estimates of contemporary dispersal were uniformly low, with most sites being estimated at >90% self-recruitment. These data imply that the majority of M. capitata colonies found at a given island/atoll across the Hawaiian Archipelago are derived from self-recruitment, and argue for more local-scale management of coral reef resources than has been considered to date.
The alveolates are composed of three major lineages, the ciliates, dinoflagellates, and apicomplexans. Together these 'protist' taxa play key roles in primary production and ecology, as well as in illness of humans and other animals. The interface between the dinoflagellate and apicomplexan clades has been an area of recent discovery, blurring the distinction between these two clades. Moreover, phylogenetic analysis has yet to determine the position of basal dinoflagellate clades hence the deepest branches of the dinoflagellate tree currently remain unresolved. Large-scale mRNA sequencing was applied to 11 species of dinoflagellates, including strains of the syndinean genera Hematodinium and Amoebophrya, parasites of crustaceans and dinoflagellates, respectively, to optimize and update the dinoflagellate tree. From the transcriptome-scale data a total of 73 ribosomal protein-coding genes were selected for phylogeny. After individual gene orthology assessment, the genes were concatenated into a >15,000 amino acid alignment with 76 taxa from dinoflagellates, apicomplexans, ciliates, and the outgroup heterokonts. Overall the tree was well resolved and supported, when the data was subsampled with gblocks or constraint trees were tested with the approximately unbiased test. The deepest branches of the dinoflagellate tree can now be resolved with strong support, and provides a clearer view of the evolution of the distinctive traits of dinoflagellates.
Parental effects are ubiquitous in nature and in many organisms play a particularly critical role in the transfer of symbionts across generations; however, their influence and relative importance in the marine environment has rarely been considered. Coral reefs are biologically diverse and productive marine ecosystems, whose success is framed by symbiosis between reef-building corals and unicellular dinoflagellates in the genus Symbiodinium. Many corals produce aposymbiotic larvae that are infected by Symbiodinium from the environment (horizontal transmission), which allows for the acquisition of new endosymbionts (different from their parents) each generation. In the remaining species, Symbiodinium are transmitted directly from parent to offspring via eggs (vertical transmission), a mechanism that perpetuates the relationship between some or all of the Symbiodinium diversity found in the parent through multiple generations. Here we examine vertical transmission in the Hawaiian coral Montipora capitata by comparing the Symbiodinium ITS2 sequence assemblages in parent colonies and the eggs they produce. Parental effects on sequence assemblages in eggs are explored in the context of the coral genotype, colony morphology, and the environment of parent colonies. Our results indicate that ITS2 sequence assemblages in eggs are generally similar to their parents, and patterns in parental assemblages are different, and reflect environmental conditions, but not colony morphology or coral genotype. We conclude that eggs released by parent colonies during mass spawning events are seeded with different ITS2 sequence assemblages, which encompass phylogenetic variability that may have profound implications for the development, settlement and survival of coral offspring.
Determining the geographic scale at which to apply ecosystem-based management (EBM) has proven to be an obstacle for many marine conservation programs. Generalizations based on geographic proximity, taxonomy, or life history characteristics provide little predictive power in determining overall patterns of connectivity, and therefore offer little in terms of delineating boundaries for marine spatial management areas. Here, we provide a case study of 27 taxonomically and ecologically diverse species (including reef fishes, marine mammals, gastropods, echinoderms, cnidarians, crustaceans, and an elasmobranch) that reveal four concordant barriers to dispersal within the Hawaiian Archipelago which are not detected in single-species exemplar studies. We contend that this multispecies approach to determine concordant patterns of connectivity is an objective and logical way in which to define the minimum number of management units and that EBM in the Hawaiian Archipelago requires at least five spatially managed regions.