Background Most animals and plants have more than one set of chromosomes and package these haplotypes into a single nucleus within each cell. In contrast, many fungal species carry multiple haploid nuclei per cell. Rust fungi are such species with two nuclei (karyons) that contain a full set of haploid chromosomes each. The physical separation of haplotypes in dikaryons means that, unlike in diploids, Hi-C chromatin contacts between haplotypes are false-positive signals. Results We generate the first chromosome-scale, fully-phased assembly for the dikaryotic leaf rust fungus Puccinia triticina and compare Nanopore MinION and PacBio HiFi sequence-based assemblies. We show that false-positive Hi-C contacts between haplotypes are predominantly caused by phase switches rather than by collapsed regions or Hi-C read mis-mappings. We introduce a method for phasing of dikaryotic genomes into the two haplotypes using Hi-C contact graphs, including a phase switch correction step. In the HiFi assembly, relatively few phase switches occur, and these are predominantly located at haplotig boundaries and can be readily corrected. In contrast, phase switches are widespread throughout the Nanopore assembly. We show that haploid genome read coverage of 30–40 times using HiFi sequencing is required for phasing of the leaf rust genome, with 0.7% heterozygosity, and that HiFi sequencing resolves genomic regions with low heterozygosity that are otherwise collapsed in the Nanopore assembly. Conclusions This first Hi-C based phasing pipeline for dikaryons and comparison of long-read sequencing technologies will inform future genome assembly and haplotype phasing projects in other non-haploid organisms.
Background Most animals and plants have more than one set of chromosomes and package these haplotypes into a single nucleus within each cell. In contrast, many fungal species carry multiple haploid nuclei per cell. Rust fungi are such species with two nuclei (karyons) that contain a full set of haploid chromosomes each. The physical separation of haplotypes in dikaryons means that, unlike in diploids, Hi-C chromatin contacts between haplotypes are false positive signals. Results We generate the first chromosome-scale, fully-phased assembly for the dikaryotic leaf rust fungus Puccinia triticina and compare Nanopore MinION and PacBio HiFi sequence-based assemblies. We show that false positive Hi-C contacts between haplotypes are predominantly caused by phase switches rather than by collapsed regions or Hi-C read mis-mappings. We introduce a method for phasing of dikaryotic genomes into the two haplotypes using Hi-C contact graphs, including a phase switch correction step. In the HiFi assembly, relatively few phase switches occur, and these are predominantly located at haplotig boundaries and can be readily corrected. In contrast, phase switches are widespread throughout the Nanopore assembly. We show that haploid genome read coverage of 30-40 times using HiFi sequencing is required for phasing of the leaf rust genome (~0.7% heterozygosity) and that HiFi sequencing resolves genomic regions with low heterozygosity that are otherwise collapsed in the Nanopore assembly. Conclusions This first Hi-C based phasing pipeline for dikaryons and comparison of long-read sequencing technologies will inform future genome assembly and haplotype phasing projects in other non-haploid organisms.
Extracting pure high-molecular weight DNA from some fungal species is difficult due to the presence of polysaccharides and potentially other compounds which biochemically mimic DNA or interfere with the DNA extraction process. Such compounds can co-elute with DNA in many extraction methods, being difficult to separate fom the DNA. Although the contaminant may not be detected by spectrophotometers or fluorometric devices, it substantially interferes with long-read DNA sequencing, such as Oxford Nanopore Technologies. To partially resolve this, a protocol is presented with some updates to current strategies and incorporates a small fragment removal step using Polyethylene Glycol. Using this protocol, we have been successfully sequencing the lentil pathogen Ascochyta lentis with a MinION (Oxford Nanopore Technologies). Sequencing yields have surpassed 13 gigabases with an N50 of approximately 15 kb. To increase sequencing output, more work is needed to identify and remove the elusive contaminants. DNA extraction modified from: https://www.protocols.io/view/high-molecular-weight-dna-extraction-from-challeng-5isg4ee PEG small fragment elimination after: https://www.protocols.io/view/size-selective-precipitation-of-dna-using-peg-amp-7erhjd6
Extracting pure high-molecular weight DNA from some fungal species is difficult due to the presence of polysaccharides and potentially other compounds which biochemically mimic DNA or interfere with the DNA extraction process. Such compounds can co-elute with DNA in many extraction methods, being difficult to separate fom the DNA. Although the contaminant may not be detected by spectrophotometers or fluorometric devices, it substantially interferes with long-read DNA sequencing, such as Oxford Nanopore Technologies. To partially resolve this, a protocol is presented with some updates to current strategies and incorporates a gel purification with a Pippin Prep (Sage Science). Using this protocol, we have been successfully sequencing the wheat stripe rust Puccinia striiformis and leaf rust Puccinia triticina with a MinION (Oxford Nanopore Technologies). Sequencing yields have surpassed 4 gigabases with an N50 of approximately 30 kb. To increase sequencing output, more work is needed to identify and remove the elusive contaminants.