Oxford Nanopore Technologies (ONT) sequencing platforms have enabled scientists to generate sequence reads of 20 kilobase or more, which has facilitated the assembly of complete genomes. Nanopore sequencing involves recording the change in electrical signals as the DNA molecule transverses the pore. Converting this signal information into bases – a process called basecalling – is challenging and require the implementation of machine learning. The models developed for ONT basecallers have continually improved. However, the mean quality of the reads generated fall below a Phred score of 20 (99% accuracy), the target quality for high-quality genome assembly, when used on challenging plant samples. These low-quality base calls limit the ability to conduct de novo genome assembly, accurate phasing of haplotypes, and long-read genotyping. To overcome these shortcomings, we fine-tuned ONT’s high accuracy basecalling model using Bonito, an AI-based deep learning basecaller, to develop crop-specific models for five important species in the Rosaceae family. These species include highly valuable tree crops such as apple, peach, pear, plum, and sweet cherry. Through the development of an automated training pipeline, we were able to achieve ∼14% higher mean and median basecall quality, and a ∼20% increase in total reads with Phred scores >15. These results were achieved without significantly affecting the read length N50s and total reads called. As a result, these new crop-specific models will enable members of the Rosaceae genomics community to produce higher quality ONT sequencing data. Furthermore, we are releasing our automated pipeline to facilitate others to train their own organism or crop-specific models. ### Competing Interest Statement The authors have declared no competing interest.
The complete genome sequences of apple, peach, and diploid strawberry - one member of each of the three main fruit-producing branches of the Rosaceae tree - were available in 2010. Despite this achievement, virtually none of this genomics knowledge was being used to assist breeding efforts of these crops. Four years later, this gap has been bridged, with genetic information routinely used in many US apple, peach, and cherry breeding programs. For example, DNA tests predict apple crispness, peach maturity date, and cherry fruit size, enabling breeders to determine the best parents to combine and the best seedlings to advance. This application significantly reduces the wasted effort to eliminate entirely poor families and reduces the costs to grow and evaluate thousands of seedlings genetically destined to have unacceptable fruit quality or maturity date. This achievement was enabled by international community efforts, including the RosBREED project, funded by the USDA-National Institute of Food and Agriculture Specialty Crop Research Initiative (SCRI). DNA tests are now applied for high-value attributes where the targeted loci explain a large proportion of the trait variation. However, limitations to widespread adoption of these predictive tests still exist. Some limitations are due to lack of knowledge, such as an understanding of genotype by environment (GxE) interactions and loci associated with variation for other valuable attributes. Technical limitations include streamlined phasing of alleles from multiple families of pedigree-connected breeding germplasm and access to suitable commercial service providers.
TreeGenes and tree fruit Genome Database Resources serve the international forestry and fruit tree genomics research communities, respectively. These databases hold similar sequence data and provide resources for the submission and recovery of this information in order to enable comparative genomics research. Large-scale genotype and phenotype projects have recently spawned the development of independent tools and interfaces within these repositories to deliver information to both geneticists and breeders. The increase in next generation sequencing projects has increased the amount of data as well as the scale of analysis that can be performed. These two repositories are now working towards a similar goal of archiving the diverse, independent data sets generated from genotype/phenotype experiments. This is achieved through focused development on data input standards (templates), pipelines for the storage and automated curation, and consistent annotation efforts through the application of widely accepted ontologies to improve the extraction and exchange of the data for comparative analysis. Efforts towards standardization are not limited to genotype/phenotype experiments but are also being applied to other data types to improve gene prediction and annotation for de novo sequencing projects. The resources developed towards these goals represent the first large-scale coordinated effort in plant databases to add informatics value to diverse genotype/phenotype experiments.
Advances in DNA sequencing technology have significantly reduced the costs associated with sequencing an organism’s genome. However, the operating costs of hardware, software, and labor to analyze the sequence data are still too high for most users to process in house. Henceforth, most of the current bioinformatics applications used by bench scientists will be accessible through a Web environment. This paper presents GenSAS, the Genome Sequence Annotation Server, a JavaScript-based framework of gene prediction and comparative sequence similarity applications for structural and functional sequence annotation. Among other web-based genome annotation pipelines, GenSAS is unique in that it offers a one-stop website with a single graphical interface for running multiple structural and functional annotation tools, visualization and manual curation of genome. We present its functionality, the technology used in implementing each functionality, and software architecture of the overall implementation.
The evolution of separate sexes (dioecy) from hermaphroditism is one of the major evolutionary transitions in plants, and this transition can be accompanied by the development of sex chromosomes. Studies in species with intermediate sexual systems are providing unprecedented insight into the initial stages of sex chromosome evolution. Here, we describe the genetic mechanism of sex determination in the octoploid, subdioecious wild strawberry, Fragaria virginiana Mill., based on a whole-genome simple sequence repeat (SSR)-based genetic map and on mapping sex determination as two qualitative traits, male and female function. The resultant total map length is 2373 c M and includes 212 markers on 42 linkage groups (mean marker spacing: 14 c M ). We estimated that approximately 70 and 90% of the total F. virginiana genetic map resides within 10 and 20 c M of a marker on this map, respectively. Both sex expression traits mapped to the same linkage group, separated by approximately 6 c M , along with two SSR markers. Together, our phenotypic and genetic mapping results support a model of gender determination in subdioecious F. virginiana with at least two linked loci (or gene regions) with major effects. Reconstruction of parental genotypes at these loci reveals that both female and hermaphrodite heterogamety exist in this species. Evidence of recombination between the sex-determining loci, an important hallmark of incipient sex chromosomes, suggest that F. virginiana is an example of the youngest sex chromosome in plants and thus a novel model system for the study of sex chromosome evolution.
A genome-wide framework physical map of peach was constructed using high-information content fingerprinting (HICF) and FPC software. The resulting HICF assembly contained 2,138 contigs composed of 15,655 clones (4.3× peach genome equivalents) from two complementary bacterial artificial chromosome libraries. The total physical length of all contigs is estimated at 303 Mb or 104.5% of the peach genome. The framework physical map is anchored on the Prunus genetic reference map and integrated with the peach transcriptome map. The physical length of anchored contigs is estimated at 45.0 Mb or 15.5% of the genome. Altogether, 2,636 markers, i.e., genetic markers, peach unigene expressed sequence tags, and gene-specific and overgo probes, were incorporated into the physical framework and supported the accuracy of contig assembly.
This research aims at the development of peach as a model genome for the identification, characterization, and cloning of important genes in Rosaceae species. We have constructed an initial physical map for peach. The map is anchored onto the general Prunus genetic map and has been integrated with a peach transcript map (http://www.rosaceae.org). Our efforts to exploit the peach genomic databases are focused on two genomic regions of interest: the distal part of linkage group 2 associated with the genes for root-knot nematode resistance and the distal part of linkage group 4 which is important for fruit characters such as freestone and melting flesh. Our approach combines linkage genetic mapping using EST based genetic markers, localization of the markers on the physical framework, and sequencing the genomic region of interest. Preliminary results illustrate the efficiency of this strategy to facilitate gene discovery in peach.
The importance of high-quality fruit and the intrinsic difficulties of breeding in a perennial species require the development and application of structural and functional genomic databases for the sustained improvement of Rosaceaous fruit crops. Identification and characterisation of genes controlling the genetic basis of the traits, and their tagging with molecular markers, permits facilitated introgression of important characters, speeding development of new breeding material combining the best traits formerly isolated in separate varieties. Currently, there are two major bottlenecks to the identification, characterisation and direct manipulation of genes important to Rosaceae crop cultivation. These are the lack of a candidate gene database and a genetic/physical map resource in which to search for the genes. In cooperation with other researchers world-wide, we are developing such databases for peach as a genome index species for the Rosaceae. The work towards this Rosaceae genome database is presented here in order to update Rosaceae community researchers on the progress of this database and its application to problems in apricot breeding and sustainability.
Pest and disease problems are important constraints of cassava production and host plant resistance is the most efficient method of combating them. Breeding for host plant resistance is considerably slowed down by the crop’s biological constraints of a long growth cycle, high levels of heterozygosity and a large genetic load. More efficient methods such as gene cloning and transgenesis are required to deploy resistance genes. To facilitate the cloning of resistance genes, bacterial artificial chromosome (BAC) library resources have been developed for cassava. Two libraries were constructed from the cassava clones, TMS 30001, resistant to the cassava mosaic disease (CMD) and the cassava bacterial blight (CBB), and MECU72, resistant to cassava white fly. The TMS30001 library has 55 296 clones with an insert size range of 40–150 kb with an average of 80 kb, while the MECU72 library consists of 92 160 clones and an insert size range of 25–250 kb average of 93 kb. Based on a genome size of 772 Mb, the TMS30001 and MECU72 libraries have a 5 and 11.3 haploid genome equivalents and a 95 and 99 chance of finding any sequence, respectively. To demonstrate the potential of the libraries, the TMS30001 library was screened by southern hybridization using a cassava analog (CBB1) of the Xa21 gene from rice that maps to a region containing a QTL for resistance to CBB as probe. Five BAC clones that hybridized to CBB1 were isolated and a Hind III fingerprint revealed 2–3 copies of the gene in individual BAC clones. A larger scale analysis of resistance gene analogs (RGAs) in cassava has also been conducted in order to understand the number and organization of RGAs. To scan for gene and repeat DNA content in the libraries, end-sequencing was performed on 2301 clones from the MECU72 library. A total of 1705 unique sequences were obtained with an average size of 715 bp. Database homology searches using BLAST revealed that 458 sequences had significant homology with known proteins and 321 with transposable elements. The use of the library in positional cloning of pest and disease resistance genes is discussed.
Modern cultivated maize (Zea mays L.)is one of the primary agronomic crops in the USA with an estimated genome size of 2500 megabases (Mb). To develop the resources for positional cloning and structural genomics in maize, we constructed a bacterial artificial chromosome (BAC) library for the inbred line B73 using the cloning enzyme Hind III. The library contains 247 680 clones (645 384‐well plates). A random sampling of 697 clones indicated an average insert size of 136 kilobase (kb) (range = 42 to 379 kb) and 0.4% empty vectors. Screening the colony filters for chloroplast DNA content indicated an exceptionally low 0.18% contamination with chloroplast DNA. Thus, the library provides 13.5 haploid genome equivalents allowing >99% probability of recovering any specific sequence of interest. High‐density filters were gridded robotically using a Genetix Q‐BOT (Hampshire, UK) in a 4 by 4 double‐spotted array on 22.5‐cm2 filters. Partial screening (6× coverage) of the library with 20 single copy probes identified an average 7.1 positive signals per probe, with a range of 3 to 15 positive signals per probe. To evaluate the utility of the library for sequence tagged connector (STC) analysis, 768 BAC clones were end sequenced in both forward and reverse directions giving a total of 1415 successful reads. End sequences were queried against SWISS‐PROT, Genbank NR, MIPS Arabidopsis, maize genomic sequence dbGSS, and maize cDNA database dbEST. Results in spreadsheet format from these searches is publicly available at the CUGI website (www.genome.clemson.edu/projects/stc/maize/ZMMBBb/).
We have constructed a bacterial artificial chromosome (BAC) library for a European honey bee strain using the cloning enzyme HindIII in order to develop resources for structural genomics research. The library contains 36,864 clones (ninety-six 384-well plates). A random sampling of 247 clones indicated an average insert size of 113 kb (range = 27 to 213 kb) and 2% empty vectors. Based on an estimated genome size of 270 Mb, this library provides approximately 15 haploid genome equivalents, allowing >99% probability of recovering any specific sequence of interest. High-density colony filters were gridded robotically using a Genetix Q-BOT in a 4 x 4 double-spotted array on 22.5-cm2 filters. Screening of the library with four mapped honey bee genomic clones and two bee cDNA probes identified an average of 21 positive signals per probe, with a range of 7-38 positive signals per probe. An additional screening was performed with nine aphid gene fragments and one Drosophila gene fragment resulting in seven of the nine aphid probes and the Drosophila probe producing positive signals with a range of 1 to 122 positive signals per probe (average of 45). To evaluate the utility of the library for sequence tagged connector analysis, 1152 BAC clones were end sequenced in both forward and reverse directions, giving a total of 2061 successful reads of high quality. End sequences were queried against SWISS-PROT, insect genomic sequence GSS, insect EST, and insect transposable element databases. Results in spreadsheet format from these searches are publicly available at the Clemson University Genomics Institute (CUGI) website in a searchable format (http://www.genome.clemson.edu/projects/stc/bee/AM__Ba/).
Rice was chosen as a model organism for genome sequencing because of its economic importance, small genome size, and syntenic relationship with other cereal species. We have constructed a bacterial artificial chromosome fingerprint-based physical map of the rice genome to facilitate the whole-genome sequencing of rice. Most of the rice genome ( approximately 90.6%) was anchored genetically by overgo hybridization, DNA gel blot hybridization, and in silico anchoring. Genome sequencing data also were integrated into the rice physical map. Comparison of the genetic and physical maps reveals that recombination is suppressed severely in centromeric regions as well as on the short arms of chromosomes 4 and 10. This integrated high-resolution physical map of the rice genome will greatly facilitate whole-genome sequencing by helping to identify a minimum tiling path of clones to sequence. Furthermore, the physical map will aid map-based cloning of agronomically important genes and will provide an important tool for the comparative analysis of grass genomes.
We have constructed a cotton (Gossypium hirsutum L.) BAC library using the elite Acala-type cultivar Maxxa. The BAC library contains 129 024 clones comprising 8.3 haploid genome equivalents based on a AD genome size of 2118 Mb. A random sampling of 435 BACs indicated an average insert length of 137 kb with a range of 80 to 275 kb, and 1.4% of the BACs do not contain inserts. Of the BAC clones in the sample 99% had an average insert length equal to or greater than 100 kb. Contamination of the genomic library with chloroplast clones was low (1.5%). To evaluate the library for genome representation and to identify clones associated with cotton fiber development, six BAC colony filters (7.2× representation) were screened with PCR-generated gene-specific probes for the GhMYB class of transcription factor genes giving an average of 12 positive signals per probe (range 6–20). To gain a glimpse into the cotton genome and evaluate the library for sequence-tagged connector (STC) development, 1536 BAC clones were end sequenced in both forward and reverse direction. The STCs were queried against the SwissProt and MIPS Arabidopsis databases and significant hits were sorted according to putative function.