With the advent of long-read DNA sequencing technologies, assembling eukaryotic genomes has become routine; however, properly phasing the maternal and paternal contributions, which is of great value for breeding programs, remains technically challenging. Here, we use the trio-binning approach to separate Oxford Nanopore reads derived from a Cannabis F1 wide cross, made between the Colombian landrace Punto Rojo and the Colorado CBD clone Cherry Pie #16. Reads were obtained from a single PromethION flow cell, generating assemblies with coverage of just 18 × per haplotype, but with good contiguity and gene completeness, demonstrating that it is a cost-effective approach for genome-wide and high-quality haplotype phasing. Evaluated through the lenses of disease resistance and secondary metabolite synthesis, both being traits of interest for the Cannabis industry, we report copy number and structural variation that, as has recently been shown for other major crops, may contribute to phenotypic variation along several relevant dimensions.
With the current speed of sequencing, there is a desire for standardized and automated genome assembly and annotation to produce high-quality genomes as input for comparative (pan)genomics. Therefore, we created a convenience pipeline using existing tools that creates annotated genome assemblies from HiFi (and optionally ultra-long ONT and/or Hi-C) reads for a set of related individuals as well as a related reference genome. Our pipeline is species-agnostic and generates an extensive quality assessment report that can be used for manual filtering and refinement of the assembly and annotation. It includes statistics for individual completeness and contamination assessments as well as a concise pangenome view. The pipeline is implemented in Snakemake and available with a GPLv3 licence at GitHub under github.com/dirkjanvw/MoGAAAP, at Zenodo under doi.org/10.5281/zenodo.14833021, and can be installed through Bioconda.
Abstract Dietary flavonoids play an important role in human nutrition and health. Flavonoid biosynthesis genes have recently been identified in lettuce (Lactuca sativa); however, few mutants have been characterized. We now report the causative mutations in Green Super Lettuce (GSL), a natural light green mutant derived from red cultivar NAR; and GSL-Dark Green (GSL-DG), an olive-green natural derivative of GSL. GSL harbors CACTA 1 (LsC1), a 3.9-kb active nonautonomous CACTA superfamily transposon inserted in the 5′ untranslated region of anthocyanidin synthase (ANS), a gene coding for a key enzyme in anthocyanin biosynthesis. Both terminal inverted repeats (TIRs) of this transposon were intact, enabling somatic excision of the mobile element, which led to the restoration of ANS expression and the accumulation of red anthocyanins in sectors on otherwise green leaves. GSL-DG harbors CACTA 2 (LsC2), a 1.1-kb truncated copy of LsC1 that lacks one of the TIRs, rendering the transposon inactive. RNA-sequencing and reverse transcription quantitative PCR of NAR, GSL, and GSL-DG indicated the relative expression level of ANS was strongly influenced by the transposon insertions. Analysis of flavonoid content indicated leaf cyanidin levels correlated positively with ANS expression. Bioinformatic analysis of the cv Salinas lettuce reference genome led to the discovery and characterization of an LsC1 transposon family with a putative transposon copy number greater than 1,700. Homologs of tnpA and tnpD, the genes encoding two proteins necessary for activation of transposition of CACTA elements, were also identified in the lettuce genome.
Flower opening and closure are traits of reproductive importance in all angiosperms because they determine the success of self- and cross-pollination. The temporal nature of this phenotype rendered it a difficult target for genetic studies. Cultivated and wild lettuce, Lactuca spp., have composite inflorescences comprised of multiple florets that open only once. Different accessions were observed to flower at different times of day. An F6 recombinant inbred line population (RIL) had been derived from accessions of L. serriola x L. sativa that originated from different environments and differed markedly for daily floral opening time. This population was used to map the genetic determinants of this trait; the floral opening time of 236 RILs was scored over a seven-hour period using time-course image series obtained by drone-based remote phenotyping on two occasions, one week apart. Floral pixels were identified from the images using a support vector machine (SVM) machine learning algorithm with an accuracy above 99%. A Bayesian inference method was developed to extract the peak floral opening time for individual genotypes from the time-stamped image data. Two independent QTLs, qDFO2.1 (Daily Floral Opening 2.1) and qDFO8.1, were discovered. Together, they explained more than 30% of the phenotypic variation in floral opening time. Candidate genes with non-synonymous polymorphisms in coding sequences were identified within the QTLs. This study demonstrates the power of combining remote imaging, machine learning, Bayesian statistics, and genome-wide marker data for studying the genetics of recalcitrant phenotypes such as floral opening time. One sentence summary Machine learning and Bayesian analyses of drone-mediated remote phenotyping data revealed two genetic loci regulating differential daily flowering time in lettuce (Lactuca spp.).
Plant mitochondrial genomes are usually assembled and displayed as circular maps based on the widely-held view across the broad community of life scientists that circular genome-sized molecules are the primary form of plant mitochondrial DNA, despite the understanding by plant mitochondrial researchers that this is an inaccurate and outdated concept. Many plant mitochondrial genomes have one or more pairs of large repeats that can act as sites for inter- or intramolecular recombination, leading to multiple alternative arrangements (isoforms). Most mitochondrial genomes have been assembled using methods unable to capture the complete spectrum of isoforms within a species, leading to an incomplete inference of their structure and recombinational activity. To document and investigate underlying reasons for structural diversity in plant mitochondrial DNA, we used long-read (PacBio) and short-read (Illumina) sequencing data to assemble and compare mitochondrial genomes of domesticated (Lactuca sativa) and wild (L. saligna and L. serriola) lettuce species. We characterized a comprehensive, complex set of isoforms within each species and compared genome structures between species. Physical analysis of L. sativa mtDNA molecules by fluorescence microscopy revealed a variety of linear, branched, and circular structures. The mitochondrial genomes for L. sativa and L. serriola were identical in sequence and arrangement and differed substantially from L. saligna, indicating that the mitochondrial genome structure did not change during domestication. From the isoforms in our data, we infer that recombination occurs at repeats of all sizes at variable frequencies. The differences in genome structure between L. saligna and the two other Lactuca species can be largely explained by rare recombination events that rearranged the structure. Our data demonstrate that representations of plant mitochondrial genomes as simple, circular molecules are not accurate descriptions of their true nature and that in reality plant mitochondrial DNA is a complex, dynamic mixture of forms.
Lettuce (Lactuca sativa) is a major crop and a member of the large, highly successful Compositae family of flowering plants. Here we present a reference assembly for the species and family. This was generated using whole-genome shotgun Illumina reads plus in vitro proximity ligation data to create large superscaffolds; it was validated genetically and superscaffolds were oriented in genetic bins ordered along nine chromosomal pseudomolecules. We identify several genomic features that may have contributed to the success of the family, including genes encoding Cycloidea-like transcription factors, kinases, enzymes involved in rubber biosynthesis and disease resistance proteins that are expanded in the genome. We characterize 21 novel microRNAs, one of which may trigger phasiRNAs from numerous kinase transcripts. We provide evidence for a whole-genome triplication event specific but basal to the Compositae. We detect 26% of the genome in triplicated regions containing 30% of all genes that are enriched for regulatory sequences and depleted for genes involved in defence.
David Bertioli and colleagues report the genomes of Arachis duranensis and Arachis ipaensis, the diploid ancestors of cultivated peanut, Arachis hypogaea. Their analyses are a first step in understanding the evolution of the peanut's tetraploid genome. Cultivated peanut (Arachis hypogaea) is an allotetraploid with closely related subgenomes of a total size of ∼2.7 Gb. This makes the assembly of chromosomal pseudomolecules very challenging. As a foundation to understanding the genome of cultivated peanut, we report the genome sequences of its diploid ancestors (Arachis duranensis and Arachis ipaensis). We show that these genomes are similar to cultivated peanut's A and B subgenomes and use them to identify candidate disease resistance genes, to guide tetraploid transcript assemblies and to detect genetic exchange between cultivated peanut's subgenomes. On the basis of remarkably high DNA identity of the A. ipaensis genome and the B subgenome of cultivated peanut and biogeographic evidence, we conclude that A. ipaensis may be a direct descendant of the same population that contributed the B subgenome to cultivated peanut.
Our ability to assemble complex genomes and construct ultradense genetic maps now allows the determination of recombination rates, translocations, and the extent of genomic collinearity between populations, species, and genera. We developed two ultradense genetic linkage maps for pepper from single-position polymorphisms (SPPs) identified de novo with a 30,173 unigene pepper genotyping array. The Capsicum frutescens × C. annuum interspecific and the C. annuum intraspecific genetic maps were constructed comprising 16,167 and 3,878 unigene markers in 2108 and 783 genetic bins, respectively. Accuracies of marker groupings and orders are validated by the high degree of collinearity between the two maps. Marker density was sufficient to locate the chromosomal breakpoint resulting in the P1/P8 translocation between C. frutescens and C. annuum to a single bin. The two maps aligned to the pepper genome showed varying marker density along the chromosomes. There were extensive chromosomal regions with suppressed recombination and reduced intraspecific marker density. These regions corresponded to the pronounced nonrecombining pericentromeric regions in tomato, a related Solanaceous species. Similar to tomato, the extent of reduced recombination appears to be more pronounced in pepper than in other plant species. Alignment of maps with the tomato and potato genomes shows the presence of previously known translocations and a translocation event that was not observed in previous genetic maps of pepper.
Of the over 50 phenotypic resistance genes mapped in lettuce, 25 colocalize to three major resistance clusters (MRC) on chromosomes 1, 2, and 4. Similarly, the majority of candidate resistance genes encoding nucleotide binding-leucine rich repeat (NLR) proteins genetically colocalize with phenotypic resistance loci. MRC1 and MRC4 span over 66 and 63 Mb containing 84 and 21 NLR-encoding genes, respectively, as well as 765 and 627 genes that are not related to NLR genes. Forward and reverse genetic approaches were applied to dissect MRC1 and MRC4. Transgenic lines exhibiting silencing were selected using silencing of β-glucuronidase as a reporter. Silencing of two of five NLR-encoding gene families resulted in abrogation of nine of 14 tested resistance phenotypes mapping to these two regions. At MRC1, members of the coiled coil-NLR-encoding RGC1 gene family were implicated in host and nonhost resistance through requirement for Dm5/8- and Dm45-mediated resistance to downy mildew caused by Bremia lactucae as well as the hypersensitive response to effectors AvrB, AvrRpm1, and AvrRpt2 of the nonpathogen Pseudomonas syringae. At MRC4, RGC12 family members, which encode toll interleukin receptor-NLR proteins, were implicated in Dm4-, Dm7-, Dm11-, and Dm44-mediated resistance to B. lactucae. Lesions were identified in the sequence of a candidate gene within dm7 loss-of-resistance mutant lines, confirming that RGC12G confers Dm7.
Premise of the study: The Compositae (Asteraceae) are a large and diverse family of plants, and the most comprehensive phylogeny to date is a meta-tree based on 10 chloroplast loci that has several major unresolved nodes. We describe the development of an approach that enables the rapid sequencing of large numbers of orthologous nuclear loci to facilitate efficient phylogenomic analyses.Methods and Results: We designed a set of sequence capture probes that target conserved orthologous sequences in the Compositae. We also developed a bioinformatic and phylogenetic workflow for processing and analyzing the resulting data. Application of our approach to 15 species from across the Compositae resulted in the production of phylogenetically informative sequence data from 763 loci and the successful reconstruction of known phylogenetic relationships across the family.Conclusions: These methods should be of great use to members of the broader Compositae community, and the general approach should also be of use to researchers studying other families.
The experimental induction of RNA silencing in plants often involves expression of transgenes encoding inverted repeat (IR) sequences to produce abundant dsRNAs that are processed into small RNAs (sRNAs). These sRNAs are key mediators of post-transcriptional gene silencing (PTGS) and determine its specificity. Despite its application in agriculture and broad utility in plant research, the mechanism of IR-PTGS is incompletely understood. We generated four sets of 60 Arabidopsis plants, each containing IR transgenes expressing different configurations of uidA and CHALCONE SYNTHASE (At-CHS) gene fragments. Levels of PTGS were found to depend on the orientation and position of the fragment in the IR construct. Deep sequencing and mapping of sRNAs to corresponding transgene-derived and endogenous transcripts identified distinctive patterns of differential sRNA accumulation that revealed similarities among sRNAs associated with IR-PTGS and endogenous sRNAs linked to uncapped mRNA decay. Detailed analyses of poly-A cleavage products from At-CHS mRNA confirmed this hypothesis. We also found unexpected associations between sRNA accumulation and the presence of predicted open reading frames in the trigger sequence. In addition, strong IR-PTGS affected the prevalence of endogenous sRNAs, which has implications for the use of PTGS for experimental or applied purposes.
Although the Compositae harbours only two major food crops, sunflower and lettuce, many other species in this family are utilized by humans and have experienced various levels of domestication. Here, we have used next-generation sequencing technology to develop 15 reference transcriptome assemblies for Compositae crops or their wild relatives. These data allow us to gain insight into the evolutionary and genomic consequences of plant domestication. Specifically, we performed Illumina sequencing of Cichorium endivia, Cichorium intybus, Echinacea angustifolia, Iva annua, Helianthus tuberosus, Dahlia hybrida, Leontodon taraxacoides and Glebionis segetum, as well 454 sequencing of Guizotia scabra, Stevia rebaudiana, Parthenium argentatum and Smallanthus sonchifolius. Illumina reads were assembled using Trinity, and 454 reads were assembled using MIRA and CAP3. We evaluated the coverage of the transcriptomes using BLASTX analysis of a set of ultra-conserved orthologs (UCOs) and recovered most of these genes (88-98%). We found a correlation between contig length and read length for the 454 assemblies, and greater contig lengths for the 454 compared with the Illumina assemblies. This suggests that longer reads can aid in the assembly of more complete transcripts. Finally, we compared the divergence of orthologs at synonymous sites (Ks) between Compositae crops and their wild relatives and found greater divergence when the progenitors were self-incompatible. We also found greater divergence between pairs of taxa that had some evidence of postzygotic isolation. For several more distantly related congeners, such as chicory and endive, we identified a signature of introgression in the distribution of Ks values.
We have generated an ultra-high-density genetic map for lettuce, an economically important member of the Compositae, consisting of 12,842 unigenes (13,943 markers) mapped in 3696 genetic bins distributed over nine chromosomal linkage groups. Genomic DNA was hybridized to a custom Affymetrix oligonucleotide array containing 6.4 million features representing 35,628 unigenes of Lactuca spp. Segregation of single-position polymorphisms was analyzed using 213 F7:8 recombinant inbred lines that had been generated by crossing cultivated Lactuca sativa cv. Salinas and L. serriola acc. US96UC23, the wild progenitor species of L. sativa. The high level of replication of each allele in the recombinant inbred lines was exploited to identify single-position polymorphisms that were assigned to parental haplotypes. Marker information has been made available using GBrowse to facilitate access to the map. This map has been anchored to the previously published integrated map of lettuce providing candidate genes for multiple phenotypes. The high density of markers achieved in this ultradense map allowed syntenic studies between lettuce and Vitis vinifera as well as other plant species.
Several applications of high throughput genome and transcriptome sequencing would benefit from a reduction of the high-copy-number sequences in the libraries being sequenced and analyzed, particularly when applied to species with large genomes. We adapted and analyzed the consequences of a method that utilizes a thermostable duplex-specific nuclease for reducing the high-copy components in transcriptomic and genomic libraries prior to sequencing. This reduces the time, cost, and computational effort of obtaining informative transcriptomic and genomic sequence data for both fully sequenced and non-sequenced genomes. It also reduces contamination from organellar DNA in preparations of nuclear DNA. Hybridization in the presence of 3 M tetramethylammonium chloride (TMAC), which equalizes the rates of hybridization of GC and AT nucleotide pairs, reduced the bias against sequences with high GC content. Consequences of this method on the reduction of high-copy and enrichment of low-copy sequences are reported for Arabidopsis and lettuce.
The widely cultivated pepper, Capsicum spp., important as a vegetable and spice crop world-wide, is one of the most diverse crops. To enhance breeding programs, a detailed characterization of Capsicum diversity including morphological, geographical and molecular data is required. Currently, molecular data characterizing Capsicum genetic diversity is limited. The development and application of high-throughput genome-wide markers in Capsicum will facilitate more detailed molecular characterization of germplasm collections, genetic relationships, and the generation of ultra-high density maps. We have developed the Pepper GeneChip® array from Affymetrix for polymorphism detection and expression analysis in Capsicum. Probes on the array were designed from 30,815 unigenes assembled from expressed sequence tags (ESTs). Our array design provides a maximum redundancy of 13 probes per base pair position allowing integration of multiple hybridization values per position to detect single position polymorphism (SPP). Hybridization of genomic DNA from 40 diverse C. annuum lines, used in breeding and research programs, and a representative from three additional cultivated species (C. frutescens, C. chinense and C. pubescens) detected 33,401 SPP markers within 13,323 unigenes. Among the C. annuum lines, 6,426 SPPs covering 3,818 unigenes were identified. An estimated three-fold reduction in diversity was detected in non-pungent compared with pungent lines, however, we were able to detect 251 highly informative markers across these C. annuum lines. In addition, an 8.7 cM region without polymorphism was detected around Pun1 in non-pungent C. annuum. An analysis of genetic relatedness and diversity using the software Structure revealed clustering of the germplasm which was confirmed with statistical support by principle components analysis (PCA) and phylogenetic analysis. This research demonstrates the effectiveness of parallel high-throughput discovery and application of genome-wide transcript-based markers to assess genetic and genomic features among Capsicum annuum.
Species of Cactaceae are well adapted to arid habitats. Determinate growth of the primary root, which involves early and complete root apical meristem (RAM) exhaustion and differentiation of cells at the root tip, has been reported for some Cactoideae species as a root adaptation to aridity. In this study, the primary root growth patterns of Cactaceae taxa from diverse habitats are classified as being determinate or indeterminate, and the molecular mechanisms underlying RAM maintenance in Cactaceae are explored. Genes that were induced in the primary root of Stenocereus gummosus before RAM exhaustion are identified.Primary root growth was analysed in Cactaceae seedlings cultivated in vertically oriented Petri dishes. Differentially expressed transcripts were identified after reverse northern blots of clones from a suppression subtractive hybridization cDNA library.All species analysed from six tribes of the Cactoideae subfamily that inhabit arid and semi-arid regions exhibited determinate primary root growth. However, species from the Hylocereeae tribe, which inhabit mesic regions, exhibited mostly indeterminate primary root growth. Preliminary results suggest that seedlings of members of the Opuntioideae subfamily have mostly determinate primary root growth, whereas those of the Maihuenioideae and Pereskioideae subfamilies have mostly indeterminate primary root growth. Seven selected transcripts encoding homologues of heat stress transcription factor B4, histone deacetylase, fibrillarin, phosphoethanolamine methyltransferase, cytochrome P450 and gibberellin-regulated protein were upregulated in S. gummosus root tips during the initial growth phase.Primary root growth in Cactoideae species matches their environment. The data imply that determinate growth of the primary root became fixed after separation of the Cactiodeae/Opuntioideae and Maihuenioideae/Pereskioideae lineages, and that the genetic regulation of RAM maintenance and its loss in Cactaceae is orchestrated by genes involved in the regulation of gene expression, signalling, and redox and hormonal responses.
Background High-resolution genetic maps are needed in many crops to help characterize the genetic diversity that determines agriculturally important traits. Hybridization to microarrays to detect single feature polymorphisms is a powerful technique for marker discovery and genotyping because of its highly parallel nature. However, microarrays designed for gene expression analysis rarely provide sufficient gene coverage for optimal detection of nucleotide polymorphisms, which limits utility in species with low rates of polymorphism such as lettuce ( Lactuca sativa ). Results We developed a 6.5 million feature Affymetrix GeneChip® for efficient polymorphism discovery and genotyping, as well as for analysis of gene expression in lettuce. Probes on the microarray were designed from 26,809 unigenes from cultivated lettuce and an additional 8,819 unigenes from four related species ( L. serriola , L. saligna , L. virosa and L. perennis ). Where possible, probes were tiled with a 2 bp stagger, alternating on each DNA strand; providing an average of 187 probes covering approximately 600 bp for each of over 35,000 unigenes; resulting in up to 13 fold redundancy in coverage per nucleotide. We developed protocols for hybridization of genomic DNA to the GeneChip® and refined custom algorithms that utilized coverage from multiple, high quality probes to detect single position polymorphisms in 2 bp sliding windows across each unigene. This allowed us to detect greater than 18,000 polymorphisms between the parental lines of our core mapping population, as well as numerous polymorphisms between cultivated lettuce and wild species in the lettuce genepool. Using marker data from our diversity panel comprised of 52 accessions from the five species listed above, we were able to separate accessions by species using both phylogenetic and principal component analyses. Additionally, we estimated the diversity between different types of cultivated lettuce and distinguished morphological types. Conclusion By hybridizing genomic DNA to a custom oligonucleotide array designed for maximum gene coverage, we were able to identify polymorphisms using two approaches for pair-wise comparisons, as well as a highly parallel method that compared all 52 genotypes simultaneously.