The basidiomycete Moniliophthora roreri causes frosty pod rot of cacao (Theobroma cacao) in the western hemisphere. Moniliophthora roreri is considered asexual and haploid throughout its hemibiotrophic life cycle. To understand the processes driving genome modification, using long-read sequencing technology, we sequenced and assembled 5 high-quality M. roreri genomes out of a collection of 99 isolates collected throughout the pathogen's range. We obtained chromosome-scale assemblies composed of 11 scaffolds. We used short-read technology to sequence the genomes of 22 similarly chosen isolates. Alignments among the 5 reference assemblies revealed inversions, translocations, and duplications between and within scaffolds. Isolates at the front of the pathogens' expanding range tend to share lineage-specific structural variants, as confirmed by short-read sequencing. We identified, for the first time, 3 new mating type A locus alleles (5 in total) and 1 new potential mating type B locus allele (3 in total). Currently, only 2 mating type combinations, A1B1 and A2B2, are known to exist outside of Colombia. A systematic survey of the M. roreri transcriptome across 2 isolates identified an expanded candidate effector pool and provided evidence that effector candidate genes unique to the Moniliophthoras are preferentially expressed during the biotrophic phase of disease. Notably, M. roreri isolates in Costa Rica carry a chromosome segment duplication that has doubled the associated gene complement and includes secreted proteins and candidate effectors. Clonal reproduction of the haploid M. roreri genome has allowed lineages with unique genome structures and compositions to dominate as it expands its range, displaying a significant founder effect.
Cacao (Theobroma cacao) is a tropical tree that produces the essential raw material for chocolate. Because yields have been stagnant, land use has expanded to provide for increasing chocolate demand. Assembled genomes of key parents could modernize breeding programs in the remote and under-resourced locations where cacao is grown. The MinION, a long read sequencer that runs off of a laptop computer, has the potential to facilitate the assembly of the complex genomes of high-yielding F-1 hybrids. Here, we validate the MinION's application to heterozygous crops by creating a de novo genome assembly of a key parent in breeding programs, the clone Pound 7. Our MinION-only assembly was 20% larger than the latest released cacao genome, with 10-fold greater contiguity, and the resolution of complex heterozygosity and repetitive elements. Polishing with Illumina short reads brought the predicted completeness of our assembly to similar levels to the previously released cacao genome assemblies. In contrast to previous cacao genome projects, our assembly required only a small scientific team and limited reagents. Our sequencing and assembly methods could easily be adopted by under-resourced breeding programs, speeding crop improvement in the developing world.
The main ingredients of chocolate are usually cocoa powder, cocoa butter, and sugar. Both the powder and the butter are extracted from the beans of the cacao tree (Theobroma cacao L.). The cocoa butter represents the fat in the beans and possesses a unique fatty acid profile that results in chocolate's characteristic texture and mouthfeel. Here, we used a linkage mapping population and phenotypic data of 3,292 samples from 420 progeny which led to the identification of 27 quantitative trait loci (QTLs) for fatty acid composition and six QTLs for fat content. Progeny showed extensive variation in fat levels and composition, with the level of palmitic acid negatively correlated to the sum of stearic acid, oleic acid, and linoleic acid. A major QTL explaining 24% of the relative level of palmitic acid was mapped to the distal end of chromosome 4, and those higher levels of palmitic acid were associated with the presence of a haplotype from the "TSH 1188" parent in the progeny. Within this region of chromosome 4 is the Thecc1EG017405 gene, an orthologue and isoform of the stearoyl-acyl carrier protein (ACP) desaturase (SAD) gene in plants, which is involved in fatty acid biosynthesis. Besides allelic differences, we also show that climate factors can change the fatty acid composition in the beans, including a significant positive correlation between higher temperatures and the higher level of palmitic acid. Moreover, we found a significant pollen donor effect from the variety "SIAL 70" which was associated with decreased palmitic acid levels.
Domestication has had a strong impact on the development of modern societies. We sequenced 200 genomes of the chocolate plant Theobroma cacao L. to show for the first time to our knowledge that a single population, the Criollo population, underwent strong domestication ~3600 years ago (95% CI: 2481–13,806 years ago). We also show that during the process of domestication, there was strong selection for genes involved in the metabolism of the colored protectants anthocyanins and the stimulant theobromine, as well as disease resistance genes. Our analyses show that domesticated populations of T. cacao (Criollo) maintain a higher proportion of high-frequency deleterious mutations. We also show for the first time the negative consequences of the increased accumulation of deleterious mutations during domestication on the fitness of individuals (significant reduction in kilograms of beans per hectare per year as Criollo ancestry increases, as estimated from a GLM, P = 0.000425).
Cacao (Theobroma cacao) is a globally important crop, and its yield is severely restricted by disease. Two of the most damaging diseases, witches' broom disease (WBD) and frosty pod rot disease (FPRD), are caused by a pair of related fungi: Moniliophthora perniciosa and Moniliophthora roreri, respectively. Resistant cultivars are the most effective long-term strategy to address Moniliophthora diseases, but efficiently generating resistant and productive new cultivars will require robust methods for screening germplasm before field testing. Marker-assisted selection (MAS) and genomic selection (GS) provide two potential avenues for predicting the performance of new genotypes, potentially increasing the selection gain per unit time. To test the effectiveness of these two approaches, we performed a genome-wide association study (GWAS) and GS on three related populations of cacao in Ecuador genotyped with a 15K single nucleotide polymorphism (SNP) microarray for three measures of WBD infection (vegetative broom, cushion broom, and chirimoya pod), one of FPRD (monilia pod) and two productivity traits (total fresh weight of pods and % healthy pods produced). GWAS yielded several SNPs associated with disease resistance in each population, but none were significantly correlated with the same trait in other populations. Genomic selection, using one population as a training set to estimate the phenotypes of the remaining two (composed of different families), varied among traits, from a mean prediction accuracy of 0.46 (vegetative broom) to 0.15 (monilia pod), and varied between training populations. Simulations demonstrated that selecting seedlings using GWAS markers alone generates no improvement over selecting at random, but that GS improves the selection process significantly. Our results suggest that the GWAS markers discovered here are not sufficiently predictive across diverse germplasm to be useful for MAS, but that using all markers in a GS framework holds substantial promise in accelerating disease-resistance in cacao.
Domestication has had a strong impact on the development of modern societies. We sequenced 200 genomes of the chocolate plant Theobroma cacao L. to show for the first time that a single population underwent strong domestication approximately 3,600 years (95% CI: 2481 – 10,903 years ago) ago, the Criollo population. We also show that during the process of domestication, there was strong selection for genes involved in the metabolism of the colored protectants anthocyanins and the stimulant theobromine, as well as disease resistance genes. Our analyses show that domesticated populations of T. cacao (Criollo) maintain a higher proportion of high frequency deleterious mutations. We also show for the first time the negative consequences the increase accumulation of deleterious mutations during domestication on the fitness of individuals (significant negative correlation between Criollo ancestry and Kg of beans per hectare per year, P = 0.000425).
Breeding programs of cacao (Theobroma cacao L.) trees share the many challenges of breeding long-living perennial crops, and genetic progress is further constrained by both the limited understanding of the inheritance of complex traits and the prevalence of technical issues, such as mislabeled individuals (off-types). To better understand the genetic architecture of cacao, in this study, 13 years of phenotypic data collected from four progeny trials in Bahia, Brazil were analyzed jointly in a multisite analysis. Three separate analyses (multisite, single site with and without off-types) were performed to estimate genetic parameters from statistical models fitted on nine important agronomic traits (yield, seed index, pod index, % healthy pods, % pods infected with witches broom, % of pods other loss, vegetative brooms, diameter, and tree height). Genetic parameters were estimated along with variance components and heritabilities from the multisite analysis, and a trial was fingerprinted with low-density SNP markers to determine the impact of off-types on estimations. Heritabilities ranged from 0.37 to 0.64 for yield and its components and from 0.03 to 0.16 for disease resistance traits. A weighted index was used to make selections for clonal evaluation, and breeding values estimated for the parental selection and estimation of genetic gain. The impact of off-types to breeding progress in cacao was assessed for the first time. Even when present at <5% of the total population, off-types altered selections by 48%, and impacted heritability estimations for all nine of the traits analyzed, including a 41% difference in estimated heritability for yield. These results show that in a mixed model analysis, even a low level of pedigree error can significantly alter estimations of genetic parameters and selections in a breeding program.
Cacao (Theobroma cacao L.) is an important cash crop in tropical regions around the world and has a rich agronomic history in South America. As a key component in the cosmetic and confectionary industries, millions of people worldwide use products made from cacao, ranging from shampoo to chocolate. An Illumina Infinity II array was created using 13,530 SNPs identified within a small diversity panel of cacao. Of these SNPs, 12,643 derive from variation within annotated cacao genes. The genotypes of 3,072 trees were obtained, including two mapping populations from Ecuador. High-density linkage maps for these two populations were generated and compared to the cacao genome assembly. Phenotypic data from these populations were combined with the linkage maps to identify the QTLs for yield and disease resistance.
Linkage disequilibrium (LD) measured over the genomes of a species can provide important indications for how future association analyses should proceed. This information can be advantageous especially for slow-growing, perennial crops such as Theobroma cacao, where experimental crosses are inherently time-consuming and logistically expensive. While LD has been evaluated in cacao, previous work has been focused on relatively narrow genetic bases. We use microsatellite marker data collected from a uniquely diverse sample of individuals broadly covering both wild and cultivated varieties to gauge the LD present in the different cacao diversity groups and populations. We find that genome-wide LD decays far more rapidly in the wild and primitive diversity groups of cacao as compared to those representing cultivated varieties. The impact that such differences can have on association analyses is demonstrated using phenotypic data on pod color and genotypic data from two cacao populations with contrasting patterns of LD decay. Our results indicate that the more rapid LD decay in wild and primitive germplasm can lead to higher-resolution mapping intervals when compared to results from cultivated germplasm. Through simulations, we demonstrate how future association mapping analyses, comprising of cacao samples with a wild or primitive background, will likely exhibit lower LD and would be more suitable for fine-scale association mapping analyses. As many traits targeted by cacao breeders are found exclusively in wild and primitive germplasm, association mapping in wild cacao populations holds significant promise for cacao improvement through marker-assisted breeding and emphasize the need to further explore the natural diversity of Amazonian cacao.
Theobroma cacao, the key ingredient in chocolate production, is one of the world's most important tree fruit crops, with ∼4,000,000 metric tons produced across 50 countries. To move towards gene discovery and marker-assisted breeding in cacao, a single-nucleotide polymorphism (SNP) identification project was undertaken using RNAseq data from 16 diverse cacao cultivars. RNA sequences were aligned to the assembled transcriptome of the cultivar Matina 1-6, and 330,000 SNPs within coding regions were identified. From these SNPs, a subset of 6,000 high-quality SNPs were selected for inclusion on an Illumina Infinium SNP array: the Cacao6kSNP array. Using Cacao6KSNP array data from over 1,000 cacao samples, we demonstrate that our custom array produces a saturated genetic map and can be used to distinguish among even closely related genotypes. Our study enhances and expands the genetic resources available to the cacao research community, and provides the genome-scale set of tools that are critical for advancing breeding with molecular markers in an agricultural species with high genetic diversity.
Background Theobroma cacao L. cultivar Matina 1-6 belongs to the most cultivated cacao type. The availability of its genome sequence and methods for identifying genes responsible for important cacao traits will aid cacao researchers and breeders. Results We describe the sequencing and assembly of the genome of Theobroma cacao L. cultivar Matina 1-6. The genome of the Matina 1-6 cultivar is 445 Mbp, which is significantly larger than a sequenced Criollo cultivar, and more typical of other cultivars. The chromosome-scale assembly, version 1.1, contains 711 scaffolds covering 346.0 Mbp, with a contig N50 of 84.4 kbp, a scaffold N50 of 34.4 Mbp, and an evidence-based gene set of 29,408 loci. Version 1.1 has 10x the scaffold N50 and 4x the contig N50 as Criollo, and includes 111 Mb more anchored sequence. The version 1.1 assembly has 4.4% gap sequence, while Criollo has 10.9%. Through a combination of haplotype, association mapping and gene expression analyses, we leverage this robust reference genome to identify a promising candidate gene responsible for pod color variation. We demonstrate that green/red pod color in cacao is likely regulated by the R2R3 MYB transcription factor TcMYB113 , homologs of which determine pigmentation in Rosaceae, Solanaceae, and Brassicaceae. One SNP within the target site for a highly conserved trans -acting siRNA in dicots, found within TcMYB113 , seems to affect transcript levels of this gene and therefore pod color variation. Conclusions We report a high-quality sequence and annotation of Theobroma cacao L. and demonstrate its utility in identifying candidate genes regulating traits.
Influenza A viruses are characterized by their ability to evade host immunity, even in vaccinated individuals. To determine how prior immunity shapes viral diversity in vivo, we studied the intra- and interhost evolution of equine influenza virus in vaccinated horses. Although the level and structure of genetic diversity were similar to those in naïve horses, intrahost bottlenecks may be more stringent in vaccinated animals, and mutations shared among horses often fall close to putative antigenic sites.
The genetic diversity present in populations of RNA viruses is likely to be strongly modulated by aspects of their life history, including mode of transmission. However, how transmission mode shapes patterns of intra- and inter-host genetic diversity, particularly when acting in combination with de novo mutation, population bottlenecks and the selection of advantageous mutations, is poorly understood. To address these issues, this study performed ultradeep sequencing of zucchini yellow mosaic virus in a wild gourd, Cucurbita pepo ssp. texana, under two infection conditions: aphid vectored and mechanically inoculated, achieving a mean coverage of approximately 10 ,000×. It was shown that mutations persisted during inter-host transmission events in both the aphid vectored and mechanically inoculated populations, suggesting that the vector-imposed transmission bottleneck is not as extreme as previously supposed. Similarly, mutations were found to persist within individual hosts, arguing against strong systemic bottlenecks. Strikingly, mutations were seen to go to fixation in the aphid-vectored plants, suggestive of a major fitness advantage, but remained at low frequency in the mechanically inoculated plants. Overall, this study highlights the utility of ultradeep sequencing in providing high-resolution data capable of revealing the nature of virus evolution, particularly as the full spectrum of genetic diversity within a population may not be uncovered without sequence coverage of at least 2500-fold.
Influenza A viruses (IAVs) cause acute, highly transmissible infections in a wide range of animal species. Understanding how these viruses are transmitted within and between susceptible host populations is critical to the development of effective control strategies. While viral gene sequences have been used to make inferences about IAV transmission dynamics at the epidemiological scale, their utility in accurately determining patterns of inter-host transmission in the short-term--i.e. who infected whom--has not been strongly established. Herein, we use intra-host sequence data from the viral HA1 (hemagglutinin) gene domain from two transmission studies employing different IAV subtypes in their natural hosts--H3N8 in horses and H1N1 in pigs-to determine how well these data recapitulate the known pattern of inter-host transmission. Although no mutations were fixed over the course of either experimental transmission chain, we show that some minor, transient alleles can provide evidence of host-to-host transmission and, importantly, can be distinguished from those that cannot.
Models of infectious disease spread that incorporate contact heterogeneity through contact networks are an important tool for epidemiologists studying disease dynamics and assessing intervention strategies. One of the challenges of contact network epidemiology has been the difficulty of collecting individual and population-level data needed to develop an accurate representation of the underlying host population's contact structure. In this study, we evaluate the utility of common epidemiological measures (R0, epidemic peak size, duration and final size) for inferring the degree of heterogeneity in a population's unobserved contact structure through a Bayesian approach. We test the method using ground truth data and find that some of these epidemiological metrics are effective at classifying contact heterogeneity. The classification is also consistent across pathogen transmission probabilities, and so can be applied even when this characteristic is unknown. In particular, the reproductive number, R0, turns out to be a poor classifier of the degree heterogeneity, while, unexpectedly, final epidemic size is a powerful predictor of network structure across the range of heterogeneity. We also evaluate our framework on empirical epidemiological data from past and recent outbreaks to demonstrate its application in practice and to gather insights about the relevance of particular contact structures for both specific systems and general classes of infectious disease. We thus introduce a simple approach that can shed light on the unobserved connectivity of a host population given epidemic data. Our study has the potential to inform future data-collection efforts and study design by driving our understanding of germane epidemic measures, and highlights a general inferential approach to learning about host contact structure in contemporary or historic populations of humans and animals.
Summary1. Maximum likelihood analyses for testing hypotheses about how rates of disparification might vary across clades can provide important insight into the evolutionary process. While the Brownie phylogenetic library can perform such analyses, it does so outside of a general scripting environment.2. We present RBrownie, an interface between the Brownie phylogenetic library and the R software environment, which provides easy access to the main methods in Brownie (see O'Meara 2008; PhD Dissertation, Nature Precedings), including discrete ancestral state reconstruction. In addition, RBrownie supplies a direct interface to Brownie, allowing advanced users to construct more complex combinations of analyses and to execute any newly added Brownie functions.3. Overall, it is a package that features evolutionary rate analyses in a flexible and familiar environment.
With more emphasis being put on global infectious disease monitoring, viral genetic data are being collected at an astounding rate, both within and without the context of a long-term disease surveillance plan. Concurrent with this increase have come improvements to the sophisticated and generalized statistical techniques used for extracting population-level information from genetic sequence data. However, little research has been done on how the collection of these viral sequence data can or does affect the efficacy of the phylogenetic algorithms used to analyse and interpret them. In this study, we use epidemic simulations to consider how the collection of viral sequence data clarifies or distorts the picture, provided by the phylogenetic algorithms, of the underlying population dynamics of the simulated viral infection over many epidemic cycles. We find that sampling protocols purposefully designed to capture sequences at specific points in the epidemic cycle, such as is done for seasonal influenza surveillance, lead to a significantly better view of the underlying population dynamics than do less-focused collection protocols. Our results suggest that the temporal distribution of samples can have a significant effect on what can be inferred from genetic data, and thus highlight the importance of considering this distribution when designing or evaluating protocols and analysing the data collected thereunder.