The classification of Nesocaryum stylosum (Boraginaceae) has remained unresolved for nearly a century. This species is endemic to Isla San Ambrosio, a small island located approximately 900 km due west of the coast of central Chile. Ivan M. Johnston transferred the species from the genus Heliotropium to the monotypic Nesocaryum in 1927 but noted that, despite its quite different and rather unique vegetative, inflorescence, and calyx morphology, its unit fruits (nutlets/eremocarps) resemble those of Cryptantha. Here, we review the morphology and taxonomic history of N. stylosum and provide DNA sequence data that support its placement in Cryptantha, for which we propose the new combination Cryptantha stylosa. Our data also support the placement of this species within the Maritimae clade, a monophyletic group of Cryptantha species that is phylogenetically distinct from the bulk of the genus. We propose an expanded membership of the Maritimae clade comprising up to 19 species: eight species (13 minimum-rank taxa) from North America and 12 species (12 minimum-rank taxa) from South America, including Cryptantha stylosa, with one taxon occurring on both continents. We further review evidence bearing on the biogeographic and evolutionary history of Cryptantha stylosa and its putative closest relatives and identify the need for additional research within the group.
Polyploidy, also known as whole-genome duplication (WGD), is a significant evolutionary force in green plants, especially angiosperms. The dynamic nature of polyploid genomes generates genetic diversity and drives the evolution of novel traits and adaptations. Pangenomics is emerging as a major frontier in plant genome research, with a rapidly growing number of pangenomes for individual species and associated analyses providing novel agronomic and evolutionary insights. Polyploid genome analysis can be confounded by intraspecific variation when relying on a single reference genome assembly. The use of pangenomes that better represent the genomic diversity of a species helps overcome this limitation. However, a major gap remains between the number of pangenomic studies in polyploid compared to diploid species, despite the widespread prevalence of WGD, limiting the potential of the pangenome framework for characterizing and understanding polyploid genomes. Furthermore, most polyploid pangenome studies have focused on domesticated crop species, and natural populations have rarely been examined. In addition to applications in crop improvement, pangenomes can provide insights into the ecological and evolutionary impact of polyploidy. Here, we summarize recent pangenome studies in polyploid plants and highlight promising topics for future research. We hope this article will encourage the growth of pangenomic studies in polyploid systems, particularly in natural populations.
Polyploidy is an important generator of evolutionary novelty across diverse groups in the Tree of Life, including many crops. However, the impact of whole-genome duplication depends on the mode of formation: doubling within a single lineage (autopolyploidy) versus doubling after hybridization between two different lineages (allopolyploidy). Researchers have historically treated these two scenarios as completely separate cases based on patterns of chromosome pairing, but these cases represent ideals on a continuum of chromosomal interactions among duplicated genomes. Understanding the history of polyploid species thus demands quantitative inferences of demographic history and rates of exchange between subgenomes. To meet this need, we developed diffusion models for genetic variation in polyploids with subgenomes that cannot be bioinformatically separated and with potentially variable inheritance patterns, implementing them in the dadi software. We validated our models using forward SLiM simulations and found that our inference approach is able to accurately infer evolutionary parameters (timing, bottleneck size) involved with the formation of auto- and allotetraploids, as well as exchange rates in segmental allotetraploids. We then applied our models to empirical data for allotetraploid shepherd's purse (Capsella bursa-pastoris), finding evidence for allelic exchange between the subgenomes. Taken together, our model provides a foundation for demographic modeling in polyploids using diffusion equations, which will help increase our understanding of the impact of demography and selection in polyploid lineages.
Most land plants are now known to be ancient polyploids that have rediploidized. Diploidization involves many changes in genome organization that ultimately restore bivalent chromosome pairing and disomic inheritance, and resolve dosage and other issues caused by genome duplication. In this review, we discuss the nature of polyploidy and its impact on chromosome pairing behavior. We also provide an overview of two major and largely independent processes of diploidization: cytological diploidization and genic diploidization/fractionation. Finally, we compare variation in gene fractionation across land plants and highlight the differences in diploidization between plants and animals. Altogether, we demonstrate recent advancements in our understanding of variation in the patterns and processes of diploidization in land plants and provide a road map for future research to unlock the mysteries of diploidization and eukaryotic genome evolution.
Inferring the frequency and mode of hybridization among closely related organisms is an important step for understanding the process of speciation and can help to uncover reticulated patterns of phylogeny more generally. Phylogenomic methods to test for the presence of hybridization come in many varieties and typically operate by leveraging expected patterns of genealogical discordance in the absence of hybridization. An important assumption made by these tests is that the data (genes or SNPs) are independent given the species tree. However, when the data are closely linked, it is especially important to consider their nonindependence. Recently, deep learning techniques such as convolutional neural networks (CNNs) have been used to perform population genetic inferences with linked SNPs coded as binary images. Here, we use CNNs for selecting among candidate hybridization scenarios using the tree topology (((P1 , P2 ), P3 ), Out) and a matrix of pairwise nucleotide divergence (dXY ) calculated in windows across the genome. Using coalescent simulations to train and independently test a neural network showed that our method, HyDe-CNN, was able to accurately perform model selection for hybridization scenarios across a wide breath of parameter space. We then used HyDe-CNN to test models of admixture in Heliconius butterflies, as well as comparing it to phylogeny-based introgression statistics. Given the flexibility of our approach, the dropping cost of long-read sequencing and the continued improvement of CNN architectures, we anticipate that inferences of hybridization using deep learning methods like ours will help researchers to better understand patterns of admixture in their study organisms.
Penstemon (Plantaginaceae), the largest genus of plants native to North America, represents a recent continental evolutionary radiation. We investigated patterns of diversification, phylogenetic relationships, and biogeography, and determined the age of the lineage using 43 nuclear gene loci. We also assessed the current taxonomic circumscription of the ca. 285 species by developing a phylogenetic taxonomic bootstrap method. Penstemon originated during the Pliocene/Pleistocene transition. Patterns of diversification and biogeography are associated with glaciation cycles during the Pleistocene, with the bulk of diversification occurring from 1.0–0.5 mya. The radiation across the North American continent tracks the advance and retreat of major and minor glaciation cycles during the past 2.5 million years with founder-event speciation contributing the most to diversification of Penstemon . Our taxonomic bootstrap analyses suggest the current circumscription of the genus is in need of revision. We propose rearrangement of subgenera, sections, and subsections based on our phylogenetic results. Given the young age and broad distribution of Penstemon across North America, it offers an excellent system for studying a rapid evolutionary radiation in a continental setting.
Summary Many crops are polyploid or have a polyploid ancestry. Recent phylogenetic analyses have found that polyploidy often preceded the domestication of crop plants. One explanation for this observation is that increased genetic diversity following polyploidy may have been important during the strong artificial selection that occurs during domestication. In order to test the connection between domestication and polyploidy, we identified and examined candidate genes associated with the domestication of the diverse crop varieties of Brassica rapa. Like all ‘diploid’ flowering plants, B. rapa has a diploidized paleopolyploid genome and experienced many rounds of whole genome duplication (WGD). We analyzed transcriptome data of more than 100 cultivated B. rapa accessions. Using a combination of approaches, we identified > 3000 candidate genes associated with the domestication of four major B. rapa crop varieties. Consistent with our expectation, we found that the candidate genes were significantly enriched with genes derived from the Brassiceae mesohexaploidy. We also observed that paleologs were significantly more diverse than non‐paleologs. Our analyses find evidence for that genetic diversity derived from ancient polyploidy played a key role in the domestication of B. rapa and provide support for its importance in the success of modern agriculture.
Despite early domestication around 3000 BC, the evolutionary history of the ancient allotetraploid species Brassica juncea (L.) Czern & Coss remains uncertain. Here, we report a chromosome-scale de novo assembly of a yellow-seeded B. juncea genome by integrating long-read and short-read sequencing, optical mapping and Hi-C technologies. Nuclear and organelle phylogenies of 480 accessions worldwide supported that B. juncea is most likely a single origin in West Asia, 8,000–14,000 years ago, via natural interspecific hybridization. Subsequently, new crop types evolved through spontaneous gene mutations and introgressions along three independent routes of eastward expansion. Selective sweeps, genome-wide trait associations and tissue-specific RNA-sequencing analysis shed light on the domestication history of flowering time and seed weight, and on human selection for morphological diversification in this versatile species. Our data provide a comprehensive insight into the origin and domestication and a foundation for genomics-based breeding of B. juncea .
Most land plants are now known to be ancient polyploids that have rediploidized. This process of diploidization involves many changes in genome organization that ultimately restores bivalent chromosome pairing, disomic inheritance, and resolves dosage and other issues caused by genome duplication. Here, we provide an overview of the variety of mechanisms involved in diploidization as well as new analyses of pairing behavior and variation in gene fractionation across land plants. Overall, we find that lineage and WGD specific attributes influence the evolutionary outcomes of WGD and the process of diploidization in plant genomes. Ultimately, many of the mechanisms and forces driving diploidization remain to be discovered. Future research that leverages variation in the patterns and processes of diploidization will be able to advance our understanding of plant genome evolution and unlock the mysteries of diploidization. INTRODUCTION A major insight from two decades of sequencing plant genomes is that most are not simply diploid, but diploidized paleopolyploid genomes. Although it has long been
Demographic inference using the site frequency spectrum (SFS) is a common way to understand historical events affecting genetic variation. However, most methods for estimating demography from the SFS assume random mating within populations, precluding these types of analyses in inbred populations. To address this issue, we developed a model for the expected SFS that includes inbreeding by parameterizing individual genotypes using beta-binomial distributions. We then take the convolution of these genotype probabilities to calculate the expected frequency of biallelic variants in the population. Using simulations, we evaluated the model’s ability to co-estimate demography and inbreeding using one- and two-population models across a range of inbreeding levels. We also applied our method to two empirical examples, American pumas ( Puma concolor ) and domesticated cabbage ( Brassica oleracea var. capitata ), inferring models both with and without inbreeding to compare parameter estimates and model fit. Our simulations showed that we are able to accurately co-estimate demographic parameters and inbreeding even for highly inbred populations ( F = 0.9). In contrast, failing to include inbreeding generally resulted in inaccurate parameter estimates in simulated data and led to poor model fit in our empirical analyses. These results show that inbreeding can have a strong effect on demographic inference, a pattern that was especially noticeable for parameters involving changes in population size. Given the importance of these estimates for informing practices in conservation, agriculture, and elsewhere, our method provides an important advancement for accurately estimating the demographic histories of these species.
Reticulate evolutionary events are hallmarks of plant phylogeny, and are increasingly recognized as common occurrences in other branches of the Tree of Life. However, inferring the evolutionary history of admixed lineages presents a difficult challenge for systematists due to genealogical discordance caused by both incomplete lineage sorting (ILS) and hybridization. Methods that accommodate both of these processes are continuing to be developed, but they often do not scale well to larger numbers of species. An additional complicating factor for many plant species is the occurrence of whole genome duplication (WGD), which can have various outcomes on the genealogical history of haplotypes sampled from the genome. In this study, we sought to investigate patterns of hybridization and WGD in two subsections from the genus Penstemon (Plantaginaceae; subsect. Humiles and Proceri ), a speciose group of angiosperms that has rapidly radiated across North America. Species in subsect. Humiles and Proceri occur primarily in the Pacific Northwest of the United States, occupying habitats such as mesic, subalpine meadows, as well as more well-drained substrates at varying elevations. Ploidy levels in the subsections range from diploid to hexaploid, and it is hypothesized that most of the polyploids are hybrids (i.e., allopolyploids). To estimate phylogeny in these groups, we first developed a method for estimating quartet concordance factors (QCFs) from multiple sequences sampled per lineage, allowing us to model all haplotypes from a polyploid. QCFs represent the proportion of gene trees that support a particular species quartet relationship, and are used for species network estimation in the program SNaQ ([Solís-Lemus & Ané. 2016][1]. PLoS Genet. 12:e1005896). Using phased haplotypes for nuclear amplicons, we inferred species trees and networks for 38 taxa from P . subsect. Humiles and Proceri . Our phylogenetic analyses recovered two clades comprising a mix of taxa from both subsections, indicating that the current taxonomy for these groups is inconsistent with our estimates of phylogeny. In addition, there was little support for hypotheses regarding the formation of putative allopolyploid lineages. Overall, we found evidence for the effects of both ILS and admixture on the evolutionary history of these species, but were able to evaluate our taxonomic hypotheses despite high levels of gene tree discordance. Our method for estimating QCFs from multiple haplotypes also allowed us to include species of varying ploidy levels in our analyses, which we anticipate will help to facilitate estimation of species networks in other plant groups as well.### Competing Interest StatementThe authors have declared no competing interest. [1]: #ref-57
Whole-genome duplications (WGDs) are prevalent throughout the evolutionary history of plants. For example, dozens of WGDs have been phylogenetically localized across the order Brassicales, specifically, within the family Brassicaceae. However, while its sister family, Cleomaceae, has also been characterized by a WGD, its placement, as well as that of other WGD events in other families in the order, remains unclear. Using phylo-transcriptomics from 74 taxa and genome survey sequencing for 66 of those taxa, we infer nuclear and chloroplast phylogenies to assess relationships among the major families of the Brassicales and within the Brassicaceae. We then use multiple methods of WGD inference to assess placement of WGD events. We not only present well-supported chloroplast and nuclear phylogenies for the Brassicales, but we also putatively place Th-α and provide evidence for previously unknown events, including one shared by at least two members of the Resedaceae, which we name Rs-α. Given its economic importance and many genomic resources, the Brassicales are an ideal group to continue assessing WGD inference methods. We add to the current conversation on WGD inference difficulties, by demonstrating that sampling is especially important for WGD identification.
PremiseEnvironmentally controlled facilities, such as growth chambers, are essential tools for experimental research. Automated, low‐cost, remote‐monitoring hardware can greatly improve both reproducibility and maintenance.Methods and ResultsUsing a Raspberry Pi computer, open‐source software, environmental sensors, and a camera, we developed Growth Monitor pi (GMpi), a cost‐effective system for monitoring growth chamber conditions. Coupled with our software, GMPi_Pack, our setup automates sensor readings, photography, and alerts when conditions fall out of range.ConclusionsGMpi offers access to environmental data logging, improving reproducibility of experiments and reinforcing the stability of controlled environmental facilities. The device is also flexible and scalable, allowing researchers the ability to customize and expand GMpi for their own needs.
Sympatric diversification is increasingly thought to have played an important role in the evolution of biodiversity around the globe. However, an in situ sympatric origin for co-distributed taxa is difficult to demonstrate empirically because different evolutionary processes can lead to similar biogeographic outcomes-especially in ecosystems with few hard barriers to dispersal that can facilitate allopatric speciation followed by secondary contact (e.g. marine habitats). Here we use a genomic (ddRADseq), model-based approach to delimit a cryptic species complex of tropical sea anemones that are co-distributed on coral reefs throughout the Tropical Western Atlantic. We use coalescent simulations in fastsimcoal2 to test competing diversification scenarios that span the allopatric-sympatric continuum. We recover support that the corkscrew sea anemone Bartholomea annulata (Le Sueur, 1817) is a cryptic species complex, co-distributed throughout its range. Simulation and model selection analyses suggest these lineages arose in the face of historical and contemporary gene flow, supporting a sympatric origin, but an alternative secondary contact model also receives appreciable model support. Leveraging the genome of Exaiptasia pallida we identify five loci under divergent selection between cryptic B. annulata lineages that fall within mRNA transcripts or CDS regions. Our study provides a rare empirical, genomic example of sympatric speciation in a tropical anthozoan-a group that includes reef-building corals. Finally, these data represent the first range-wide molecular study of any tropical sea anemone, underscoring that anemone diversity is under described in the tropics, and highlighting the need for additional systematic studies into these ecologically and economically important species.
Motivation Genotyping and parameter estimation using high throughput sequencing data are everyday tasks for population geneticists, but methods developed for diploids are typically not applicable to polyploid taxa. This is due to their duplicated chromosomes, as well as the complex patterns of allelic exchange that often accompany whole genome duplication (WGD) events. For WGDs within a single lineage (autopolyploids), inbreeding can result from mixed mating and/or double reduction. For WGDs that involve hybridization (allopolyploids), alleles are typically inherited through independently segregating subgenomes. Results We present two new models for estimating genotypes and population genetic parameters from genotype likelihoods for auto‐ and allopolyploids. We then use simulations to compare these models to existing approaches at varying depths of sequencing coverage and ploidy levels. These simulations show that our models typically have lower levels of estimation error for genotype and parameter estimates, especially when sequencing coverage is low. Finally, we also apply these models to two empirical datasets from the literature. Overall, we show that the use of genotype likelihoods to model non‐standard inheritance patterns is a promising approach for conducting population genomic inferences in polyploids. Availability and implementation A C ++ program, EBG, is provided to perform inference using the models we describe. It is available under the GNU GPLv3 on GitHub: https://github.com/pblischak/polyploid‐genotyping.
The analysis of hybridization and gene flow among closely related taxa is a common goal for researchers studying speciation and phylogeography. Many methods for hybridization detection use simple site pattern frequencies from observed genomic data and compare them to null models that predict an absence of gene flow. The theory underlying the detection of hybridization using these site pattern probabilities exploits the relationship between the coalescent process for gene trees within population trees and the process of mutation along the branches of the gene trees. For certain models, site patterns are predicted to occur in equal frequency (i.e., their difference is 0), producing a set of functions called phylogenetic invariants. In this article, we introduce HyDe, a software package for detecting hybridization using phylogenetic invariants arising under the coalescent model with hybridization. HyDe is written in Python and can be used interactively or through the command line using pre-packaged scripts. We demonstrate the use of HyDe on simulated data, as well as on two empirical data sets from the literature. We focus in particular on identifying individual hybrids within population samples and on distinguishing between hybrid speciation and gene flow. HyDe is freely available as an open source Python package under the GNU GPL v3 on both GitHub (https://github.com/pblischak/HyDe) and the Python Package Index (PyPI: https://pypi.python.org/pypi/phyde).
PREMISE OF THE STUDY:Targeted enrichment strategies for phylogenomic inference are a time- and cost-efficient way to collect DNA sequence data for large numbers of individuals at multiple, independent loci. Automated and reproducible processing of these data is a crucial step for researchers conducting phylogenetic studies.METHODS AND RESULTS:We present Fluidigm2PURC, an open source Python utility for processing paired-end Illumina data from double-barcoded PCR amplicons. In combination with the program PURC (Pipeline for Untangling Reticulate Complexes), our scripts process raw FASTQ files for analysis with PURC and use its output to infer haplotypes for diploids, polyploids, and samples with unknown ploidy. We demonstrate the use of the pipeline with an example data set from the genus Thalictrum (Ranunculaceae).CONCLUSIONS:Fluidigm2PURC is freely available for Unix-like operating systems on GitHub (https://github.com/pblischak/fluidigm2purc) and for all operating systems through Docker (https://hub.docker.com/r/pblischak/fluidigm2purc).
Duplication events are regarded as sources of evolutionary novelty, but our understanding of general trends for the long-term trajectory of additional genomic material is still lacking. Organisms with a history of whole genome duplication (WGD) offer a unique opportunity to study potential trends in the context of gene retention and/or loss, gene and network dosage, and changes in gene expression. In this review, we discuss the prevalence of polyploidy across the tree of life, followed by an overview of studies investigating genome evolution and gene expression. We then provide an overview of methods in network biology, phylogenomics, and population genomics that are critical for advancing our understanding of evolution post-WGD, highlighting the need for models that can accommodate polyploids. Finally, we close with a brief note on the importance of random processes in the evolution of polyploids with respect to neutral versus selective forces, ancestral polymorphisms, and the formation of autopolyploids versus allopolyploids.