Understanding the molecular basis of adaptation and the genetic architecture of complex traits are longstanding goals in biology. One problem impeding this understanding is the complexity of continental populations, with their complicated demographic histories, gene flow and secondary contact. In contrast, island populations represent simpler systems where uncovering the genetic basis of complex traits and tracing how traits built up is much more tractable. In Arabidopsis thaliana, the Cape Verde Islands populations represent a case of long-range colonization and adaptation to a divergent selective regime. Here, we describe the development and testing of a new multiparent intercross doubled haploid population of A. thaliana from the Cape Verde Islands. This population balances the representation of natural diversity and overcomes the shortcomings of existing resources, such as biparental recombinant inbred lines and genome-wide association populations. Specifically, it captures variation that segregates within the archipelago but is fixed on individual islands. We mapped the genetic basis of flowering time, rosette size, and photosystem II efficiency (ΦPSII) in this inter-island intercross population, representing traits that we hypothesized may be evolving under strong selection during the colonization of the archipelago. We identified functional loci underlying these traits, including FRI K232X and FLC R3X for flowering time, and IRT1 G130X for ΦPSII and rosette size. Our multiparent intercross population complements existing mapping resources and provides a robust framework for investigating the genetic basis of complex traits in A. thaliana. This work emphasizes the value of island systems and complementary approaches for advancing our understanding of genetic adaptation.
Photosynthesis is acknowledged as a potential target to increase crop yield. Improved photosynthesis may be achieved by conventional breeding, exploiting the available natural genetic variation for photosynthesis traits. This approach is challenging for crops due to limitations in high-throughput photosynthesis phenotyping, the highly polygenic nature of photosynthesis, and its strongly dynamic response to environmental changes. Recent advancements in phenomics make accurate and detailed photosynthesis phenotyping more feasible, with the model species Arabidopsis thaliana paving the way for applications in crops. In this study, we examined photosynthesis parameters over time in the global Arabidopsis HapMap diversity panel exposed to three conditions: optimal nutrient supply, low phosphorus supply and low nitrogen supply. Combined with two previous studies on photosynthesis in response to low temperature, and to a one-step change in irradiance from low light to high light, five high-quality datasets were systematically analysed using the same approach (with one million-maker set, uni- and multi-variate analyses). Our findings emphasize the genetic complexity of photosynthesis, detecting hundreds of significant quantitative trait loci, only a small number of which are robust, and of which most are condition specific. Robust loci, found in multiple conditions, exemplify those suited for conferring higher all-round photosynthesis, and targets for marker-assisted selection, contributing to environmental resilience, while the multitude of small-effect conditional loci suggest that genomic selection approaches may be more suited to improve crop photosynthesis.
Making sense of whole-genome polymorphism data is challenging, but it is essential for overcoming the biases in SNP data. Here we analyze 27 genomes of Arabidopsis thaliana to illustrate these issues. Genome size variation is mostly due to tandem repeat regions that are difficult to assemble. However, while the rest of the genome varies little in length, it is full of structural variants, mostly due to transposon insertions. Because of this, the pangenome coordinate system grows rapidly with sample size and ultimately becomes 70% larger than the size of any single genome, even for n = 27. Finally, we show how short-read data are biased by read mapping. SNP calling is biased by the choice of reference genome, and both transcriptome and methylome profiling results are affected by mapping reads to a reference genome rather than to the genome of the assayed individual.
Our view of genetic polymorphism is shaped by methods that provide a limited and reference-biased picture. Long-read sequencing technologies, which are starting to provide nearly complete genome sequences for population samples, should solve the problem—except that characterizing and making sense of non-SNP variation is difficult even with perfect sequence data. Here, we analyze 27 genomes of Arabidopsis thaliana in an attempt to address these issues, and illustrate what can be learned by analyzing whole-genome polymorphism data in an unbiased manner. Estimated genome sizes range from 135 to 155 Mb, with differences almost entirely due to centromeric and rDNA repeats. The completely assembled chromosome arms comprise roughly 120 Mb in all accessions, but are full of structural variants, many of which are caused by insertions of transposable elements (TEs) and subsequent partial deletions of such insertions. Even with only 27 accessions, a pan-genome coordinate system that includes the resulting variation ends up being 40% larger than the size of any one genome. Our analysis reveals an incompletely annotated mobile-ome: our ability to predict what is actually moving is poor, and we detect several novel TE families. In contrast to this, the genic portion, or “gene-ome”, is highly conserved. By annotating each genome using accession-specific transcriptome data, we find that 13% of all genes are segregating in our 27 accessions, but that most of these are transcriptionally silenced. Finally, we show that with short-read data we previously massively underestimated genetic variation of all kinds, including SNPs—mostly in regions where short reads could not be mapped reliably, but also where reads were mapped incorrectly. We demonstrate that SNP-calling errors can be biased by the choice of reference genome, and that RNA-seq and BS-seq results can be strongly affected by mapping reads to a reference genome rather than to the genome of the assayed individual. In conclusion, while whole-genome polymorphism data pose tremendous analytical challenges, they will ultimately revolutionize our understanding of genome evolution.### Competing Interest StatementD.W.∼holds equity in Computomics, which advises plant breeders. D.W.∼also consults for KWS SE, a plant breeder and seed producer with activities throughout the world. J.F. is an employee of Tropic TI, Lda. All other authors declare no competing interests.
The environments in which plant species evolved are now generally understood to be dynamic rather than static. Photosynthesis has to operate within these dynamic environments, such as sudden changes to light intensities. Plants have evolved photoprotection mechanisms that prevent damage caused by sudden changes to high light intensities. The extent of genetic variation within plants species to deal with these dynamic light conditions remains largely unexplored. Here we show that one accession of A. thaliana has a more efficient photoprotection mechanism in dynamic light conditions, compared to six other accessions. The construction of a doubled haploid population and subsequent phenotyping in a dynamically controlled high-throughput system reveals up to 15 QTLs for photoprotection. Identifying the causal gene underlying one of the major QTLs shows that an allelic variant of cpFtsY results in more efficient photoprotection under high and fluctuating light intensities. Further analyses reveal this allelic variant to be overprotecting, reducing biomass in a range of dynamic environmental conditions. This suggests that within nature, adaptation can occur to more stressful environments and that revealing the causal genes and mechanisms can help improve the general understanding of photosynthetic functioning. The other QTLs possess different photosynthetic properties, and thus together they show how there is ample intraspecific genetic variation for photosynthetic functioning in dynamic environments. With photosynthesis being one of the last unimproved components of crop yield, this amount of genetic variation for photosynthesis forms excellent input for breeding approaches. In these breeding approaches, the interactions with the environmental conditions should however be precisely assessed. Doing so correctly, allows us to tap into nature’s solution to challenging environmental conditions.
Understanding how populations adapt to abrupt environmental change is necessary to predict responses to future challenges, but identifying specific adaptive variants, quantifying their responses to selection and reconstructing their detailed histories is challenging in natural populations. Here, we use Arabidopsis from the Cape Verde Islands as a model to investigate the mechanisms of adaptation after a sudden shift to a more arid climate. We find genome-wide evidence of adaptation after a multivariate change in selection pressures. In particular, time to flowering is reduced in parallel across islands, substantially increasing fitness. This change is mediated by convergent de novo loss of function of two core flowering time genes: FRI on one island and FLC on the other. Evolutionary reconstructions reveal a case where expansion of the new populations coincided with the emergence and proliferation of these variants, consistent with models of rapid adaptation and evolutionary rescue.
Most well-characterized cases of adaptation involve single genetic loci. Theory suggests that multilocus adaptive walks should be common, but these are challenging to identify in natural populations. Here, we combine trait mapping with population genetic modeling to show that a two-step process rewired nutrient homeostasis in a population of Arabidopsis as it colonized the base of an active stratovolcano characterized by extremely low soil manganese (Mn). First, a variant that disrupted the primary iron (Fe) uptake transporter gene (IRT1) swept quickly to fixation in a hard selective sweep, increasing Mn but limiting Fe in the leaves. Second, multiple independent tandem duplications occurred at NRAMP1 and together rose to near fixation in the island population, compensating the loss of IRT1 by improving Fe homeostasis. This study provides a clear case of a multilocus adaptive walk and reveals how genetic variants reshaped a phenotype and spread over space and time.
Objectives Lathyrus tuberosus is a nitrogen-fixing member of the Fabaceae which forms protein-rich tubers. To aid future domestication programs for this legume plant and facilitate evolutionary studies of tuber formation, we have generated a draft genome assembly based on Pacific Biosciences sequence reads. Data description Genomic DNA from L. tuberosus was sequenced with PacBio’s HiFi sequencing chemistry generating 12.8 million sequence reads with an average read length of 14 kb (approximately 180 Gb of sequence data). The reads were assembled to give a draft genome of 6.8 Gb in 1353 contigs with an N50 contig length of 11.1 Mb. The GC content of the genome assembly was 38.3%. BUSCO analysis of the genome assembly indicated a genome completeness of at least 96%. The genome sequence will be a valuable resource, for example, in assessing genomic consequences of domestication efforts and developing marker sets for breeding programs. The L. tuberosus genome will also aid in the analysis of the evolutionary history of plants within the nitrogen-fixing Fabaceae family and in understanding the molecular basis of tuber evolution.
Urbanization transforms environments in ways that alter biological evolution. We examined whether urban environmental change drives parallel evolution by sampling 110,019 white clover plants from 6169 populations in 160 cities globally. Plants were assayed for a Mendelian antiherbivore defense that also affects tolerance to abiotic stressors. Urban-rural gradients were associated with the evolution of clines in defense in 47% of cities throughout the world. Variation in the strength of clines was explained by environmental changes in drought stress and vegetation cover that varied among cities. Sequencing 2074 genomes from 26 cities revealed that the evolution of urban-rural clines was best explained by adaptive evolution, but the degree of parallel adaptation varied among cities. Our results demonstrate that urbanization leads to adaptation at a global scale.
Discoveries of adaptive gene knockouts and widespread losses of complete genes have in recent years led to a major rethink of the early view that loss-of-function alleles are almost always deleterious. Today, surveys of population genomic diversity are revealing extensive loss-of-function and gene content variation, yet the adaptive significance of much of this variation remains unknown. Here we examine the evolutionary dynamics of adaptive loss of function through the lens of population genomics and consider the challenges and opportunities of studying adaptive loss-of-function alleles using population genetics models. We discuss how the theoretically expected existence of allelic heterogeneity, defined as multiple functionally analogous mutations at the same locus, has proven consistent with empirical evidence and why this impedes both the detection of selection and causal relationships with phenotypes. We then review technical progress towards new functionally explicit population genomic tools and genotype-phenotype methods to overcome these limitations. More broadly, we discuss how the challenges of studying adaptive loss of function highlight the value of classifying genomic variation in a way consistent with the functional concept of an allele from classical population genetics.
Assessment of the impact of variation in chloroplast and mitochondrial DNA (collectively termed the plasmotype) on plant phenotypes is challenging due to the difficulty in separating their effect from nuclear-derived variation (the nucleotype). Haploid-inducer lines can be used as efficient plasmotype donors to generate new plasmotype–nucleotype combinations (cybrids)1. We generated a panel comprising all possible cybrids of seven Arabidopsis thaliana accessions and extensively phenotyped these lines for 1,859 phenotypes under both stable and fluctuating conditions. We show that natural variation in the plasmotype results in both additive and epistatic effects across all phenotypic categories. Plasmotypes that induce more additive phenotypic changes also cause more epistatic effects, suggesting a possible common basis for both additive and epistatic effects. On average, epistatic interactions explained twice as much of the variance in phenotypes as additive plasmotype effects. The impact of plasmotypic variation was also more pronounced under fluctuating and stressful environmental conditions. Thus, the phenotypic impact of variation in plasmotypes is the outcome of multi-level nucleotype–plasmotype–environment interactions and, as such, the plasmotype is likely to serve as a reservoir of variation that is predominantly exposed under certain conditions. The production of cybrids using haploid inducers is a rapid and precise method for assessment of the phenotypic effects of natural variation in organellar genomes. It will facilitate efficient screening of unique nucleotype–plasmotype combinations to both improve our understanding of natural variation in these combinations and identify favourable combinations to enhance plant performance. Plants store the vast majority of their DNA in the nucleus, like all other eukaryotes, but also possess two sets of organellar genomes in mitochondria and plastids. Now, researchers have employed haploid inducers to generate reciprocal cybrids to disentangle the specific contributions of organellar variations to plant performance.
Trichomes are distinctive features of plant stems and leaves. Their primary role is to protect plants—functioning as physical barriers, stinging hairs, or chemical factories that produce a diverse array of compounds. At the extreme, sundews ( Drosera sp) evolved trichomes that capture and digest
Photosynthesis is the gateway of the Sun's energy into the biosphere and the source of the ozone layer; thus it is both provider and protector of life as we know it. Despite its pivotal role we know surprisingly little about the genetic basis of variation in photosynthesis and the selective pressures giving rise to or maintaining this variation. In this review, I will briefly summarise our current knowledge of intraspecific and interspecific variation in photosynthesis to understand the main selective constraints on photosynthesis and what this means for the future of nature and agriculture in a changing world.
Meiotic crossovers (COs) ensure proper chromosome segregation and redistribute the genetic variation that is transmitted to the next generation. Large populations and the demand for genome-wide, fine-scale resolution challenge existing methods for CO identification. Taking advantage of linked-read sequencing, we develop a highly efficient method for genome-wide identification of COs at kilobase resolution in pooled recombinants. We first test this method using a pool of Arabidopsis F2 recombinants, and recapitulate results obtained from the same plants using individual whole-genome sequencing. By applying this method to a pool of pollen DNA from an F1 plant, we establish a highly accurate CO landscape without generating or sequencing a single recombinant plant. The simplicity of this approach enables the simultaneous generation and analysis of multiple CO landscapes, accelerating the pace at which mechanisms for the regulation of recombination can be elucidated through efficient comparisons of genotypic and environmental effects on recombination.
Background Photosynthesis underpins plant productivity and yet is notoriously sensitive to small changes in environmental conditions, meaning that quantitation in nature across different time scales is not straightforward. The 'dynamic' changes in photosynthesis (i.e. the kinetics of the various reactions of photosynthesis in response to environmental shifts) are now known to be important in driving crop yield. Scope It is known that photosynthesis does not respond in a timely manner, and even a small temporal 'mismatch' between a change in the environment and the appropriate response of photosynthesis toward optimality can result in a fall in productivity. Yet the most commonly measured parameters are still made at steady state or a temporary steady state (including those for crop breeding purposes), meaning that new photosynthetic traits remain undiscovered. Conclusions There is a great need to understand photosynthesis dynamics from a mechanistic and biological viewpoint especially when applied to the field of 'phenomics' which typically uses large genetically diverse populations of plants. Despite huge advances in measurement technology in recent years, it is still unclear whether we possess the capability of capturing and describing the physiologically relevant dynamic features of field photosynthesis in sufficient detail. Such traits are highly complex, hence we dub this the 'photosynthome'. This review sets out the state of play and describes some approaches that could be made to address this challenge with reference to the relevant biological processes involved.
Over the past 20 y, many studies have examined the history of the plant ecological and molecular model, Arabidopsis thaliana, in Europe and North America. Although these studies informed us about the recent history of the species, the early history has remained elusive. In a large-scale genomic analysis of African A. thaliana, we sequenced the genomes of 78 modern and herbarium samples from Africa and analyzed these together with over 1,000 previously sequenced Eurasian samples. In striking contrast to expectations, we find that all African individuals sampled are native to this continent, including those from sub-Saharan Africa. Moreover, we show that Africa harbors the greatest variation and represents the deepest history in the A. thaliana lineage. Our results also reveal evidence that selfing, a major defining characteristic of the species, evolved in a single geographic region, best represented today within Africa. Demographic inference supports a model in which the ancestral A. thaliana population began to split by 120-90 kya, during the last interglacial and Abbassia pluvial, and Eurasian populations subsequently separated from one another at around 40 kya. This bears striking similarities to the patterns observed for diverse species, including humans, implying a key role for climatic events during interglacial and pluvial periods in shaping the histories and current distributions of a wide range of species.
Plants are powerful models for the study of adaptive evolution. Since they are rooted in place, they must directly face environmental insults, making adaptation to local conditions vital. In addition to adaptation to natural conditions, some plant species have held a central role in human subsistence over the past several thousand years. In these species, humans exerted strong selective pressures on traits of agricultural importance. Recently, an increasing number of studies have aimed to identify the genomic basis of adaptation. These studies have provided insights into the mechanisms through which the raw materials of adaptation were introduced as well as the modes of adaptation in wild and domesticated species.
Strong selection on a beneficial mutation can cause a selective sweep, which fixes the mutation in the population and reduces the genetic variation in the region flanking the mutation [1-3]. These flanking regions have increased in frequency due to their physical association with the selected loci, a phenomenon called "genetic hitchhiking'' [4]. Theoretically, selection could extend the hitchhiking to unlinked parts of the genome, to the point that selection on organelles affects nuclear genome diversity. Such indirect selective sweeps have never been observed in nature. Here we show that strong selection on a chloroplast gene in the wild plant species Arabidopsis thaliana has caused widespread and lasting hitchhiking of the whole nuclear genome. The selected allele spread more than 400 km along the British railway network, reshaping the genetic composition of local populations. This demonstrates that selection on organelle genomes can significantly reduce nuclear genetic diversity in natural populations. We expect that organelle-mediated genetic draft is a more common occurrence than previously realized and needs to be considered when studying genome evolution.
Introduction Batch effects in large untargeted metabolomics experiments are almost unavoidable, especially when sensitive detection techniques like mass spectrometry (MS) are employed. In order to obtain peak intensities that are comparable across all batches, corrections need to be performed. Since non-detects, i.e., signals with an intensity too low to be detected with certainty, are common in metabolomics studies, the batch correction methods need to take these into account.Objectives This paper aims to compare several batch correction methods, and investigates the effect of different strategies for handling non-detects.Methods Batch correction methods usually consist of regression models, possibly also accounting for trends within batches. To fit these models quality control samples (QCs), injected at regular intervals, can be used. Also study samples can be used, provided that the injection order is properly randomized. Normalization methods, not using information on batch labels or injection order, can correct for batch effects as well. Introducing two easy-to-use quality criteria, we assess the merits of these batch correction strategies using three large LC-MS and GC-MS data sets of samples from Arabidopsis thaliana.Results The three data sets have very different characteristics, leading to clearly distinct behaviour of the batch correction strategies studied. Explicit inclusion of information on batch and injection order in general leads to very good corrections; when enough QCs are available, also general normalization approaches perform well. Several approaches are shown to be able to handle non-detects-replacing them with very small numbers such as zero seems the worst of the approaches considered.Conclusion The use of quality control samples for batch correction leads to good results when enough QCs are available. If an experiment is properly set up, batch correction using the study samples usually leads to a similar high-quality correction, but has the advantage that more metabolites are corrected. The strategy for handling non-detects is important: choosing small values like zero can lead to suboptimal batch corrections.
Background Recent advances in genome sequencing technologies have shifted the research bottleneck in plant sciences from genotyping to phenotyping. This shift has driven the development of phenomics, high-throughput non-invasive phenotyping technologies. Results We describe an automated high-throughput phenotyping platform, the Phenovator, capable of screening 1440 Arabidopsis plants multiple times per day for photosynthesis, growth and spectral reflectance at eight wavelengths. Using this unprecedented phenotyping capacity, we have been able to detect significant genetic differences between Arabidopsis accessions for all traits measured, across both temporal and environmental scales. The high frequency of measurement allowed us to observe that heritability was not only trait specific, but for some traits was also time specific. Conclusions Such continuous real-time non-destructive phenotyping will allow detailed genetic and physiological investigations of the kinetics of plant homeostasis and development. The success and ultimate outcome of a breeding program will depend greatly on the genetic variance which is sampled. Our observation of temporal fluctuations in trait heritability shows that the moment of measurement can have lasting consequences. Ultimately such phenomic level technologies will provide more dynamic insights into plant physiology, and the necessary data for the omics revolution to reach its full potential.