A growing body of evidence suggests that gene flow between closely related species is a widespread phenomenon. Alleles that introgress from one species into a close relative are typically neutral or deleterious, but sometimes confer a significant fitness advantage. Given the potential relevance to speciation and adaptation, numerous methods have therefore been devised to identify regions of the genome that have experienced introgression. Recently, supervised machine learning approaches have been shown to be highly effective for detecting introgression. One especially promising approach is to treat population genetic inference as an image classification problem, and feed an image representation of a population genetic alignment as input to a deep neural network that distinguishes among evolutionary models (i.e. introgression or no introgression). However, if we wish to investigate the full extent and fitness effects of introgression, merely identifying genomic regions in a population genetic alignment that harbor introgressed loci is insufficient-ideally we would be able to infer precisely which individuals have introgressed material and at which positions in the genome. Here we adapt a deep learning algorithm for semantic segmentation, the task of correctly identifying the type of object to which each individual pixel in an image belongs, to the task of identifying introgressed alleles. Our trained neural network is thus able to infer, for each individual in a two-population alignment, which of those individual's alleles were introgressed from the other population. We use simulated data to show that this approach is highly accurate, and that it can be readily extended to identify alleles that are introgressed from an unsampled "ghost" population, performing comparably to a supervised learning method tailored specifically to that task. Finally, we apply this method to data from Drosophila, showing that it is able to accurately recover introgressed haplotypes from real data. This analysis reveals that introgressed alleles are typically confined to lower frequencies within genic regions, suggestive of purifying selection, but are found at much higher frequencies in a region previously shown to be affected by adaptive introgression. Our method's success in recovering introgressed haplotypes in challenging real-world scenarios underscores the utility of deep learning approaches for making richer evolutionary inferences from genomic data.
Abstract Rapid evolution may play an important role in the range expansion of invasive species and modify forecasts of invasion, which are the backbone of land management strategies. However, losses of genetic variation associated with colonization bottlenecks may constrain trait and niche divergence at leading range edges, thereby impacting management decisions that anticipate future range expansion. The spatial and temporal scales over which adaptation contributes to invasion dynamics remain unresolved. We leveraged detailed records of the ~130‐year invasion history of the invasive polyploid plant, leafy spurge (Euphorbia virgata), across ~500 km in Minnesota, U.S.A. We examined the consequences of range expansion for population genomic diversity, niche breadth, and the evolution of germination behavior. Using genotyping‐by‐sequencing, we found some population structure in the range core, where introduction occurred, but panmixia among all other populations. Range expansion was accompanied by only modest losses in sequence diversity, with small, isolated populations at the leading edge harboring similar levels of diversity to those in the range core. The climatic niche expanded during most of the range expansion, and the niche of the range core was largely non‐overlapping with the invasion front. Ecological niche models indicated that mean temperature of the warmest quarter was the strongest determinant of habitat suitability and that populations at the leading edge had the lowest habitat suitability. Guided by these findings, we tested for rapid evolution in germination behavior over the time course of range expansion using a common garden experiment and temperature manipulations. Germination behavior diverged from the early to late phases of the invasion, with populations from later phases having higher dormancy at lower temperatures. Our results suggest that trait evolution may have contributed to niche expansion during invasion and that distribution models, which inform future management planning, may underestimate invasion potential without accounting for evolution.
The increasing incidence of bovine congestive heart failure (BCHF) in feedlot cattle poses a significant challenge to the beef industry from economic loss, reduced performance, and reduced animal welfare attributed to cardiac insufficiency. Changes to cardiac morphology as well as abnormal pulmonary arterial pressure (PAP) in cattle of mostly Angus ancestry have been recently characterized. However, congestive heart failure affecting cattle late in the feeding period has been an increasing problem and tools are needed for the industry to address the rate of mortality in the feedlot for multiple breeds. At harvest, a population of 32,763 commercial fed cattle were phenotyped for cardiac morphology with associated production data collected from feedlot processing to harvest at a single feedlot and packing plant in the Pacific Northwest. A sub-population of 5,001 individuals were selected for low-pass genotyping to estimate variance components and genetic correlations between heart score and the production traits observed during the feeding period. At harvest, the incidence of a heart score of 4 or 5 in this population was approximately 4.14%, indicating a significant proportion of feeder cattle are at risk of cardiac mortality before harvest. Heart scores were also significantly and positively correlated with the percentage Angus ancestry observed by genomic breed percentage analysis. The heritability of heart score measured as a binary (scores 1 and 2 = 0, scores 4 and 5 = 1) trait was 0.356 in this population, which indicates development of a selection tool to reduce the risk of congestive heart failure as an EPD (expected progeny difference) is feasible. Genetic correlations of heart score with growth traits and feed intake were moderate and positive (0.289-0.460). Genetic correlations between heart score and backfat and marbling score were -0.120 and -0.108, respectively. Significant genetic correlation to traits of high economic importance in existing selection indexes explain the increased rate of congestive heart failure observed over time. These results indicate potential to implement heart score observed at harvest as a phenotype under selection in genetic evaluation in order to reduce feedlot mortality due to cardiac insufficiency and improve overall cardiopulmonary health in feeder cattle.
Following the discovery of western corn rootworm (WCR; Diabrotica virgifera virgifera) populations resistant to the Bacillus thuringiensis (Bt) protein Cry3Bb1, resistance was genetically mapped to a single locus on WCR chromosome 8 and linked SNP markers were shown to correlate with the frequency of resistance among field-collected populations from the US Corn Belt. The purpose of this paper is to further investigate the relationship between one of these resistance-linked markers and the causal resistance locus. Using data from laboratory bioassays and field experiments, we show that one allele of the resistance-linked marker increased in frequency in response to selection, but was not perfectly linked to the causal resistance allele. By coupling the response to selection data with a genetic model of the linkage between the marker and the causal allele, we developed a model that allowed marker allele frequencies to be mapped to causal allele frequencies. We then used this model to estimate the resistance allele frequency distribution in the US Corn Belt based on collections from 40 populations. These estimates suggest that chromosome 8 Cry3Bb1 resistance allele frequency was generally low (<10%) for 65% of the landscape, though an estimated 13% of landscape has relatively high (>25%) resistance allele frequency.
Variation in complex traits is the result of contributions from many loci of small effect. Based on this principle, genomic prediction methods are used to make predictions of breeding value for an individual using genome-wide molecular markers. In breeding, genomic prediction models have been used in plant and animal breeding for almost two decades to increase rates of genetic improvement and reduce the length of artificial selection experiments. However, evolutionary genomics studies have been slow to incorporate this technique to select individuals for breeding in a conservation context or to learn more about the genetic architecture of traits, the genetic value of missing individuals or microevolution of breeding values. Here, we outline the utility of genomic prediction and provide an overview of the methodology. We highlight opportunities to apply genomic prediction in evolutionary genetics of wild populations and the best practices when using these methods on field-collected phenotypes.
Plant breeders face multiple global challenges that affect food security, productivity, accessibility, and nutritional quality. One major challenge for plant breeders is developing environmentally resilient crop cultivars in response to rapid shifts in cultivation conditions and resources due to climate change. Plant breeders rely on different crop genetic resources, breeding tools, and methods to incorporate genetic diversity into commercialized cultivars. Breeders use genetic diversity to develop new cultivars with improved agronomics, such as higher yield, biotic and abiotic stress tolerance, and to improve the nutritional quality of foods for a growing world population. Plant breeders perform the essential task of strategic integration of new genetic diversity while preserving important economic traits of individual crops such as relative maturity (maize, Zea mays L.), fruit type (tomato, Lycopersicon esculentum Mill.), plant type (lettuce Lactuca sativa L.), and habitat type (canola, Brassica napus L.) that are highly specialized for specific consumer preferences or market needs. This review provides an industry perspective on how genetic diversity is incorporated for crop improvement by (a) using a real-life example to highlight the vast amount of genetic diversity that exists in plants, (b) providing a conceptual example to illustrate strategic challenges a breeder faces while incorporating diversity, (c) describing how and why it can a decade or more to incorporate diversity into commercialized cultivars, even when advanced tools and technologies are used, and (d) sharing factors that plant breeders consider when applying various tools, including genome editing, at different stages of plant breeding.
A challenge to improve an integrative phenotype, like yield, is the interaction between the broad range of possible molecular and physiological traits that contribute to yield and the multitude of potential environmental conditions in which they are expressed. This study collected data on 31 phenotypic traits, 83 annotated metabolites, and nearly 22,000 transcripts from a set of 57 diverse, commercially relevant maize hybrids across three years in central U.S. Corn Belt environments. Although variability in characteristics created a complex picture of how traits interact produce yield, phenotypic traits and gene expression were more consistent across environments, while metabolite levels showed low repeatability. Phenology traits, such as green leaf number and grain moisture and whole plant nitrogen content showed the most consistent correlation with yield. A machine learning predictive analysis of phenotypic traits revealed that ear traits, phenology, and root traits were most important to predicting yield. Analysis suggested little correlation between biomass traits and yield, suggesting there is more of a sink limitation to yield under the conditions studied here. This work suggests that continued improvement of maize yields requires a strong understanding of baseline variation of plant characteristics across commercially-relevant germplasm to drive strategies for consistently improving yield.
Understanding genomic structural variation such as inversions and translocations is a key challenge in evolutionary genetics. We develop a novel statistical approach to comparative genetic mapping to detect large-scale structural mutations from low-level sequencing data. The procedure, called Genome Order Optimization by Genetic Algorithm (GOOGA), couples a Hidden Markov Model with a Genetic Algorithm to analyze data from genetic mapping populations. We demonstrate the method using both simulated data (calibrated from experiments on Drosophila melanogaster) and real data from five distinct crosses within the flowering plant genus Mimulus. Application of GOOGA to the Mimulus data corrects numerous errors (misplaced sequences) in the M. guttatus reference genome and confirms or detects eight large inversions polymorphic within the species complex. Finally, we show how this method can be applied in genomic scans to improve the accuracy and resolution of Quantitative Trait Locus (QTL) mapping.
Population-scale genomic data sets have given researchers incredible amounts of information from which to infer evolutionary histories. Concomitant with this flood of data, theoretical and methodological advances have sought to extract information from genomic sequences to infer demographic events such as population size changes and gene flow among closely related populations/species, construct recombination maps, and uncover loci underlying recent adaptation. To date, most methods make use of only one or a few summaries of the input sequences and therefore ignore potentially useful information encoded in the data. The most sophisticated of these approaches involve likelihood calculations, which require theoretical advances for each new problem, and often focus on a single aspect of the data (e.g., only allele frequency information) in the interest of mathematical and computational tractability. Directly interrogating the entirety of the input sequence data in a likelihood-free manner would thus offer a fruitful alternative. Here, we accomplish this by representing DNA sequence alignments as images and using a class of deep learning methods called convolutional neural networks (CNNs) to make population genetic inferences from these images. We apply CNNs to a number of evolutionary questions and find that they frequently match or exceed the accuracy of current methods. Importantly, we show that CNNs perform accurate evolutionary model selection and parameter estimation, even on problems that have not received detailed theoretical treatments. Thus, when applied to population genetic alignments, CNNs are capable of outperforming expert-derived statistical methods and offer a new path forward in cases where no likelihood approach exists.
Fall armyworm, Spodoptera frugiperda (J.E. Smith) is a major lepidopteran pest of maize in Brazil and its control particularly relies on the use of genetically engineered crops expressing Bacillus thuringiensis (Bt) toxins such as Cry1F. However, control failures compromising the efficacy of this technology have been reported in many regions in Brazil, but the mechanism of Cry1F resistance in Brazilian fall armyworm populations remained elusive. Here we investigated the molecular mechanism of Cry1F resistance in two field-collected strains of S. frugiperda from Brazil exhibiting high levels of Cry1F resistance. We first rigorously evaluated several candidate reference genes for normalization of gene expression data across strains, larval instars and gut tissues, and identified ribosomal proteins L10, L17 and RPS3A to be most suitable. We then investigated the expression pattern of ten potential Bt toxin receptors/enzymes in both neonates and 2nd instar gut tissue of Cry1F resistant fall armyworm strains compared to a susceptible strain. Next we sequenced the ATP-dependent Binding Cassette subfamily C2 gene (ABCC2) and identified three mutated sites present in ABCC2 of both Cry1F resistant strains: two of them, a GY deletion (positions 788-789) and a P799 K/R amino acid substitution, located in a conserved region of ABCC2 extracellular loop 4 (EC4) and another amino acid substitution, G1088D, but in a less conserved region. We further characterized the role of the novel mutations present in EC4 by functionally expressing both wild type and mutated ABCC2 transporters in insect cell lines, and confirmed a critical role of both sites for Cry1F binding by cell viability assays. Finally, we assessed the frequency of the mutant alleles by pooled population sequencing and pyrosequencing in 40 fall armyworm populations collected from maize fields in different regions in Brazil. We found that the GY deletion being present at high frequency. However we also observed many rare alleles which disrupt residues between sites 783-799, and their diversity and abundance in field collected populations lends further support to the importance of the EC4 domain for Cry1F toxicity.
The use of Bt proteins in crops has revolutionized insect pest management by offering effective season-long control. However, field-evolved resistance to Bt proteins threatens their utility and durability. A recent example is field-evolved resistance to Cry1Fa and Cry1A.105 in fall armyworm (Spodoptera frugiperda). This resistance has been detected in Puerto Rico, mainland USA, and Brazil. A S. frugiperda population with suspected resistance to Cry1Fa was sampled from a maize field in Puerto Rico and used to develop a resistant lab colony. The colony demonstrated resistance to Cry1Fa and partial cross-resistance to Cry1A.105 in diet bioassays. Using genetic crosses and proteomics, we show that this resistance is due to loss-of-function mutations in the ABCC2 gene. We characterize two novel mutant alleles from Puerto Rico. We also find that these alleles are absent in a broad screen of partially resistant Brazilian populations. These findings confirm that ABCC2 is a receptor for Cry1Fa and Cry1A.105 in S. frugiperda, and lay the groundwork for genetically enabled resistance management in this species, with the caution that there may be several distinct ABCC2 resistances alleles in nature.
The use of dsRNA to control insect pests via the RNA interference (RNAi) pathway is being explored by researchers globally. However, with every new class of insect control compounds, the evolution of insect resistance needs to be considered, and understanding resistance mechanisms is essential in designing durable technologies and effective resistance management strategies. To gain insight into insect resistance to dsRNA, a field screen with subsequent laboratory selection was used to establish a population of DvSnf7 dsRNA-resistant western corn rootworm, Diabrotica virgifera virgifera, a major maize insect pest. WCR resistant to ingested DvSnf7 dsRNA had impaired luminal uptake and resistance was not DvSnf7 dsRNA-specific, as indicated by cross resistance to all other dsRNAs tested. No resistance to the Bacillus thuringiensis Cry3Bb1 protein was observed. DvSnf7 dsRNA resistance was inherited recessively, located on a single locus, and autosomal. Together these findings will provide insights for dsRNA deployment for insect pest control.
Understanding genomic structural variation such as inversions and translocations is a key challenge in evolutionary genetics. In this paper, we tackle this challenge by developing a novel statistical approach to comparative genetic mapping. The procedure couples a Hidden Markov Model with a Genetic Algorithm to detect large-scale structural variation using low-level sequencing data from multiple genetic mapping populations. We demonstrate the method using five distinct crosses within the flowering plant genus Mimulus . The synthesis of data from these experiments is first used to correct numerous errors (misplaced sequences) in the M. guttatus reference genome. Second, we confirm and/or detect eight large inversions polymorphic within the M. guttatus species complex. Finally, we show how this method can be applied in genomic scans to improve the accuracy and resolution of Quantitative Trait Locus (QTL) mapping. AUTHOR SUMMARY Genome sequences have proved to be a critical experimental resource for genetic research in many species. However, in some species there is considerable variation in genomic organization, making a single reference genome sequence inadequate. This variation can cause issues in interpreting genomic signals, such as those coming from trait mapping. We introduce a new statistical method and computational tools that use linkage information to reorganize a single reference genome to 1) repair genome assembly errors, and 2) identify variation between individuals or populations of the same species. Using this method we can create a new genome order that improves upon the reference genome. We apply this method to five crosses among plants in the Mimulus guttatus species complex. In this system we detect eight large chromosomal inversions and improve the resolution of a trait mapping study. This work highlights the utility of our method, and indicates how others studying diverse species might use them to improve their own research.
Background: Gene duplication is prevalent in many species and can result in coding and regulatory divergence. Gene duplications can be classified as whole genome duplication (WGD), tandem and inserted (non-syntenic). In maize, WGD resulted in the subgenomes maize1 and maize2, of which maize1 is considered the dominant subgenome. However, the landscape of co-expression network divergence of duplicate genes in maize is still largely uncharacterized.Results: To address the consequence of gene duplication on co-expression network divergence, we developed a gene co-expression network from RNA-seq data derived from 64 different tissues/stages of the maize reference inbred-B73. WGD, tandem and inserted gene duplications exhibited distinct regulatory divergence. Inserted duplicate genes were more likely to be singletons in the co-expression networks, while WGD duplicate genes were likely to be co-expressed with other genes. Tandem duplicate genes were enriched in the co-expression pattern where co-expressed genes were nearly identical for the duplicates in the network. Older gene duplications exhibit more extensive co-expression variation than younger duplications. Overall, non-syntenic genes primarily from inserted duplications show more co-expression divergence. Also, such enlarged co-expression divergence is significantly related to duplication age. Moreover, subgenome dominance was not observed in the co-expression networks - maize1 and maize2 exhibit similar levels of intra subgenome correlations. Intriguingly, the level of inter subgenome co-expression was similar to the level of intra subgenome correlations, and genes from specific subgenomes were not likely to be the enriched in co-expression network modules and the hub genes were not predominantly from any specific subgenomes in maize.Conclusions: Our work provides a comprehensive analysis of maize co-expression network divergence for three different types of gene duplications and identifies potential relationships between duplication types, duplication ages and co-expression consequences.
GO Enrichment for duplicate genes with different co-expression types. (XLSX 21 kb)
Western corn rootworm (WCR) is a major maize (Zea mays L.) pest leading to annual economic losses of more than 1 billion dollars in the United States. Transgenic maize expressing insecticidal toxins derived from the bacterium Bacillus thuringiensis (Bt) are widely used for the management of WCR. However, cultivation of Bt-expressing maize places intense selection pressure on pest populations to evolve resistance. Instances of resistance to Bt toxins have been reported in WCR. Developing genetic markers for resistance will help in characterizing the extent of existing issues, predicting where future field failures may occur, improving insect resistance management strategies, and in designing and sustainably implementing forthcoming WCR control products. Here, we discover and validate genetic markers in WCR that are associated with resistance to the Cry3Bb1 Bt toxin. A field-derived WCR population known to be resistant to the Cry3Bb1 Bt toxin was used to generate a genetic map and to identify a genomic region associated with Cry3Bb1 resistance. Our results indicate that resistance is inherited in a nearly recessive manner and associated with a single autosomal linkage group. Markers tightly linked with resistance were validated using WCR populations collected from Cry3Bb1 maize fields showing significant WCR damage from across the US Corn Belt. Two markers were found to be correlated with both diet (R-2 = 0.14) and plant (R-2 = 0.23) bioassays for resistance. These results will assist in assessing resistance risk for different WCR populations, and can be used to improve insect resistance management strategies.
Significance Comparative analyses of central molecular networks uncover variation that can be targeted by biomedical research to develop insights and interventions into disease. The insulin/insulin-like signaling and target of rapamycin (IIS/TOR) molecular network regulates metabolism, growth, and aging. With the development of new molecular resources for reptiles, we show that genes in IIS/TOR are rapidly evolving within amniotes (mammals and reptiles, including birds). Additionally, we find evidence of natural selection that diversified the hormone-receptor binding relationships that initiate IIS/TOR signaling. Our results uncover substantial variation in the IIS/TOR network within and among amniotes and provide a critical step to unlocking information on vertebrate patterns of genetic regulation of metabolism, modes of reproduction, and rates of aging.
The insulin/insulin-like signaling and target of rapamycin (IIS/TOR) network regulates lifespan and reproduction, as well as metabolic diseases, cancer, and aging. Despite its vital role in health, comparative analyses of IIS/TOR have been limited to invertebrates and mammals. We conducted an extensive evolutionary analysis of the IIS/TOR network across 66 amniotes with 18 newly generated transcriptomes from nonavian reptiles and additional available genomes/transcriptomes. We uncovered rapid and extensive molecular evolution between reptiles (including birds) and mammals: (i) the IIS/TOR network, including the critical nodes insulin receptor substrate (IRS) and phosphatidylinositol 3-kinase (PI3K), exhibit divergent evolutionary rates between reptiles and mammals; (ii) compared with a proxy for the rest of the genome, genes of the IIS/TOR extracellular network exhibit exceptionally fast evolutionary rates; and (iii) signatures of positive selection and coevolution of the extracellular network suggest reptile- and mammal-specific interactions between members of the network. In reptiles, positively selected sites cluster on the binding surfaces of insulin-like growth factor 1 (IGF1), IGF1 receptor (IGF1R), and insulin receptor (INSR); whereas in mammals, positively selected sites clustered on the IGF2 binding surface, suggesting that these hormone-receptor binding affinities are targets of positive selection. Further, contrary to reports that IGF2R binds IGF2 only in marsupial and placental mammals, we found positively selected sites clustered on the hormone binding surface of reptile IGF2R that suggest that IGF2R binds to IGF hormones in diverse taxa and may have evolved in reptiles. These data suggest that key IIS/TOR paralogs have sub- or neofunctionalized between mammals and reptiles and that this network may underlie fundamental life history and physiological differences between these amniote sister clades.