Parasites, including pathogens, can adapt to better exploit their hosts on many scales, ranging from within an infection of a single individual to series of infections spanning multiple host species. However, little is known about how the genomes of parasites in natural communities evolve when they face diverse hosts. We investigated how Bartonella bacteria that circulate in rodent communities in the dunes of the Negev Desert in Israel adapt to different species of rodent hosts. We propagated 15 Bartonella populations through infections of either a single host species (Gerbillus andersoni or Gerbillus pyramidum) or alternating between the two. After 20 rodent passages, strains with de novo mutations replaced the ancestor in most populations. Mutations in two mononucleotide simple sequence repeats (SSRs) that caused frameshifts in the same adhesin gene dominated the evolutionary dynamics. They appeared exclusively in populations that encountered G. andersoni and altered the dynamics of infections of this host. Similar SSRs in other genes are conserved and exhibit ON/OFF variation in Bartonella isolates from the Negev Desert dunes. Our results suggest that SSR-based contingency loci could be important not only for rapidly and reversibly generating antigenic variation to escape immune responses but that they may also mediate the evolution of host specificity.
Supplementary Table 1 from Xenoestrogen-Induced Epigenetic Repression of microRNA-9-3 in Breast Epithelial Cells
Supplementary Table 2 from Xenoestrogen-Induced Epigenetic Repression of microRNA-9-3 in Breast Epithelial Cells
Supplementary Methods, Figures 1-5, Tables 1-5 from Breast Cancer–Associated Fibroblasts Confer AKT1-Mediated Epigenetic Silencing of Cystatin M in Epithelial Cells
Supplementary Figures 1-6 from Epigenetic Repression of <i>microRNA-129-2</i> Leads to Overexpression of <i>SOX4</i> Oncogene in Endometrial Cancer
Supplementary Figures 1-6, Tables 1-15, Methods from Epigenetic Silencing Mediated through Activated PI3K/AKT Signaling in Breast Cancer
Supplementary Figures 1-6 from Epigenetic Repression of microRNA-129-2 Leads to Overexpression of SOX4 Oncogene in Endometrial Cancer
Understanding how cells are likely to evolve can guide medical interventions and bioengineering efforts that must contend with unwanted mutations. The adaptome of a cell-the neighborhood of genetic changes that are most likely to drive adaptation in a given environment-can be mapped by tracking rare beneficial variants during the early stages of clonal evolution. We used multiplex adaptome capture sequencing (mAdCap-seq), a procedure that combines unique molecular identifiers and hybridization-based enrichment, to characterize mutations in eight Escherichia coli genes known to be under selection in a laboratory environment. We tracked 301 mutations at frequencies as low as 0.01% and inferred the fitness effects of 240 of these mutations. There were distinct molecular signatures of selection on protein structure and function for the three genes with the most beneficial mutations. Our results demonstrate how mAdCap-seq can be used to deeply profile a targeted portion of a cell's adaptome.
Clonal populations of cells continuously evolve new genetic diversity, but it takes a significant amount of time for the progeny of a single cell with a new beneficial mutation to outstrip both its ancestor and competitors to fully dominate a population. If these driver mutations can be discovered earlier—while they are still extremely rare—and profiled in large numbers, it may be possible to anticipate the future evolution of similar cell populations. For example, one could diagnose the likely course of incipient diseases, such as cancer and bacterial infections, and better judge which treatments will be effective, by tracking rare drug-resistant variants. To test this approach, we replayed the first 500 generations of a >70,000-generation Escherichia coli experiment and examined the trajectories of new mutations in eight genes known to be under positive selection in this environment in six populations. By employing a deep sequencing procedure using unique molecular identifiers and target enrichment we were able to track 301 beneficial mutations at frequencies as low as 0.01% and infer the fitness effects of 240 of these. Distinct molecular signatures of selection on protein structure and function were evident for the three genes in which beneficial mutations were most common ( nadR, pykF, and topA ). We detected mutations hundreds of generations before they became dominant and tracked beneficial alleles in genes that were not mutated in the long-term experiment until thousands of generations had passed. This type of targeted adaptome sequencing approach could function as an early warning system to inform interventions that aim to prevent undesirable evolution.
The potency and indiscriminate nature of formaldehyde reactivity upon biological molecules make it a universal stressor. However, some organisms such as Methylorubrum extorquens possess means to rapidly and effectively mitigate formaldehyde-induced damage. EfgA is a recently identified formaldehyde sensor predicted to halt translation in response to elevated formaldehyde as a means to protect cells. Herein, we investigate growth and changes in gene expression to understand how M. extorquens responds to formaldehyde with and without the EfgA-formaldehyde-mediated translational response, and how this mechanism compares to antibiotic-mediated translation inhibition. These distinct mechanisms of translation inhibition have notable differences: they each involve different specific players and in addition, formaldehyde also acts as a general, multi-target stressor and a potential carbon source. We present findings demonstrating that in addition to its characterized impact on translation, functional EfgA allows for a rapid and robust transcriptional response to formaldehyde and that removal of EfgA leads to heightened proteotoxic and genotoxic stress in the presence of increased formaldehyde levels. We also found that many downstream consequences of translation inhibition were shared by EfgA-formaldehyde- and kanamycin-mediated translation inhibition. Our work uncovered additional layers of regulatory control enacted by functional EfgA upon experiencing formaldehyde stress, and further demonstrated the importance this protein plays at both transcriptional and translational levels in this model methylotroph.
Experimental studies of evolution using microbes have a long tradition, and these studies have increased greatly in number and scope in recent decades. Most such experiments have been short in duration, typically running for weeks or months. A venerable exception, the long-term evolution experiment (LTEE) with Escherichia coli has continued for 30 years and 70,000 bacterial generations. The LTEE has become one of the cornerstones of the field of experimental evolution, in general, and the BEACON Center for the Study of Evolution in Action, in particular. Science laboratories and experiments usually have finite lifespans, but we hope that the LTEE can continue far into the future. There are practical issues associated with maintaining such a long-term experiment. One issue, which we address here, is whether key measurements made at one time and place are reproducible, within reasonable limits, at other times and places. This issue comes to the forefront when one considers moving an experiment like the LTEE from one lab to another. To that end, the Barrick lab at The University of Texas at Austin, measured the fitness values of samples from the 12 LTEE populations at 2,000, 10,000, and 50,000 generations and compared the new data to data previously obtained at Michigan State University. On balance, the datasets agree very well. More generally, this finding shows the value of simplicity in experimental design, such as using a chemically defined growth medium and appropriately storing samples from microbiological experiments. Even so, one must be vigilant in checking assumptions and procedures given the potential for uncontrolled factors (e.g., water quality) to affect outcomes. This vigilance is perhaps especially important for a trait like fitness, which integrates all aspects of organismal performance and may therefore be sensitive to any number of subtle environmental influences.
Transmissible plasmids spread genes encoding antibiotic resistance and other traits to new bacterial species. Here we report that laboratory populations of Escherichia coli with a newly acquired IncQ plasmid often evolve 'satellite plasmids' with deletions of accessory genes and genes required for plasmid replication. Satellite plasmids are molecular parasites: their presence reduces the copy number of the full-length plasmid on which they rely for their continued replication. Cells with satellite plasmids gain an immediate fitness advantage from reducing burdensome expression of accessory genes. Yet, they maintain copies of these genes and the complete plasmid, which potentially enables them to benefit from and transmit the traits they encode in the future. Evolution of satellite plasmids is transient. Cells that entirely lose accessory gene function or plasmid mobility dominate in the long run. Satellite plasmids also evolve in Snodgrassella alvi colonizing the honey bee gut, suggesting that this mechanism may broadly contribute to the importance of IncQ plasmids as agents of bacterial gene transfer in nature.
Unwanted evolution of designed DNA sequences limits metabolic and genome engineering efforts. Engineered functions that are burdensome to host cells and slow their replication are rapidly inactivated by mutations, and unplanned mutations with unpredictable effects often accumulate alongside designed changes in large-scale genome editing projects. We developed a directed evolution strategy, Periodic Reselection for Evolutionarily Reliable Variants (PResERV), to discover mutations that prolong the function of a burdensome DNA sequence in an engineered organism. Here, we used PResERV to isolate Escherichia coli cells that replicate ColE1-type plasmids with higher fidelity. We found mutations in DNA polymerase I and in RNase E that reduce plasmid mutation rates by 6- to 30-fold. The PResERV method implicitly selects to maintain the growth rate of host cells, and high plasmid copy numbers and gene expression levels are maintained in some of the evolved E. coli strains, indicating that it is possible to improve the genetic stability of cellular chassis without encountering trade-offs in other desirable performance characteristics. Utilizing these new antimutator E. coli and applying PResERV to other organisms in the future promises to prevent evolutionary failures and unpredictability to provide a more stable genetic foundation for synthetic biology.
Significance Organisms evolve and adapt via changes in their genomes that improve survival and reproduction in the context of their environment. Few experiments have examined how these genomic signatures of adaptation, which may favor mutations in certain genes or molecular pathways, vary across a set of similar environments that have both shared and distinctive characteristics. We sequenced complete genomes from 30 Escherichia coli lineages that evolved for 2,000 generations in one of five environments that differed only in the temperatures they experienced. Particular “signature” genes acquired mutations in these bacteria in response to selection imposed by specific temperature treatments. Thus, it is sometimes possible to predict aspects of the environment recently experienced by microbial populations from changes in their genome sequences.
Adaptation by natural selection depends on the rates, effects and interactions of many mutations, making it difficult to determine what proportion of mutations in an evolving lineage are beneficial. Here we analysed 264 complete genomes from 12 Escherichia coli populations to characterize their dynamics over 50,000 generations. The populations that retained the ancestral mutation rate support a model in which most fixed mutations are beneficial, the fraction of beneficial mutations declines as fitness rises, and neutral mutations accumulate at a constant rate. We also compared these populations to mutation-accumulation lines evolved under a bottlenecking regime that minimizes selection. Nonsynonymous mutations, intergenic mutations, insertions and deletions are overrepresented in the long-term populations, further supporting the inference that most mutations that reached high frequency were favoured by selection. These results illuminate the shifting balance of forces that govern genome evolution in populations adapting to a new environment.
New mutations leading to structural variation (SV) in genomes—in the form of mobile element insertions, large deletions, gene duplications, and other chromosomal rearrangements—can play a key role in microbial evolution. Yet, SV is considerably more difficult to predict from short-read genome resequencing data than single-nucleotide substitutions and indels (SN), so it is not yet routinely identified in studies that profile population-level genetic diversity over time in evolution experiments. We implemented an algorithm for detecting polymorphic SV as part of the breseq computational pipeline. This procedure examines split-read alignments, in which the two ends of a single sequencing read match disjoint locations in the reference genome, in order to detect structural variants and estimate their frequencies within a sample. We tested our algorithm using simulated Escherichia coli data and then applied it to 500- and 1000-generation population samples from the Lenski E. coli long-term evolution experiment (LTEE). Knowledge of genes that are targets of selection in the LTEE and mutations present in previously analyzed clonal isolates allowed us to evaluate the accuracy of our procedure. Overall, SV accounted for ~25% of the genetic diversity found in these samples. By profiling rare SV, we were able to identify many cases where alternative mutations in key genes transiently competed within a single population. We also found, unexpectedly, that mutations in two genes that rose to prominence at these early time points always went extinct in the long term. Because it is not limited by the base-calling error rate of the sequencing technology, our approach for identifying rare SV in whole-population samples may have a lower detection limit than similar predictions of SNs in these data sets. We anticipate that this functionality of breseq will be useful for providing a more complete picture of genome dynamics during evolution experiments with haploid microorganisms.
ABSTRACT Large-scale rearrangements may be important in evolution because they can alter chromosome organization and gene expression in ways not possible through point mutations. In a long-term evolution experiment, twelve Escherichia coli populations have been propagated in a glucose-limited environment for over 25 years. We used whole-genome mapping (optical mapping) combined with genome sequencing and PCR analysis to identify the large-scale chromosomal rearrangements in clones from each population after 40,000 generations. A total of 110 rearrangement events were detected, including 82 deletions, 19 inversions, and 9 duplications, with lineages having between 5 and 20 events. In three populations, successive rearrangements impacted particular regions. In five populations, rearrangements affected over a third of the chromosome. Most rearrangements involved recombination between insertion sequence (IS) elements, illustrating their importance in mediating genome plasticity. Two lines of evidence suggest that at least some of these rearrangements conferred higher fitness. First, parallel changes were observed across the independent populations, with ~65% of the rearrangements affecting the same loci in at least two populations. For example, the ribose-utilization operon and the manB-cpsG region were deleted in 12 and 10 populations, respectively, suggesting positive selection, and this inference was previously confirmed for the former case. Second, optical maps from clones sampled over time from one population showed that most rearrangements occurred early in the experiment, when fitness was increasing most rapidly. However, some rearrangements likely occur at high frequency and may have simply hitchhiked to fixation. In any case, large-scale rearrangements clearly influenced genomic evolution in these populations. IMPORTANCE Bacterial chromosomes are dynamic structures shaped by long histories of evolution. Among genomic changes, large-scale DNA rearrangements can have important effects on the presence, order, and expression of genes. Whole-genome sequencing that relies on short DNA reads cannot identify all large-scale rearrangements. Therefore, deciphering changes in the overall organization of genomes requires alternative methods, such as optical mapping. We analyzed the longest-running microbial evolution experiment (more than 25 years of evolution in the laboratory) by optical mapping, genome sequencing, and PCR analyses. We found multiple large genome rearrangements in all 12 independently evolving populations. In most cases, it is unclear whether these changes were beneficial themselves or, alternatively, hitchhiked to fixation with other beneficial mutations. In any case, many genome rearrangements accumulated over decades of evolution, providing these populations with genetic plasticity reminiscent of that observed in some pathogenic bacteria.
Background: Mutations that alter chromosomal structure play critical roles in evolution and disease, including in the origin of new lifestyles and pathogenic traits in microbes. Large-scale rearrangements in genomes are often mediated by recombination events involving new or existing copies of mobile genetic elements, recently duplicated genes, or other repetitive sequences. Most current software programs for predicting structural variation from short-read DNA resequencing data are intended primarily for use on human genomes. They typically disregard information in reads mapping to repeat sequences, and significant post-processing and manual examination of their output is often required to rule out false-positive predictions and precisely describe mutational events.Results: We have implemented an algorithm for identifying structural variation from DNA resequencing data as part of the breseq computational pipeline for predicting mutations in haploid microbial genomes. Our method evaluates the support for new sequence junctions present in a clonal sample from split-read alignments to a reference genome, including matches to repeat sequences. Then, it uses a statistical model of read coverage evenness to accept or reject these predictions. Finally, breseq combines predictions of new junctions and deleted chromosomal regions to output biologically relevant descriptions of mutations and their effects on genes. We demonstrate the performance of breseq on simulated Escherichia coli genomes with deletions generating unique breakpoint sequences, new insertions of mobile genetic elements, and deletions mediated by mobile elements. Then, we reanalyze data from an E. coli K-12 mutation accumulation evolution experiment in which structural variation was not previously identified. Transposon insertions and large-scale chromosomal changes detected by breseq account for similar to 25% of spontaneous mutations in this strain. In all cases, we find that breseq is able to reliably predict structural variation with modest read-depth coverage of the reference genome (>40-fold).Conclusions: Using breseq to predict structural variation should be useful for studies of microbial epidemiology, experimental evolution, synthetic biology, and genetics when a reference genome for a closely related strain is available. In these cases, breseq can discover mutations that may be responsible for important or unintended changes in genomes that might otherwise go undetected.
Next-generation DNA sequencing (NGS) can be used to reconstruct eco-evolutionary population dynamics and to identify the genetic basis of adaptation in laboratory evolution experiments. Here, we describe how to run the open-source breseq computational pipeline to identify and annotate genetic differences found in whole-genome and whole-population NGS data from haploid microbes where a high-quality reference genome is available. These methods can also be used to analyze mutants isolated in genetic screens and to detect unintended mutations that may occur during strain construction and genome editing.
Significance Unexpected evolutionary innovations that lead to qualitatively new traits may result from complex genetic and ecological interactions that develop over long timescales. In a 25-y evolution experiment with Escherichia coli , a rare metabolic innovation arose that allowed a previously untapped resource to be exploited. By dissecting the genetics of this trait using a recursive genomewide recombination and sequencing method (REGRES), we identified a key mutation that converts a rudimentary form of the innovation into a refined trait that confers a decisive competitive advantage. The effects of this mutation demonstrate how improvement of an emergent trait can be as important to its eventual success as earlier mutations or environmental conditions that may have been necessary for it to evolve in the first place.