The Saccharomyces Genome Database (SGD) is one of the longest-running and most consequential biological databases in the world. Founded in the early 1990s at Stanford University under the visionary leadership of David Botstein and developed under the long-term technical direction of J. Michael Cherry, SGD has served for more than three decades not only as the authoritative knowledge center for the budding yeast Saccharomyces cerevisiae, but also as the source for much of the fundamentals of eukaryotic biology. This history traces the arc of a remarkable intellectual and scientific project: beginning with the challenge of building the very first integrated eukaryotic genome database and evolving across thirty years into a global knowledge hub for genetics, functional genomics, and human disease research. The history is organized chronologically, with each section highlighting the central ideas, technical developments, and concrete accomplishments of that period.
The Candida Genome Database (CGD; www.candidagenome.org ) is both a model organism database and a fungal pathogen database. As a model organism database, CGD stores data for Candida albicans , which serves as a model species both for other Candida spp. and for non- Candida fungi that form biofilms and undergo routine morphogenic switching. As a fungal pathogen database, CGD now hosts locus pages for six species of the best-studied pathogenic fungi in the Candida group. Pathogenic Candida species have become increasingly drug resistant and there is thus a pressing need for research into basic Candida biology, epidemiology, phylogeny, and potential new antifungals, as well as a single location where all of the available data are collected, curated, and made easily searchable. CGD curates the gene-based Candida experimental literature in real time, extracting, organizing and standardizing gene annotations. CGD also links clinical data on disease to relevant Literature Topics to improve searchability for clinical researchers. Because CGD curates the literature for multiple species and most research focuses on aspects related to pathogenicity, we focus our curation efforts on assigning Literature Topic tags, collecting detailed mutant phenotype data, and assigning controlled Gene Ontology terms with accompanying evidence codes. Our Summary pages for each locus include the primary name and all aliases for that locus, a description of the gene and/or gene product, detailed ortholog information with links, a synteny view, a JBrowse window with a visual view of the gene on its chromosome, links to Phenotype, Gene Ontology, Interactions, and Expression pages, as well as sequence information, references cited on the summary page itself, and any locus notes. The database also serves as a community hub, where we link to various types of reference material of relevance to Candida researchers, including colleague information, news, and notice of upcoming meetings. We routinely survey the community to learn how the field is evolving and how needs may have changed. Here we describe CGD's new modern web interface and multiple new tools that have been added in the last 6 months, allowing, among other things, users to better understand the available expression data for a locus and seamlessly switch between species for a given locus.
We introduce SCRIVENER (sequential conjugation and recombination for in vivo elongation of nucleotides with low errors), an in vivo DNA assembly platform that streamlines and scales DNA engineering. SCRIVENER combines bacterial conjugation, in vivo DNA cutting, and homologous recombination to stitch DNA blocks together by mating E. coli in large arrays or pools. This approach is simpler, cheaper, and higher throughput than methods requiring DNA to be moved in and out of cells. We performed over 5,000 assemblies with 2 to 19 blocks (240 bp-12 kb) and assembled constructs up to 81 kb with high fidelity. Most errors are deletions between long repeats, but SCRIVENER minimizes their impact by enabling high-replication assembly and sequence verification at a nominal additional cost per replicate. The platform enables combinatorial library construction and DNA block reuse without PCR and is therefore a powerful tool to accelerate DNA design-build-test-learn cycles. A record of this paper's transparent peer review process is included in the supplemental information.
Medications administered over long durations, such as phenothiazine antipsychotics, accumulate in the gut at concentrations that affect microbial growth. However, the bacterial features influencing sensitivity to these non-antibiotics remain poorly understood. Bacterial capsular polysaccharides (CPSs) typically confer protection against environmental stressors, including chemical, viral, and immunological pressures within the gut. But their roles under non-antibiotic drug pressure are unknown. Here, we show that the K5 CPS of Escherichia coli Nissle 1917 ( EcN ) sensitizes it to thioridazine (TDZ) and related antipsychotics. Among a panel of E. coli strains grown in minimal medium, EcN exhibited the highest sensitivity to TDZ. Experimentally evolving EcN under gut-relevant TDZ concentrations selected for resistant populations with convergent variants affecting the CPS locus, and TDZ-resistant clones correspondingly had lower CPS expression than their drug-susceptible counterparts. Genetic, transcriptomic, and phenotypic analyses confirmed that the K5 CPS enhances, rather than mitigates, TDZ sensitivity. These findings demonstrate that canonically protective surface structures can become vulnerabilities under non-antibiotic pharmaceutical pressure. Human medications may therefore inadvertently shape the expression and evolution of bacterial surface structures in the gastrointestinal tract, challenging presumptions of CPS-mediated environmental protection.
The Candida Genome Database (CGD; www.candidagenome.org) is both a model organism database and a fungal pathogen database. As a model organism database, CGD stores data for Candida albicans , which serves as a model species both for other Candida spp. and for non- Candida fungi that form biofilms and undergo routine morphogenic switching. As a fungal pathogen database, CGD now hosts locus pages for six species of the best-studied pathogenic fungi in the Candida group. Pathogenic Candida species have become increasingly drug resistant and there is thus a pressing need for research into basic Candida biology, epidemiology, phylogeny, and potential new antifungals, as well as a single location where all of the available data are collected, curated, and made easily searchable. CGD curates the gene-based Candida experimental literature in real time, extracting, organizing and standardizing gene annotations. CGD also links clinical data on disease to relevant Literature Topics to improve searchability for clinical researchers. Because CGD curates the literature for multiple species and most research focuses on aspects related to pathogenicity, we focus our curation efforts on assigning Literature Topic tags, collecting detailed mutant phenotype data, and assigning controlled Gene Ontology terms with accompanying evidence codes. Our Summary pages for each locus include the primary name and all aliases for that locus, a description of the gene and/or gene product, detailed ortholog information with links, a synteny view, a JBrowse window with a visual view of the gene on its chromosome, links to Phenotype, Gene Ontology, Interactions, and Expression pages, as well as sequence information, references cited on the summary page itself, and any locus notes. The database also serves as a community hub, where we link to various types of reference material of relevance to Candida researchers, including colleague information, news, and notice of upcoming meetings. We routinely survey the community to learn how the field is evolving and how needs may have changed. Here we describe CGDs new modern web interface and multiple new tools that have been added in the last 6 months, allowing, among other things, users to better understand the available expression data for a locus and seamlessly switch between species for a given locus.
Stationary phase in yeast and other microorganisms begins when a limiting nutrient in the environment is exhausted and cell division ceases. Most cells subsequently enter quiescence and lose viability. In spent media, without metabolic byproducts being diluted, cellular processes can modify the environment and cause the relative growth rates of different genotypes to vary over the course of stationary phase. In this work we experimentally evolve S. cerevisiae in batch culture, varying the time spent in stationary phase between growth cycles. We measure the relative fitness of the resulting adaptive clones across a range of environments: with different amounts of time in stationary phase and in two different carbon sources. By comparing the inferred performance (relative growth rate during a period of the growth cycle) of a mutant to that of its ancestor, we can estimate the effects of each observed mutation on performance during various phases of growth. We show that when an adaptive mutation emerges in growth cycles that include a stationary phase, its effect on stationary phase performance is largely independent of the type of carbon source provided. However, for the same group of mutants, mutational effects on performance in early stationary phase are negatively correlated with those effects in late stationary phase, suggesting a trade-off. We also show that increased intervals of stationary phase result in larger fitness effects of adaptive mutations and distinct routes of adaptation. Together, these results demonstrate that stationary phase consists of more than one distinct fitness-related phenotype, and that the phenotypes that allow for high performance in the first few days of stationary phase trade off with those that allow for high performance in later stationary phase.
Killer yeasts, such as the K1 killer strain of S. cerevisiae, express a secreted anti-competitive toxin whose production and propagation require the presence of two vertically-transmitted dsRNA viruses. In sensitive cells lacking killer virus infection, toxin binding to the cell wall results in ion pore formation, disruption of osmotic homeostasis, and cell death. However, the exact mechanism(s) of K1 toxin killing activity, how killer yeasts are immune to their own toxin, and which factors could influence adaptation and resistance to K1 toxin within formerly sensitive populations are still unknown. Here, we describe the state of knowledge about K1 killer toxin, including current models of toxin processing and killing activity, and a summary of known modifiers of K1 toxin immunity and resistance. In addition, we discuss two key signaling pathways, HOG (high osmolarity glycerol) and CWI (cell wall integrity), whose involvement in an adaptive response to K1 killer toxin in sensitive cells has been previously documented but requires further study. As both host-virus and sensitive-killer competition have been documented in killer systems like K1, further characterization of K1 killer yeasts may provide a useful model system for study of both intracellular genetic conflict and counter-adaptation between competing sensitive and killer populations.
During mitosis, stable but dynamic interactions between the centromere DNA and kinetochore complex enable accurate and efficient chromosome segregation. Even though many proteins of the kinetochore are highly conserved, centromeres are among the fastest evolving regions within a genome, showing extensive variation even on short evolutionary timescales. Here, we sought to understand how new types of centromeres emerge and reach fixation by mapping centromere evolution across 138 budding yeast species and over 2,500 natural strain isolates. We show that new centromeres spread progressively via drift and subsequent selection, and that the kinetochore interface, which is evolving slowly in relative terms, determines which new centromere variants are tolerated. Together, our findings provide insight into the evolutionary constraints and trajectories shaping centromere evolution. ### Competing Interest Statement The authors have declared no competing interest.
The Candida Genome Database (CGD; www.candidagenome.org) is unique in being both a model organism database and a fungal pathogen database. As a fungal pathogen database, CGD hosts locus pages for 5 species of the best-studied pathogenic fungi in the Candida group. As a model organism database, the species Candida albicans serves as a model both for other Candida spp. and for non-Candida fungi that form biofilms and undergo routine morphogenic switching from the planktonic form to the filamentous form, which is not done by other model yeasts. As pathogenic Candida species have become increasingly drug resistant, the high lethality of invasive candidiasis in immunocompromised people is increasingly alarming. There is a pressing need for additional research into basic Candida biology, epidemiology and phylogeny, and potential new antifungals. CGD serves the needs of this diverse research community by curating the entire gene-based Candida experimental literature as it is published, extracting, organizing, and standardizing gene annotations. Gene pages were added for the species Candida auris, recent outbreaks of which have been labeled an "urgent" threat. Most recently, we have begun linking clinical data on disease to relevant Literature Topics to improve searchability for clinical researchers. Because CGD curates for multiple species and most research focuses on aspects related to pathogenicity, we focus our curation efforts on assigning Literature Topic tags, collecting detailed mutant phenotype data, and assigning controlled Gene Ontology terms with accompanying evidence codes. Our Summary pages for each feature include the primary name and all aliases for that locus, a description of the gene and/or gene product, detailed ortholog information with links, a JBrowse window with a visual view of the gene on its chromosome, summarized phenotype, Gene Ontology, and sequence information, references cited on the summary page itself, and any locus notes. The database serves as a community hub, where we link to various types of reference material of relevance to Candida researchers, including colleague information, news, and notice of upcoming meetings. We routinely survey the community to learn how the field is evolving and how needs may have changed. For example, we asked our users which species we should next add to CGD, and the clear answer was Candida tropicalis. A key future challenge is management of the flood of high-throughput expression data to make it as useful as possible to as many researchers as possible. The central challenge for any community database is to turn data into knowledge, which the community can access, use, and build upon.
DNA can be engineered to produce new biologics, gene therapies, and cellular therapies, and to reprogram organisms. Having the ability to engineer DNA at scale can accelerate the development of these applications. Existing technologies excel at short oligonucleotide synthesis by chemical or enzymatic methods (up to 2000 bp) and intermediate-size DNA assembly (up to 5-7 kb). Yet synthesizing sequence-validated longer DNA (>10 kb) and/or constructing highly complex combinatorial DNA libraries at scale remains a significant challenge, due largely to technical and cost barriers. Inspired by recent studies on an in vivo DNA processing platform for megabase-long DNA assembly and on high-throughput sequence verification, we discuss how these platforms may be used to achieve DNA engineering at scale.
The prototypic crAssphage (Carjivirus communis) is an abundant, prevalent, and persistent human gut bacteriophage, yet it remains uncultured and its lifestyle uncharacterized. C. communis does not readily plaque, suggesting a largely non-lytic lifestyle. Here, we find that C. communis is a linear phage-plasmid that stably persists extrachromosomally within its host. Plasmid and phage-related genes are transcribed, and multiple putative replication origins may initiate replication for multiple lifestyles and genome conformations, including both circular and linear formations. Leveraging these findings, we use a plaque-free culturing approach to measure C. communis replication on prevalent gut bacteria, notably Phocaeicola vulgatus, P. dorei, and Bacteroides stercoris, revealing a broad host range. C. communis persists without causing major cell lysis events or integrating into host chromosomes. Taken together, C. communis' ability to switch between phage and plasmid lifestyles within a wide range of hosts may contribute to its widespread presence in human gut microbiomes.
The fitness of a genotype is defined as its lifetime reproductive success, with fitness itself being a composite trait likely dependent on many underlying phenotypes. Measuring fitness is important for understanding how alteration of different cellular components affects a cell’s ability to reproduce. Here, we describe an improved approach, implemented in Python, for estimating fitness in high throughput via pooled competition assays.
During mitosis, stable but dynamic interactions between centromere DNA and the kinetochore complex enable accurate and efficient chromosome segregation. Even though many proteins of the kinetochore are highly conserved1,2, centromeres are among the fastest evolving regions in a genome3,4, showing extensive variation even on short evolutionary timescales. Here we sought to understand how organisms evolve completely new sets of centromeres that still effectively engage with the kinetochore machinery by identifying and tracking thousands of centromeres across two major fungal clades, including more than 2,500 natural strain isolates and representing over 1,000 million years of evolution. We show that new centromeres spread progressively via drift and subsequent selection and that the kinetochore, which is evolving slowly in relative terms, appears to act as a filter to determine which new centromere variants are tolerated. Together, our findings provide insight into the evolutionary constraints and trajectories shaping centromere evolution.
Evolution of microbes under laboratory selection produces genetically diverse populations, owing to the continuous input of mutations and to competition among lineages. Whole-genome whole-population sequencing makes it possible to identify mutations arising in such populations, to use them to discern functional modules where adaptation occurs, and then map gene structure–function relationships. Here, we report on the use of this approach, adaptive genetics, to discover targets of selection and the mutational consequences thereof in E. coli evolving under chronic nutrient limitation. Replicate bacterial populations were cultured for ≥ 300 generations in glucose limited chemostats and sequenced every 50 generations at 1000X-coverage, enabling identification of mutations that rose to ≥ 1
Budding yeast (Saccharomyces cerevisiae) is the most extensively characterized eukaryotic model organism and has long been used to gain insight into the fundamentals of genetics, cellular biology, and the functions of specific genes and proteins. The Saccharomyces Genome Database (SGD) is a scientific resource that provides information about the genome and biology of S. cerevisiae. For more than 30 years, SGD has maintained the genetic nomenclature, chromosome maps, and functional annotation for budding yeast along with search and analysis tools to explore these data. Here, we describe recent updates at SGD, including the 2 most recent reference genome annotation updates, expanded biochemical pathway representation, changes to SGD search and data files, and other enhancements to the SGD website and user interface. These activities are part of our continuing effort to promote insights gained from yeast to enable the discovery of functional relationships between sequence and gene products in fungi and higher eukaryotes.
Sequence verification of plasmid DNA is critical for many cloning and molecular biology workflows. To leverage high-throughput sequencing, several methods have been developed that add a unique DNA barcode to individual samples prior to pooling and sequencing. However, these methods require an individual plasmid extraction and/or in vitro barcoding reaction for each sample processed, limiting throughput and adding cost. Here, we develop an arrayed in vivo plasmid barcoding platform that enables pooled plasmid extraction and library preparation for Oxford Nanopore sequencing. This method has a high accuracy and recovery rate, and greatly increases throughput and reduces cost relative to other plasmid barcoding methods or Sanger sequencing. We use in vivo barcoding to sequence verify >45,000 plasmids and show that the method can be used to transform error-containing dispersed plasmid pools into sequence-perfect arrays or well-balanced pools. In vivo barcoding does not require any specialized equipment beyond a low-overhead Oxford Nanopore sequencer, enabling most labs to flexibly process hundreds to thousands of plasmids in parallel.
The prototypic crAssphage (Carjivirus communis) is one of the most abundant, prevalent, and persistent gut bacteriophages, yet it remains uncultured and its lifestyle uncharacterized. For the last decade, crAssphage has escaped plaque-dependent culturing efforts, leading us to investigate alternative lifestyles that might explain its widespread success. Through genomic analyses and culturing, we find that crAssphage uses a phage-plasmid lifestyle to persist extrachromosomally. Plasmid-related genes are more highly expressed than those implicated in phage maintenance. Leveraging this finding, we use a plaque-free culturing approach to measure crAssphage replication in culture with Phocaeicola vulgatus, Phocaeicola dorei, and Bacteroides stercoris, revealing a broad host range. We demonstrate that crAssphage persists with its hosts in culture without causing major cell lysis events or integrating into host chromosomes. The ability to switch between phage and plasmid lifestyles within a wide range of hosts contributes to the prolific nature of crAssphage in the human gut microbiome.
Evolution by natural selection is expected to be a slow and gradual process. In particular, the mutations that drive evolution are predicted to be small and modular, incrementally improving a small number of traits. However, adaptive mutations identified early in microbial evolution experiments, cancer, and other systems often provide substantial fitness gains and pleiotropically improve multiple traits at once. We asked whether such pleiotropically adaptive mutations are common throughout adaptation or are instead a rare feature of early steps in evolution that tend to target key signaling pathways. To do so, we conducted barcoded second-step evolution experiments initiated from 5 first-step mutations identified from a prior yeast evolution experiment. We then isolated hundreds of second-step mutations from these evolution experiments, measured their fitness and performance in several growth phases, and conducted whole genome sequencing of the second-step clones. Here, we found that while the vast majority of mutants isolated from the first-step of evolution in this condition show patterns of pleiotropic adaptation—improving both performance in fermentation and respiration growth phases—second-step mutations show a shift towards modular adaptation, mostly improving respiration performance and only rarely improving fermentation performance. We also identified a shift in the molecular basis of adaptation from genes in cellular signaling pathways towards genes involved in respiration and mitochondrial function. Our results suggest that the genes in cellular signaling pathways may be more likely to provide large, adaptively pleiotropic benefits to the organism due to their ability to coherently affect many phenotypes at once. As such, these genes may serve as the source of pleiotropic adaptation in the early stages of evolution, and once these become exhausted, organisms then adapt more gradually, acquiring smaller, more modular mutations.
Candida glabrata is a fungal microbe associated with multiple vertebrate microbiomes and their terrestrial environments. In humans, the species has emerged as an opportunistic pathogen that now ranks as the second- leading cause of candidiasis in Europe and North America (Beardsley et al . Med Mycol 2024, 62). People at highest risk of infection include the elderly, immunocompromised individuals and/or long- term residents of hospital and assisted- living facilities. C. glabrata is intrinsically drug- resistant, metabolically versatile and able to avoid detection by the immune system. Analyses of its 12.3 Mb genome indicate a stable pangenome Marcet- Houben et al . ( BMC Biol 2022, 20) and phylogenetic affinity with Saccharomyces cerevisiae. Recent phylogenetic analyses suggest reclassifying C. glabrata as Nakaseomyces glabratus Lakashima and Sugita ( Med Mycol J 2022, 63: 119-132).
The eukaryotic cell division machinery must rapidly and reproducibly duplicate and partition the cell's chromosomes in a carefully coordinated process. However, chromosome numbers vary dramatically between genomes, even on short evolutionary timescales. We sought to understand how the mitotic machinery senses and responds to karyotypic changes by using a series of budding yeast strains in which the native chromosomes have been successively fused. Using a combination of cell biological profiling, genetic engineering and experimental evolution, we show that chromosome fusions are well tolerated up until a critical point. Cells with fewer than five centromeres lack the necessary number of kinetochore-microtubule attachments needed to counter outward forces in the metaphase spindle, triggering the spindle assembly checkpoint and prolonging metaphase. Our findings demonstrate that spindle architecture is a constraining factor for karyotype evolution. Helsen et al. use experimental evolution and chromosome engineering to probe the link between karyotype changes and the cell division machinery. They conclude that spindle organization dictates the available trajectories for karyotype evolution.