Despite ongoing efforts to study CRISPR systems, the evolutionary origins giving rise to reprogrammable RNA-guided mechanisms remain poorly understood. Here, we describe an integrated sequence/structure evolutionary tracing approach to identify the ancestors of the RNA-targeting CRISPR-Cas13 system. We find that Cas13 likely evolved from AbiF, which is encoded by an abortive infection-linked gene that is stably associated with a conserved non-coding RNA (ncRNA). We further characterize a miniature Cas13, classified here as Cas13e, which serves as an evolutionary intermediate between AbiF and other known Cas13s. Despite this relationship, we show that their functions substantially differ. Whereas Cas13e is an RNA-guided RNA-targeting system, AbiF is a toxin-antitoxin (TA) system with an RNA antitoxin. We solve the structure of AbiF using cryoelectron microscopy (cryo-EM), revealing basic structural alterations that set Cas13s apart from AbiF. Finally, we map the key structural changes that enabled a non-guided TA system to evolve into an RNA-guided CRISPR system.
Supplementary file 1: BLASTN results for protospacers search This file contains the output of the BLASTN search of the spacers in the same genomic partition. See materials and methods for more details regarding the BLASTN search. All sequences between two adjacent repeats (less than 60bp apart) were retrieved and marked with the first repeat coordinates, spacer coordinates, and “Array” flag. For all single repeat-like sequences, 30bp up/downstream as potential spacers and marked with the repeat coordinates, spacer coordinates, and “SingleUpstream” or “SingleDownstream” flags.
The TnpB proteins are transposon-associated RNA-guided nucleases that are among the most abundant proteins encoded in bacterial and archaeal genomes, but whose functions in the transposon life cycle remain unknown. TnpB appears to be the evolutionary ancestor of Cas12, the effector nuclease of type V CRISPR-Cas systems. We performed a comprehensive census of TnpBs in archaeal and bacterial genomes and constructed a phylogenetic tree on which we mapped various features of these proteins. In multiple branches of the tree, the catalytic site of the TnpB nuclease is rearranged, demonstrating structural and probably biochemical malleability of this enzyme. We identified numerous cases of apparent recruitment of TnpB for other functions of which the most common is the evolution of type V CRISPR-Cas effectors on about 50 independent occasions. In many other cases of more radical exaptation, the catalytic site of the TnpB nuclease is apparently inactivated, suggesting a regulatory function, whereas in others, the activity appears to be retained, indicating that the recruited TnpB functions as a nuclease, for example, as a toxin. These findings demonstrate remarkable evolutionary malleability of the TnpB scaffold and provide extensive opportunities for further exploration of RNA-guided biological systems as well as multiple applications.
CRISPR- cas loci typically contain CRISPR arrays with unique spacers separating direct repeats. Spacers along with portions of adjacent repeats are transcribed and processed into CRISPR(cr) RNAs that target complementary sequences (protospacers) in mobile genetic elements, resulting in cleavage of the target DNA or RNA. Additional, standalone repeats in some CRISPR- cas loci produce distinct cr-like RNAs implicated in regulatory or other functions. We developed a computational pipeline to systematically predict crRNA-like elements by scanning for standalone repeat sequences that are conserved in closely related CRISPR- cas loci. Numerous crRNA-like elements were detected in diverse CRISPR-Cas systems, mostly, of type I, but also subtype V-A. Standalone repeats often form mini-arrays containing two repeat-like sequence separated by a spacer that is partially complementary to promoter regions of cas genes, in particular cas8 , or cargo genes located within CRISPR-Cas loci, such as toxins-antitoxins. We show experimentally that a mini-array from a type I-F1 CRISPR-Cas system functions as a regulatory guide. We also identified mini-arrays in bacteriophages that could abrogate CRISPR immunity by inhibiting effector expression. Thus, recruitment of CRISPR effectors for regulatory functions via spacers with partial complementarity to the target is a common feature of diverse CRISPR-Cas systems.
CRISPR-cas loci typically contain CRISPR arrays with unique spacers separating direct repeats. Spacers along with portions of adjacent repeats are transcribed and processed into CRISPR(cr) RNAs that target complementary sequences (protospacers) in mobile genetic elements, resulting in cleavage of the target DNA or RNA. Additional, standalone repeats in some CRISPR-cas loci produce distinct cr-like RNAs implicated in regulatory or other functions. We developed a computational pipeline to systematically predict crRNA-like elements by scanning for standalone repeat sequences that are conserved in closely related CRISPR-cas loci. Numerous crRNA-like elements were detected in diverse CRISPR-Cas systems, mostly, of type I, but also subtype V-A. Standalone repeats often form mini-arrays containing two repeat-like sequence separated by a spacer that is partially complementary to promoter regions of cas genes, in particular cas8, or cargo genes located within CRISPR-Cas loci, such as toxins-antitoxins. We show experimentally that a mini-array from a type I-F1 CRISPR-Cas system functions as a regulatory guide. We also identified mini-arrays in bacteriophages that could abrogate CRISPR immunity by inhibiting effector expression. Thus, recruitment of CRISPR effectors for regulatory functions via spacers with partial complementarity to the target is a common feature of diverse CRISPR-Cas systems.
aNational Center for Biotechnology Information, National Library of Medicine, Bethesda, Maryland, USA bBroad Institute of MIT and Harvard, Cambridge, Massachusetts, USA cHHMI, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA dMcGovern Institute for Brain Research, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA eDepartment of Brain and Cognitive Sciences, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA fDepartment of Biological Engineering, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA
CrAssphage is the most abundant human-associated virus and the founding member of a large group of bacteriophages, discovered in animal-associated and environmental metagenomes, that infect bacteria of the phylum Bacteroidetes. We analyze 4907 Circular Metagenome Assembled Genomes (cMAGs) of putative viruses from human gut microbiomes and identify nearly 600 genomes of crAss-like phages that account for nearly 87% of the DNA reads mapped to these cMAGs. Phylogenetic analysis of conserved genes demonstrates the monophyly of crAss-like phages, a putative virus order, and of 5 branches, potential families within that order, two of which have not been identified previously. The phage genomes in one of these families are almost twofold larger than the crAssphage genome (145-192 kilobases), with high density of self-splicing introns and inteins. Many crAss-like phages encode suppressor tRNAs that enable read-through of UGA or UAG stop-codons, mostly, in late phage genes. A distinct feature of the crAss-like phages is the recurrent switch of the phage DNA polymerase type between A and B families. Thus, comparative genomic analysis of the expanded assemblage of crAss-like phages reveals aspects of genome architecture and expression as well as phage biology that were not apparent from the previous work on phage genomics.
CRISPR-Cas systems provide RNA-guided adaptive immunity in prokaryotes. We report that the multisubunit CRISPR effector Cascade transcriptionally regulates a toxin-antitoxin RNA pair, CreTA. CreT (Cascade-repressed toxin) is a bacteriostatic RNA that sequesters the rare arginine tRNA(UCU) (transfer RNA with anticodon UCU). CreA is a CRISPR RNA-resembling antitoxin RNA, which requires Cas6 for maturation. The partial complementarity between CreA and the creT promoter directs Cascade to repress toxin transcription. Thus, CreA becomes antitoxic only in the presence of Cascade. In CreTA-deleted cells, cascade genes become susceptible to disruption by transposable elements. We uncover several CreTA analogs associated with diverse archaeal and bacterial CRISPR-cas loci. Thus, toxin-antitoxin RNA pairs can safeguard CRISPR immunity by making cells addicted to CRISPR-Cas, which highlights the multifunctionality of Cas proteins and the intricate mechanisms of CRISPR-Cas regulation.
Transposition is a major mechanism of horizontal gene mobility in prokar-yotes. However, exploration of the genes mobilized by transposons (cargo) is hampered by the difficulty in delineating integrated transposons from their surrounding genetic context. Here, we present a computational approach that allowed us to identify the boundaries of 6,549 Tn7-like transposons. We found that 96% of these transposons carry at least one cargo gene. Delineation of distinct communities in a gene-sharing network demonstrates how transposons function as a conduit of genes between phylo-genetically distant hosts. Comparative analysis of the cargo genes reveals significant enrichment of mobile genetic elements (MGEs) nested within Tn7-like transposons, such as insertion sequences and toxin-antitoxin modules, and of genes involved in recombination, anti-MGE defense, and antibiotic resistance. More unexpectedly, cargo also includes genes encoding central carbon metabolism enzymes. Twenty-two Tn7-like transposons carry both an anti-MGE defense system and antibiotic resistance genes, illustrating how bacteria can overcome these combined pressures upon acquisition of a single transposon. This work substantially expands the distribution of Tn7-like transpo-sons, defines their evolutionary relationships, and provides a large-scale functional clas-sification of prokaryotic genes mobilized by transposition. IMPORTANCE Transposons are major vehicles of horizontal gene transfer that, in addi-tion to genes directly involved in transposition, carry cargo genes. However, charac-terization of these genes is hampered by the difficulty of identification of transposon boundaries. We developed a computational approach for detecting transposon ends and applied it to perform a comprehensive census of the cargo genes of Tn7-like transposons, a large class of bacterial mobile genetic elements (MGE), many of which employ a unique, CRISPR-mediated mechanism of site-specific transposition. The cargo genes encompass a striking diversity of MGE, defense, and antibiotic resistance systems. Unexpectedly, we also identified cargo genes encoding metabolic enzymes. Thus, Tn7-like transposons mobilize a vast repertoire of genes that can have multiple effects on the host bacteria.
CRISPR-Cas are adaptive immune systems that degrade foreign genetic elements in archaea and bacteria. In carrying out their immune functions, CRISPR-Cas systems heavily rely on RNA components. These CRISPR (cr) RNAs are repeat-spacer units that are produced by processing of pre-crRNA, the transcript of CRISPR arrays, and guide Cas protein(s) to the cognate invading nucleic acids, enabling their destruction. Several bioinformatics tools have been developed to detect CRISPR arrays based solely on DNA sequences, but all these tools employ the same strategy of looking for repetitive patterns, which might correspond to CRISPR array repeats. The identified patterns are evaluated using a fixed, built-in scoring function, and arrays exceeding a cut-off value are reported. Here, we instead introduce a data-driven approach that uses machine learning to detect and differentiate true CRISPR arrays from false ones based on several features. Our CRISPR detection tool, CRISPRidentify, performs three steps: detection, feature extraction and classification based on manually curated sets of positive and negative examples of CRISPR arrays. The identified CRISPR arrays are then reported to the user accompanied by detailed annotation. We demonstrate that our approach identifies not only previously detected CRISPR arrays, but also CRISPR array candidates not detected by other tools. Compared to other methods, our tool has a drastically reduced false positive rate. In contrast to the existing tools, our approach not only provides the user with the basic statistics on the identified CRISPR arrays but also produces a certainty score as a practical measure of the likelihood that a given genomic region is a CRISPR array.
Type IV CRISPR-Cas are a distinct variety of highly derived CRISPR-Cas systems that appear to have evolved from type III systems through the loss of the target-cleaving nuclease and partial deterioration of the large subunit of the effector complex. All known type IV CRISPR-Cas systems are encoded on plasmids, integrative and conjugative elements (ICEs), or prophages, and are thought to contribute to competition between these elements, although the mechanistic details of their function remain unknown. There is a clear parallel between the compositions and likely origin of type IV and type I systems recruited by Tn7-like transposons and mediating RNA-guided transposition. We investigated the diversity and evolutionary relationships of type IV systems, with a focus on those in Acidithiobacillia, where this variety of CRISPR is particularly abundant and always found on ICEs. Our analysis revealed remarkable evolutionary plasticity of type IV CRISPR-Cas systems, with adaptation and ancillary genes originating from different ancestral CRISPR-Cas varieties, and extensive gene shuffling within the type IV loci. The adaptation module and the CRISPR array apparently were lost in the type IV ancestor but were subsequently recaptured by type IV systems on several independent occasions. We demonstrate a high level of heterogeneity among the repeats with type IV CRISPR arrays, which far exceed the heterogeneity of any other known CRISPR repeats and suggest a unique adaptation mechanism. The spacers in the type IV arrays, for which protospacers could be identified, match plasmid genes, in particular those encoding the conjugation apparatus components. Both the biochemical mechanism of type IV CRISPR-Cas function and their role in the competition among mobile genetic elements remain to be investigated.
Background: Double-stranded DNA bacteriophages (dsDNA phages) play pivotal roles in structuring human gut microbiomes; yet, the gut virome is far from being fully characterized, and additional groups of phages, including highly abundant ones, continue to be discovered by metagenome mining. A multilevel framework for taxonomic classification of viruses was recently adopted, facilitating the classification of phages into evolutionary informative taxonomic units based on hallmark genes. Together with advanced approaches for sequence assembly and powerful methods of sequence analysis, this revised framework offers the opportunity to discover and classify unknown phage taxa in the human gut. Results: A search of human gut metagenomes for circular contigs encoding phage hallmark genes resulted in the identification of 3738 apparently complete phage genomes that represent 451 putative genera. Several of these phage genera are only distantly related to previously identified phages and are likely to found new families. Two of the candidate families, "Flandersviridae" and "Quimbyviridae", include some of the most common and abundant members of the human gut virome that infect Bacteroides, Parabacteroides, and Prevotella. The third proposed family, "Gratiaviridae," consists of less abundant phages that are distantly related to the families Autographiviridae, Drexlerviridae, and Chaseviridae. Analysis of CRISPR spacers indicates that phages of all three putative families infect bacteria of the phylum Bacteroidetes. Comparative genomic analysis of the three candidate phage families revealed features without precedent in phage genomes. Some "Quimbyviridae" phages possess Diversity-Generating Retroelements (DGRs) that generate hypervariable target genes nested within defense-related genes, whereas the previously known targets of phage-encoded DGRs are structural genes. Several "Flandersviridae" phages encode enzymes of the isoprenoid pathway, a lipid biosynthesis pathway that so far has not been known to be manipulated by phages. The "Gratiaviridae" phages encode a HipA-family protein kinase and glycosyltransferase, suggesting these phages modify the host cell wall, preventing superinfection by other phages. Hundreds of phages in these three and other families are shown to encode catalases and iron-sequestering enzymes that can be predicted to enhance cellular tolerance to reactive oxygen species. Conclusions: Analysis of phage genomes identified in whole-community human gut metagenomes resulted in the delineation of at least three new candidate families of Caudovirales and revealed diverse putative mechanisms underlying phage-host interactions in the human gut. Addition of these phylogenetically classified, diverse, and distinct phages to public databases will facilitate taxonomic decomposition and functional characterization of human gut viromes.
CRISPR-Cas defense systems opened up the field of genome editing due to the ease with which effector Cas nucleases can be programmed with guide RNAs to access desirable genomic sites. Type II-A SpCas9 from Streptococcus pyogenes was the first Cas9 nuclease used for genome editing and it remains the most popular enzyme of its class. Nevertheless, SpCas9 has some drawbacks including a relatively large size and restriction to targets flanked by an 'NGG' PAM sequence. The more compact Type II-C Cas9 orthologs can help to overcome the size limitation of SpCas9. Yet, only a few Type II-C nucleases were fully characterized to date. Here, we characterized two Cas9 II-C orthologs, DfCas9 from Defluviimonas sp.20V17and PpCas9 from Pasteurella pneumotropica. Both DfCas9 and PpCas9 cleave DNA in vitro and have novel PAM requirements. Unlike DfCas9, the PpCas9 nuclease is active in human cells. This small nuclease requires an 'NNNNRTT' PAM orthogonal to that of SpCas9 and thus potentially can broaden the range of Cas9 applications in biomedicine and biotechnology.
The CRISPR-Cas are adaptive bacterial and archaeal immunity systems that have been harnessed for the development of powerful genome editing and engineering tools. In the incessant host-parasite arms race, viruses evolved multiple anti-defense mechanisms including diverse anti-CRISPR proteins (Acrs) that specifically inhibit CRISPR-Cas and therefore have enormous potential for application as modulators of genome editing tools. Most Acrs are small and highly variable proteins which makes their bioinformatic prediction a formidable task. We present a machine-learning approach for comprehensive Acr prediction. The model shows high predictive power when tested against an unseen test set and was employed to predict 2,500 candidate Acr families. Experimental validation of top candidates revealed two unknown Acrs (AcrIC9, IC10) and three other top candidates were coincidentally identified and found to possess anti-CRISPR activity. These results substantially expand the repertoire of predicted Acrs and provide a resource for experimental Acr discovery.
Bacteria and archaea evolve under constant pressure from numerous, diverse viruses and thus have evolved multiple defense systems. The CRISPR-Cas are adaptive immunity systems that have been harnessed for the development of the new generation of genome editing and engineering tools. In the incessant host-parasite arms race, viruses evolved multiple anti-defense mechanisms including numerous, diverse anti-CRISPR proteins (Acrs) that can inhibit CRISPR-Cas and therefore have enormous potential for application as modulators of genome editing tools. Most Acrs are small, highly variable proteins which makes their prediction a formidable task. We developed a machine learning approach for comprehensive Acr prediction. The model showed high predictive power when tested against an unseen test set that included several families of recently discovered Acrs and was employed to predict 2,500 novel candidate Acr families. An examination of the top candidates confirms that they possess typical Acr features. One of the top candidates was independently tested and found to possess anti-CRISPR activity (AcrIIA12). We provide a web resource ( http://acrcatalog.pythonanywhere.com/ ) to access the predicted Acrs sequences and annotation. The results of this analysis expand the repertoire of predicted Acrs almost by two orders of magnitude and provide a rich resource for experimental Acr discovery.
CrAssphage is the most abundant virus identified in the human gut virome and the founding member of a large group of bacteriophages that infect bacteria of the phylum Bacteroidetes and have been discovered by metagenomics of both animal-associated and environmental habitats. By analysis of circular contigs from human gut microbiomes, we identified nearly 600 genomes of crAss-like phages. Phylogenetic analysis of conserved genes demonstrates the monophyly of crAss-like phages, which can be expected to become a new order of viruses, and of 5 distinct branches, likely, families within that order. Two of these putative families have not been identified previously. The phages in one of these groups have large genomes (145-192 kilobases) and contain an unprecedented high density of self-splicing introns and inteins. Many crAss-like phages encode suppressor tRNAs that enable readthrough of UGA or UAG stop-codons, mostly, in late phage genes, which could represent a distinct anti-defense strategy. Another putative anti-defense mechanism that might target an unknown defense system in Bacteroidetes inhibiting phage DNA replication involves multiple switches of the phage DNA polymerase type between A and B families. Thus, comparative genomic analysis of the expanded assemblage of crAss-like phages reveals several unusual features of genome architecture and expression as well as phage biology that were not apparent from the previous crAssphage analyses.### Competing Interest StatementThe authors have declared no competing interest.
CRISPR-Cas systems typically consist of a CRISPR array and cas genes that are organized in one or more operons. However, a substantial fraction of CRISPR arrays are not adjacent to cas genes. Definitive identification of such isolated CRISPR arrays runs into the problem of false-positives, with unrelated types of repetitive sequences mimicking CRISPR. We developed a computational pipeline to eliminate false CRISPR predictions and found that up to 25% of the CRISPR arrays in complete bacterial and archaeal genomes are located away from cas genes. Most of the repeats in these isolated arrays are identical to repeats in cas-adjacent CRISPR arrays in the same or closely related genomes, indicating an evolutionary relationship between isolated arrays and arrays in typical CRISPR-cas loci. The spacers in isolated CRISPR arrays show nearly as many matches to viral genomes as spacers from complete CRISPR-cas loci, suggesting that the isolated arrays were either functionally active recently or continue to function. Reconstruction of evolutionary events in closely related bacterial genomes suggests three routes of evolution of isolated CRISPR arrays: (1) loss of cas genes in a CRISPR-cas locus, (2) de novo generation of arrays from off-target spacer integration into sequences resembling the corresponding repeats, and (3) transfer by mobile genetic elements. Both combination of de novo emerging arrays with cas genes and regain of cas genes by isolated arrays via recombination likely contribute to functional diversification in CRISPR-Cas evolution.
The principal function of archaeal and bacterial CRISPR-Cas systems is antivirus adaptive immunity. However, recent genome analyses identified a variety of derived CRISPR-Cas variants at least some of which appear to perform different functions. Here, we describe a unique repertoire of CRISPR-Cas-related systems that we discovered by searching archaeal metagenome-assemble genomes of the Asgard superphylum. Several of these variants contain extremely diverged homologs of Cas1, the integrase involved in CRISPR adaptation as well as casposon transposition. Strikingly, the diversity of Cas1 in Asgard archaea alone is greater than that detected so far among the rest of archaea and bacteria. The Asgard CRISPR-Cas derivatives also encode distinct forms of Cas4, Cas5, and Cas7 proteins, and/or additional nucleases. Some of these systems are predicted to perform defense functions, but possibly not programmable ones, whereas others are likely to represent previously unknown mobile genetic elements.
CRISPR arrays contain spacers, some of which are homologous to genome segments of viruses and other parasitic genetic elements and are employed as portion of guide RNAs to recognize and specifically inactivate the target genomes. However, the fraction of the spacers in sequenced CRISPR arrays that reliably match protospacer sequences in genomic databases is small, leaving the question of the origin(s) open for the great majority of the spacers. Here, we extend the spacer analysis by examining the distribution of partial matches (matching k -mers) between spacers and genomes of viruses infecting the given host as well as the host genomes themselves. The results indicate that most of the spacers originate from the host-specific viromes, whereas self-targeting is strongly selected against. However, we present evidence that the vast majority of the viruses comprising the viromes currently remain unknown although they are likely to be related to identified viruses.