The endosomal sorting complexes required for transport (ESCRT) is a membrane-remodelling machinery that mediates membrane constriction and scission. Most enveloped eukaryotic viruses exploit host ESCRT during membrane budding, but none is known to encode ESCRT components. Here, we discovered homologues of two core ESCRT components, ESCRT-III and ATPase Vps4, in mirusviruses and some viruses in the phylum Nucleocytoviricota that infect unicellular eukaryotes. Mirusviruses seem to have hijacked host ESCRT machinery early in their evolution.
Abstract Recent extensive metatranscriptome mining vastly expanded the range of apparently covalently closed circular (ccc) RNA replicons. A notable family of such replicons is Obelisks, ~1 kilobase (kb) cccRNAs encoding a protein with a unique fold, Oblin-1, and detected in diverse metatranscriptomes. To identify potential cccRNAs in a sequence similarity–independent manner, we adopt the Fragmented and primer-Ligated DsRNA Sequencing (FLDS) method to selectively sequence double-stranded (ds) RNAs, replicative intermediates of RNA replicons. We focus on candidates with predicted extensive intramolecular base-pairing, a hallmark of viroid-like elements. Using FLDS, we explore metatranscriptomes from acidic hot springs in Japan and discover a distinct family of Obelisks apparently associated with thermoacidophilic bacteria (Hot spring Obelisks, HsObs). Despite lacking sequence similarity to known Oblins, HsObs share key features, including ~1 kb genome size, rod-like RNA secondary structure, and the predicted fold of the encoded protein, HsOblin. A comprehensive metatranscriptome search for Oblin-1 and HsOblin homologs expands Obelisk diversity about two-fold, revealing multiple subfamilies sharing the same core fold,. some of which are also predicted to encode additional small proteins with simple alpha-helical folds. These findings highlight Obelisks as widespread and overlooked components of microbial ecosystems, expanding understanding of viroid-like RNA replicon diversity and evolution.
The International Committee on Taxonomy of Viruses (ICTV) holds a ratification vote annually following the review of newly proposed taxa by ICTV Study Groups and members of the virology community. This article reports changes to the taxonomy of viruses infecting archaea that were approved and ratified by the ICTV in March 2025. Six new families of head-tailed viruses expanded the order Caudoviricetes (realm Duplodnaviria); one new family of filamentous viruses was added to the order Ligamenvirales (realm Adnaviria); one new family of viruses with pleomorphic virions was included within a new phylum, new order and new class in the kingdom Trapavirae (realm Monodnaviria); finally, three new families were created for spindle-shaped viruses that remain unassigned to higher level taxa. The 25 new species represent viruses infecting a broad range of archaea, including members of the classes Archaeoglobi, Bathyarchaeia, Methanobacteria, Methanomicrobia, Nitrososphaeria and Poseidoniia. Most of these viruses have been discovered by metagenomics in samples derived from diverse environments, including ambient and extreme marine ecosystems, the gastrointestinal tract of humans and animals, anaerobic digesters and terrestrial hot springs. Following this taxonomic update, archaeal viruses are officially classified into a total of 163 virus species in 94 genera within 62 families.
Extensive metatranscriptome mining has recently vastly expanded the range of covalently closed circular (ccc) RNA replicons. A notable group of such replicons are Obelisks, cccRNAs of about 1 kilobase (kb) encoding a protein with a unique fold, Oblin-1, and detected in a broad variety of metatranscriptomes, in particular, those from the human gastrointestinal tract. We adopted Fragmented and primer-Ligated DsRNA Sequencing (FLDS) method to selectively sequence double-stranded (ds) RNAs, the replicative intermediates of RNA replicons, and to identify cccRNAs among the resulting sequences. From these data, we selected cccRNAs with predicted extensive intramolecular base-pairing, a hallmark of viroid-like elements. ch We employed FLDS to explore metatranscriptomes from acidic hot springs in Japan and discovered a distinct family of Obelisks probably associated with thermoacidophilic bacteria (Hot spring Obelisks, HsObs). The proteins encoded by HsObs, HsOblins, show no significant sequence similarity to previously identified Oblin-1 proteins, but are predicted to adopt a closely similar structure. A comprehensive search of metagenomes for Oblin-1 and HsOblin homologs substantially expanded this family of Obelisk-encoded proteins revealing several distinct subfamilies that share the same core fold. A cccRNA encoding an HsOblin homolog was also detected in a Yellowstone hot spring metatranscriptome. Apart from Oblin-1, some subfamilies of Obelisks were predicted to encode additional small proteins with simple alpha-helical folds.
Metatranscriptome sequencing dramatically expanded the known diversity of the global RNA virome and, in particular, suggested several new candidate phyla in riboviruses. Using a double-stranded RNA (dsRNA) sequencing, here, we report five complete, bisegmented RNA genomes of a putative phylum group, paraxenoviruses, identified from marine environments. Phylogenetic analysis of the RNA-directed RNA polymerases of paraxenoviruses demonstrated their affinity with the ribovirus order Durnavirales within the class Duplopiviricetes of the phylum Pisuviricota. The order Durnavirales includes families Cystoviridae that consists of well-characterized dsRNA bacteriophages and less thoroughly studied Picobirnaviridae that are also suspected to infect bacteria. Consistently, modeling and analysis of the structure of the predicted capsid protein (CP) of several paraxenoviruses revealed similarity to picobirnavirus CP although the paraxenovirus CP is much larger and contains unique structural elaborations. Taken together, these affinities suggest that paraxenoviruses represent a distinct family within Durnavirales, which we provisionally name "Paraxenoviridae". Both genomic segments in Picobirnaviridae and "Paraxenoviridae" encompass multiple open reading frames, each preceded by a typical bacterial ribosome-binding site, strongly suggesting that these families consist of bacterial viruses. Search for homologs of paraxenovirus genes shows widespread distribution of this virus group in the global ocean, suggesting an important contribution to marine microbial ecosystems. Our findings further expand the diversity and ecological role of the bacterial RNA virome, reveal extensive structural variability of RNA viral CPs, and demonstrate the common ancestry of several distinct families of bacterial viruses with dsRNA genomes.
Mirusviruses infect unicellular eukaryotes and are related to tailed bacteriophages and herpesviruses. Here we expand the known diversity of mirusviruses by screening diverse metagenomic assemblies and characterizing 1,202 non-redundant environmental genomes. Mirusviricota comprises a highly diversified phylum of large and giant eukaryotic viruses that rivals the evolutionary scope and functional complexity of nucleocytoviruses. Critically, major Mirusviricota lineages lack essential genes encoding components of the replication and transcription machineries and, concomitantly, encompass numerous spliceosomal introns that are enriched in virion morphogenesis genes. These features point to multiple transitions from cytoplasmic to nuclear reproduction during mirusvirus evolution. Many mirusvirus introns encode diverse homing endonucleases, suggestive of a previously undescribed mechanism promoting the horizontal mobility of spliceosomal introns. Available metatranscriptomes reveal long-range trans-splicing in a virion morphogenesis gene. Collectively, our data strongly suggest that nuclei of unicellular eukaryotes across marine and freshwater ecosystems worldwide are a major niche for replication of intron-rich mirusviruses.
Similar to many eukaryotes, the thermoacidophilic archaeon Saccharolobus islandicus follows a defined cell cycle program, with two growth phases, G1 and G2, interspersed by a chromosome replication phase (S), and followed by genome segregation and cytokinesis (M-D) phases. To study whether and which other processes are cell cycle-coordinated, we synchronized cultures of S. islandicus and performed an in-depth transcriptomic analysis of samples enriched in cells undergoing the M-G1, S, and G2 phases, providing a holistic view of the S. islandicus cell cycle. We show that diverse metabolic pathways, protein synthesis, cell motility and even antiviral defense systems, are expressed in a cell cycle-dependent fashion. Moreover, application of a transcriptome deconvolution method defined sets of phase-specific signature genes, whose peaks of expression roughly matched those of yeast homologs. Collectively, our data elucidates the complexity of the S. islandicus cell cycle, suggesting that it more closely resembles the cell cycle of certain eukaryotes than previously appreciated.
Fragmented and primer Ligated DsRNA Sequencing (FLDS) was used to reconstruct five complete, bisegmented RNA genomes of paraxenoviruses, a group of viruses that was previously identified in the ocean and that based on the analysis of partial genomes was proposed to represent a putative new phylum within the kingdom Orthornavirae of the realm Riboviria. Phylogenetic analysis of the RNA-directed RNA polymerases of paraxenoviruses demonstrated their affinity with the ribovirus order Durnavirales within the class Duplopiviricetes of the phylum Pisuviricota. The order Durnavirales includes families Cystoviridae that consists of well-characterized dsRNA bacteriophages and less thoroughly studied Picobirnaviridae that are also suspected to infect bacteria. Consistently, modeling and analysis of the structure of the predicted capsid protein (CP) of several paraxenoviruses revealed similarity to picobirnavirus CP although the paraxenovirus CP is much larger and contains unique structural elaborations. Taken together, these affinities suggest that paraxenoviruses represent a distinct family within Durnavirales, which we provisionally name "Paraxenoviridae". Both genomic segments in Picobirnaviridae and "Paraxenoviridae" encompass multiple open reading frames, each preceded by a typical bacterial ribosomebinding site, strongly suggesting that these families consist of bacterial viruses. Search for homologs of paraxenovirus genes shows widespread distribution of this virus group in the global ocean, suggesting a potential important contribution to marine microbial ecosystems. Our findings further expand the diversity and ecological role of the bacterial RNA virome, reveal extensive structural variability of RNA viral capsid proteins, and demonstrate the common ancestry of several distinct families of bacterial viruses with dsRNA genomes.
Prokaryotic cells employ multiple protective layers crucial for defense, structural integrity, and cellular interactions in the environment. Archaea often feature an S-layer, with some species possessing additional and remarkably resistant sheaths. The archaeal sheath has been studied in Methanothrix and Methanospirillum, revealing a complex structure consisting of amyloid proteins organized into rings. Here, we conducted a comprehensive survey of sheath-forming proteins (SH proteins) across archaeal genomes. Structural modeling reveals a rich diversity of SH proteins, indicating the presence of a sheath in members of the TACK superphylum (Thermoprotei), as well as in the methanotrophic ANME-1. SH proteins are present in up to 40 copies per genome and display diverse domain arrangements suggesting multifunctional roles within the sheath, and potential involvement in cell-cell interaction with syntrophic partners. We uncover a complex evolutionary dynamic, indicating active exchange of SH proteins in archaeal communities. We find that viruses infecting sheathed archaea encode a diversity of SH-like proteins and we use them as markers to identify 580 vOTUs potentially associated with sheathed archaea. Structural modeling suggests that viral SH proteins can form complexes with the host SH proteins. We propose a previously unreported egress strategy where the expression of viral SH-like proteins may disrupt the integrity of the host sheath and facilitate viral exit during lysis. Together, our results significantly expand knowledge of the diversity and evolution of the archaeal sheath, which has been largely understudied but might have an important role in shaping microbial communities.
The cell cycle is a series of events that occur from the moment of cell birth to cell division. In eukaryotes, cell growth, genome replication, genome segregation, and cytokinesis are strictly coordinated, defining discrete cell cycle phases. In contrast, these key processes may occur concurrently in bacteria. Thermoacidophilic archaea in the genus Saccharolobus follow a defined cell cycle program, with the first pre-replicative growth (G1) phase, followed by the chromosome replication (S) phase, the second growth (G2) phase, and rapid genome segregation (M) and cytokinesis (D) phases. However, whether other processes, such as metabolism, catabolism, protein translation, and antiviral defense also occur at specific cell cycle phases, as in eukaryotes, or are active throughout the cell cycle, as in bacteria, remains unclear. To address this question, we synchronized cultures of S. islandicus and performed an in-depth transcriptomic analysis of samples enriched in cells undergoing the M-G1, S, and G2 phases. Differential gene expression and consensus gene co-expression network analyses provided a holistic view of the S. islandicus cell cycle. In addition to the core transcriptome network, which is expressed throughout the cell cycle, we show that diverse metabolic pathways, protein synthesis, cell motility and even antiviral defense systems, are expressed in a cell cycle dependent fashion. Our data also refines understanding of the processes previously known to be linked to the cell cycle, such as DNA replication. We show that most DNA replication genes are expressed prior to the S phase, during the M-G1, whereas expression of the major chromatin genes, and accordingly, chromatinization are concomitant with replication. A statistical model was used to define sets of signature genes characteristic of each of the analyzed cell cycle phases, emphasizing transcriptional stratification of the phases. Signature genes are more conserved across Thermoproteota than non-signature genes and their peak expression, especially for the M-G1 and G2 specific genes, matches that of homologs in yeast. Collectively, our data elucidate the complexity of the S. islandicus cell cycle and suggest that it more closely resembles the cell cycle of eukaryotes than previously appreciated. ### Competing Interest Statement The authors have declared no competing interest.
Type VI CRISPR-Cas systems are among the few CRISPR varieties that target exclusively RNA. The CRISPR RNA–guided, sequence-specific binding of target RNAs, such as phage transcripts, activates the type VI effector, Cas13. Once activated, Cas13 causes collateral RNA cleavage, which induces bacterial cell dormancy, thus protecting the host population from the phage spread. We show here that the principal form of collateral RNA degradation elicited by Leptotrichia shahii Cas13a expressed in Escherichia coli cells is the cleavage of anticodons in a subset of transfer RNAs (tRNAs) with uridine-rich anticodons. This tRNA cleavage is accompanied by inhibition of protein synthesis, thus providing defense from the phages. In addition, Cas13a-mediated tRNA cleavage indirectly activates the RNases of bacterial toxin-antitoxin modules cleaving messenger RNA, which could provide a backup defense. The mechanism of Cas13a-induced antiphage defense resembles that of bacterial anticodon nucleases, which is compatible with the hypothesis that type VI effectors evolved from an abortive infection module encompassing an anticodon nuclease.
The human gut virome, which is mainly composed of bacteriophages, also includes viruses infecting archaea, yet their role remains poorly understood due to lack of isolates. Here, we characterize a temperate archaeal virus (MSTV1) infecting Methanobrevibacter smithii, the dominant methanogenic archaeon of the human gut. The MSTV1 genome is integrated in the host chromosome as a provirus which is sporadically induced, resulting in virion release. Using cryo-electron tomography, we capture several intracellular virion assembly intermediates and confirm that only a small fraction of the host population actively produces virions in vitro. Similar low frequency of induction is observed in a mouse colonization model, using mice harboring a stable consortium of 12 bacterial species (OMM12). Transcriptomic analysis suggests a regulatory lysogeny-lysis switch involving an interplay between viral proteins to maintain virus-host equilibrium, ensuring host survival and viral persistence. Thus, our study sheds light on archaeal virus-host interactions and highlights similarities with bacteriophages in establishing stable coexistence with their hosts in the gut. The human gut virome includes understudied viruses that infect archaea. Here, Baquero et al. characterize a temperate archaeal virus that infects the dominant methanogenic archaeon of the human gut, shedding light on archaeal virus-host interactions and highlighting similarities with gut bacteriophages in establishing stable coexistence with their hosts.
Abstract Mobile genetic elements (MGEs), especially viruses, have a major impact on microbial communities. Methanogenic archaea play key environmental and economical roles, being the main producers of methane -a potent greenhouse gas and an energy source. They are widespread in diverse anoxic artificial and natural environments, including animal gut microbiomes. However, their viruses remain vastly unknown. Here, we carried out a global investigation of MGEs in 3436 genomes and metagenome-assembled genomes covering all known diversity of methanogens and using a newly assembled CRISPR database consisting of 60,000 spacers of methanogens, the most extensive collection to date. We obtained 248 high-quality (pro)viral and 63 plasmid sequences assigned to hosts belonging to nine main orders of methanogenic archaea, including the first MGEs of Methanonatronarchaeales, Methanocellales and Methanoliparales archaea. We found novel CRISPR arrays in ‘Ca. Methanomassiliicoccus intestinalis’ and ‘Ca. Methanomethylophilus’ genomes with spacers targeting small ssDNA viruses of the Smacoviridae, supporting and extending the hypothesis of an interaction between smacoviruses and gut associated Methanomassiliicoccales. Gene network analysis shows that methanogens encompass a unique and interconnected MGE repertoire, including novel viral families belonging to head-tailed Caudoviricetes, but also icosahedral and archaeal-specific pleomorphic, spherical, and spindle (pro)viruses. We reveal well-delineated modules for virus-host interaction, genome replication and virion assembly, and a rich repertoire of defense and counter-defense systems suggesting a highly dynamic and complex network of interactions between methanogens and their MGEs. We also identify potential conjugation systems composed of VirB4, VirB5 and VirB6 proteins encoded on plasmids and (pro)viruses of Methanosarcinales, the first report in Euryarchaeota. We identified 15 new families of viruses infecting Methanobacteriales, the most prominent archaea in the gut microbiome. These encode a large repertoire of protein domains for recognizing and cleaving pseudomurein for viral entry and egress, suggesting convergent adaptation of bacterial and archaeal viruses to the presence of a cell wall. Finally, we highlight an enrichment of glycan-binding domains (immunoglobulin-like (Ig-like)/Flg_new) and diversity-generating retroelements (DGRs) in viruses from gut-associated methanogens, suggesting a role in adaptation to host environments and remarkable convergence with phages infecting gut-associated bacteria. Our work represents an important step toward the characterization of the vast repertoire of MGEs associated with methanogens, including a better understanding of their role in regulating their communities globally and the development of much-needed genetic tools.
Methanogenic archaea are major producers of methane, a potent greenhouse gas and biofuel, and are widespread in diverse environments, including the animal gut. The ecophysiology of methanogens is likely impacted by viruses, which remain, however, largely uncharacterized. Here we carried out a global investigation of viruses associated with all current diversity of methanogens by assembling an extensive CRISPR database consisting of 156,000 spacers. We report 282 high-quality (pro)viral and 205 virus-like/plasmid sequences assigned to hosts belonging to ten main orders of methanogenic archaea. Viruses of methanogens can be classified into 87 families, underscoring a still largely undiscovered genetic diversity. Viruses infecting gut-associated archaea provide evidence of convergence in adaptation with viruses infecting gut-associated bacteria. These viruses contain a large repertoire of lysin proteins that cleave archaeal pseudomurein and are enriched in glycan-binding domains (Ig-like/Flg_new) and diversity-generating retroelements. The characterization of this vast repertoire of viruses paves the way towards a better understanding of their role in regulating methanogen communities globally, as well as the development of much-needed genetic tools.
Asgardarchaeota harbour many eukaryotic signature proteins and are widely considered to represent the closest archaeal relatives of eukaryotes. Whether similarities between Asgard archaea and eukaryotes extend to their viromes remains unknown. Here we present 20 metagenome-assembled genomes of Asgardarchaeota from deep-sea sediments of the basin off the Shimokita Peninsula, Japan. By combining a CRISPR spacer search of metagenomic sequences with phylogenomic analysis, we identify three family-level groups of viruses associated with Asgard archaea. The first group, verdandiviruses, includes tailed viruses of the class Caudoviricetes (realm Duplodnaviria); the second, skuldviruses, consists of viruses with predicted icosahedral capsids of the realm Varidnaviria; and the third group, wyrdviruses, is related to spindle-shaped viruses previously identified in other archaea. More than 90% of the proteins encoded by these viruses of Asgard archaea show no sequence similarity to proteins encoded by other known viruses. Nevertheless, all three proposed families consist of viruses typical of prokaryotes, providing no indication of specific evolutionary relationships between viruses infecting Asgard archaea and eukaryotes. Verdandiviruses and skuldviruses are likely to be lytic, whereas wyrdviruses potentially establish chronic infection and are released without host cell lysis. All three groups of viruses are predicted to play important roles in controlling Asgard archaea populations in deep-sea ecosystems. Analysis of CRISPR spacers in Asgardarchaeota metagenomes reveals three family-level groups of viruses associated with these microbial eukaryotes.
Type VI CRISPR-Cas systems are the only CRISPR variety that cleaves exclusively RNA 1,2 . In addition to the CRISPR RNA (crRNA)-guided, sequence-specific binding and cleavage of target RNAs, such as phage transcripts, the type VI effector, Cas13, causes collateral RNA cleavage, which induces bacterial cell dormancy, thus protecting the host population from phage spread 3,4 . We show here that the principal form of collateral RNA degradation elicited by Cas13a protein from Leptotrichia shahii upon target RNA recognition is the cleavage of anticodons of multiple tRNA species, primarily those with anticodons containing uridines. This tRNA cleavage is necessary and sufficient for bacterial dormancy induction by Cas13a. In addition, Cas13a activates the RNases of bacterial toxin-antitoxin modules, thus indirectly causing mRNA and rRNA cleavage, which could provide a back-up defense mechanism. The identified mode of action of Cas13a resembles that of bacterial anticodon nucleases involved in antiphage defense 5 , which is compatible with the hypothesis that type VI effectors evolved from an abortive infection module 6,7 encompassing an anticodon nuclease.
For Type I CRISPR-Cas systems, a mode of CRISPR adaptation named priming has been described. Priming allows specific and highly efficient acquisition of new spacers from DNA recognized (primed) by the Cascade-crRNA (CRISPR RNA) effector complex. Recognition of the priming protospacer by Cascade-crRNA serves as a signal for engaging the Cas3 nuclease-helicase required for both interference and primed adaptation, suggesting the existence of a primed adaptation complex (PAC) containing the Cas1-Cas2 adaptation integrase and Cas3. To detect this complex in vivo, we here performed chromatin immunoprecipitation with Cas3-specific and Cas1-specific antibodies using cells undergoing primed adaptation. We found that prespacers are bound by both Cas1 (presumably, as part of the Cas1-Cas2 integrase) and Cas3, implying direct physical association of the interference and adaptation machineries as part of PAC.
CRISPR (clustered regularly interspaced short palindromic repeats) Cas (CRISPR-associated) systems provide prokaryotes with efficient protection against foreign nucleic acid invaders. We have recently demonstrated the defensive interference function of a CRISPR-Cas system from Clostridioides (Clostridium) difficile, a major human enteropathogen, and showed that it could be harnessed for efficient genome editing in this bacterium. However, molecular details are still missing on CRISPR-Cas function for adaptation and sequence requirements for both interference and new spacer acquisition in this pathogen. Despite accumulating knowledge on the individual CRISPR-Cas systems in various prokaryotes, no data are available on the adaptation process in bacterial type I-B CRISPR-Cas systems. Here, we report the first experimental evidence that the C. difficile type I-B CRISPR-Cas system acquires new spacers upon overexpression of its adaptation module. The majority of new spacers are derived from a plasmid expressing Cas proteins required for adaptation or from regions of the C difficile genome where generation of free DNA termini is expected. Results from protospacer-adjacent motif (PAM) library experiments and plasmid conjugation efficiency assays indicate that C. difficile CRISPR-Cas requires the YCN consensus PAM for efficient interference. We revealed a functional link between the adaptation and interference machineries, since newly adapted spacers are derived from sequences associated with a CCN PAM, which fits the interference consensus. The definition of functional PAMs and establishment of relative activity levels of each of the multiple C. difficile CRISPR arrays in present study are necessary for further CRISPR-based biotechnological and medical applications involving this organism. IMPORTANCE CRISPR Cas systems provide prokaryotes with adaptive immunity for defense against foreign nucleic acid invaders, such as viruses or phages and plasmids. The CRISPR-Cas systems are highly diverse, and detailed studies of individual CRISPR-Cas subtypes are important for our understanding of various aspects of microbial adaptation strategies and for the potential applications. The significance of our work is in providing the first experimental evidence for type I-B CRISPR-Cas system adaptation in the emerging human enteropathogen Clostridioides difficile. This bacterium needs to survive in phage-rich gut communities, and its active CRISPR-Cas system might provide efficient antiphage defense by acquiring new spacers that constitute memory for further invader elimination. Our study also reveals a functional link between the adaptation and interference CRISPR machineries. The definition of all possible functional trinucleotide motifs upstream protospacers within foreign nucleic acid sequences is important for CRISPR-based genome editing in this pathogen and for developing new drugs against C. difficile infections.
Saccharolobus (formerly Sulfolobus) shibatae B12, isolated from a hot spring in Beppu, Japan in 1982, is one of the first hyperthermophilic and acidophilic archaeal species to be discovered. It serves as a natural host to the extensively studied spindle-shaped virus SSV1, a prototype of the Fuselloviridae family. Two additional Sa. shibatae strains, BEU9 and S38A, sensitive to viruses of the families Lipothrixviridae and Portogloboviridae, respectively, have been isolated more recently. However, none of the strains has been fully sequenced, limiting their utility for studies on archaeal biology and virus-host interactions. Here, we present the complete genome sequences of all three Sa. shibatae strains and explore the rich diversity of their integrated mobile genetic elements (MGE), including transposable insertion sequences, integrative and conjugative elements, plasmids, and viruses, some of which were also detected in the extrachromosomal form. Analysis of related MGEs in other Sulfolobales species and patterns of CRISPR spacer targeting revealed a complex network of MGE distributions, involving horizontal spread and relatively frequent host switching by MGEs over large phylogenetic distances, involving species of the genera Saccharolobus, Sulfurisphaera and Acidianus. Furthermore, we characterize a remarkable case of a virus-to-plasmid transition, whereby a fusellovirus has lost the genes encoding for the capsid proteins, while retaining the replication module, effectively becoming a plasmid.
CRISPR arrays are prokaryotic genomic loci consisting of repeat sequences alternating with unique spacers acquired from foreign nucleic acids. As one of the fastest-evolving parts of the genome, CRISPR arrays can be used to differentiate closely related prokaryotic lineages and track individual strains in prokaryotic communities. However, the assembly of full-length CRISPR arrays sequences remains a problem. Here, we developed SCRAMBLER, a tool that includes several pipelines for assembling CRISPR arrays from high-throughput short-read sequencing data. We assessed its performance with model data sets (Escherichia coli strains containing different CRISPR arrays and imitating prokaryotic communities of different complexities) and intestinal microbiomes of extant and extinct pachyderms. Evaluation of SCRAMBLER's performance using model data sets demonstrated its ability to assemble CRISPR arrays correctly from reads containing pairs of spacers, yielding a precision rate of >80% and a recall rate of 60-85% when checked against ground-truth data. Likewise, SCRAMBLER successfully assembled CRISPR arrays from the environmental samples, as attested by their matching with database entries. SCRAMBLER, an open-source software (github.com/biolab-tools/SCRAMBLER), can facilitate analysis of the composition and dynamics of CRISPR arrays in complex communities.