The breadth of life's diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 × 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the world's public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.
Machine learning has generated millions of high-quality predicted protein structures, creating a need for computationally efficient structure search algorithms and robust estimates of statistical significance at this scale. We show that unrelated proteins have a universal tendency towards convergent evolution of secondary and tertiary motifs, causing an excess of high-scoring false positive alignments. We investigate popular structure search and alignment algorithms, finding that previous methods routinely overestimate significance by up to six orders of magnitude. To address these issues, and to accommodate recent innovations in search algorithm design, we describe a novel method for estimating statistical significance. We show that its E -values are accurate, scale successfully with database size, and are robust against the (generally unknown) diversity of folds in the database. We implement our approach in an online structure search service based on Reseek at https://reseek.online .
Protein multiple alignment is an essential step in many bioinformatics analysis such as phylogenetic tree estimation, HMM construction and critical residue identification. Structure is conserved between distantly-related proteins where amino acid similarity is weak or undetectable, suggesting that structure-informed sequence alignments might offer advantages over alignments constructed from amino acid sequences alone. The advent of the AI folding era has unleashed millions of high-quality predicted structures, motivating the development and assessment of scalable multiple structure alignment (MStA) methods. Here, we describe Muscle-3D, a new MStA algorithm combining a rich sequence representation of structure context, the Reseek "mega-alphabet", with state-of-the art alignment techniques from Muscle5 including a posterior decoding pair-HMM, consistency transformation, iterative refinement and ensemble construction. We show that Muscle-3D readily scales to thousands of structures. Comparative validation on several benchmark datasets using different quality metrics shows Muscle-3D to be among the higher-scoring methods, but we find that algorithm rankings from different metrics disagree despite low P-values according to the Wilcoxon rank-sum test. We suggest that these conflicts arise from the inherently fuzzy nature of structural alignment, and argue that a universal standard of MStA accuracy is not possible in principle. We describe contact map profiles for visualizing variation in inter-residue distances, and introduce a novel measure of local conformation similarity, LDDT-muw. Muscle-3D software is available at https://github.com/rcedgar/muscle. ### Competing Interest Statement The authors have declared no competing interest.
Infection with resistant bacteria has become an ever-increasing problem in modern medical practice. Bacteremia is a serious and potentially lethal condition that can lead to sepsis without early intervention. Currently, broad-spectrum antibiotics are prescribed until bacteria can be identified through blood cultures, a process that can take 2-3 days and is unable to provide quantitative information. Staphylococcus aureus (S. aureus) is a leading cause of bacteremia, and methicillin-resistant S. aureus (MRSA) accounts for more than a third of the cases. Other bacteria such as Clostridium difficile, Acinetobacter baumannii, and Carbapenem-resistant Enterobacteriaceae are becoming more prevalent and antibiotic-resistant. Rapid diagnostics for each of these superbugs has been a priority for health organizations around the world. Bacteriophages have evolved for millions of years to develop exquisite specificity in target binding using their host attachment proteins. Bacteriophages are viruses that infect bacteria. Bacteriophages use tail spikes, specialized attachment proteins, to bind specifically to their target bacterial cell surface proteins. We use bacteriophages and parts of bacteriophages as specific tags coupled with photoacoustic flow cytometry for the detection and quantification of bacteria. In photoacoustic flow cytometry, laser light is absorbed by particles under flow, and the ultrasound waves generated on the release of the energy are detected. Photoacoustics involves the detection of ultrasound waves resulting from laser irradiation. In photoacoustic flow cytometry, pulsed laser light is delivered to a sample flowing past a focused transducer, and particles that absorb laser light create a photoacoustic response. Bacteria can be tagged with dyed bacteriophage and processed through a photoacoustic flow cytometer where they are detected by the acoustic response. In this chapter, we describe the procedure and methods used to accomplish this. Often the limiting factor for the treatment of patients is the time spent waiting for results. It is our hope that the work presented in this chapter can be a foundation for future work and provide an ability to detect bacterial pathogens in blood cultures. Bacterial plate cultures and Gram staining are nineteenth-century technologies that have been the gold standards for decades, but current trends in resistant bacteria have necessitated a move toward more rapid and quantifiable diagnostic tools.
Here, we describe “obelisks,” a class of heritable RNA elements sharing several properties: (1) apparently circular RNA ∼1 kb genome assemblies, (2) predicted rod-like genome-wide secondary structures, and (3) open reading frames encoding a novel “Oblin” protein superfamily. A subset of obelisks includes a variant hammerhead self-cleaving ribozyme. Obelisks form their own phylogenetic group without detectable similarity to known biological agents. Surveying globally, we identified 29,959 distinct obelisks (clustered at 90% sequence identity) from diverse ecological niches. Obelisks are prevalent in human microbiomes, with detection in ∼7% (29/440) and ∼50% (17/32) of queried stool and oral metatranscriptomes, respectively. We establish Streptococcus sanguinis as a cellular host of a specific obelisk and find that this obelisk’s maintenance is not essential for bacterial growth. Our observations identify obelisks as a class of diverse RNAs of yet-to-be-determined impact that have colonized and gone unnoticed in human and global microbiomes.
Recent breakthroughs in protein fold prediction from amino acid sequences have unleashed a deluge of new structures, raising new opportunities for expanding insights into the universe of proteins and pursuing practical applications in bio-engineering and therapeutics while also presenting new challenges to protein search and analysis algorithms. Here, I describe Reseek, a protein alignment algorithm which doubles sensitivity in protein homolog detection compared to state-of-the-art methods including DALI, TM-align and Foldseek, with improved speed over Foldseek, the fastest previous method. Reseek is based on alignment of sequences where each residue in the protein backbone is represented by a letter in a novel “mega-alphabet” of 85,899,345,920 (∼ 1011) distinct states. Code is available at . ### Competing Interest Statement The author is a co-founder of GigaMune, Inc, a biotechnology company pursuing research related to this article.
[This corrects the article DOI: 10.7717/peerj.14055.].
MOTIVATION:Recent breakthroughs in protein fold prediction from amino acid sequences have unleashed a deluge of new structures, presenting new opportunities and challenges to bioinformatics. RESULTS:Reseek is a novel protein structure alignment algorithm based on sequence alignment where each residue in the protein backbone is represented by a letter in a "mega-alphabet" of 85 899 345 920 (∼1011) distinct states. Reseek achieves substantially improved sensitivity to remote homologs compared to state-of-the-art methods including DALI, TMalign, and Foldseek, with comparable speed to Foldseek, the fastest previous method. Scaling to large databases of AI-predicted folds is analyzed. Foldseek E-values are shown to be under-estimated by several orders of magnitude, while Reseek E-values are in good agreement with measured error rates. AVAILABILITY AND IMPLEMENTATION:https://github.com/rcedgar/reseek.
IntroductionAntibiotic resistance in bacterial species constitutes a growing problem in the clinical management of infections. Not only does it limit therapeutic options, but application of ineffective antibiotics allows resistant species to progress prior to prescribing more effective treatment to patients. Methicillin resistance in Staphylococcus aureus is a major problem in clinical infections as it is the most common hospital acquired infection.MethodsWe developed a photoacoustic flow cytometer using engineered bacteriophage as probes for rapid determination of methicillin resistance in Staphylococcus aureus with thirteen clinical samples obtained from keratitis patients. This method irradiates cells under flow with 532 nm laser light and selectively generates acoustic waves in labeled bacterial cells, thus enabling detection and enumeration of them. Staphylococcus aureus isolates were classified from culture isolation as either methicillin resistant or susceptible using cefoxitin disk diffusion testing. The photoacoustic method enumerates bacterial cells before and after treatment with antibiotics. Decreasing counts of bacteria after treatment indicate susceptible strains. We quantified the bacterial cells in the treated and untreated samples.ResultsUsing k-means clustering on the data, we achieved 100% concordance with the classification of Staphylococcus aureus resistance using culture.DiscussionPhotoacoustics can be used to differentiate methicillin resistant and susceptible strains of bacteria from ocular infections. This method may be generalized to other bacterial species using appropriate bacteriophages and testing for resistance using other antibiotics.
Photoacoustic flow cytometry is a method to detect rare analytes in fluids. We developed photoacoustic flow cytometry to detect pathological cells in body fluids, such as circulating tumor cells or bacteria in blood. In order to induce specific optical absorption in bacteria, we use modified bacteriophage that precisely target bacterial species or subspecies for rapid identification. In order to reduce detection variability and to halt the lytic lifescycle that results in lysis of the bacteria, we attached dyed latex microspheres to the tail fibers of bacteriophage that retained the bacterial recognition binding sites. We tested these microsphere complexes using Salmonella enterica (Salmonella) and Escherichia coli (E. coli) bacteria and found robust and specific detection of targeted bacteria. In our work we used LT2, a strain of Salmonella, against K12, a strain of E. coli. Using Det7, a bacteriophage that binds to LT2 and not to K12, we detected an average of 109.3±9.0 of LT2 versus 2.0±1.7 of K12 using red microspheres and 86.7±13.2 of LT2 versus 0.3±0.6 of K12 using blue microspheres. These results confirmed our ability to selectively detect bacterial species using photoacoustic flow cytometry.
Earth's life may have originated as self-replicating RNA, and it has been argued that RNA viruses and viroid-like elements are remnants of such pre-cellular RNA world. RNA viruses are defined by linear RNA genomes encoding an RNA-dependent RNA polymerase (RdRp), whereas viroid-like elements consist of small, single-stranded, circular RNA genomes that, in some cases, encode paired self-cleaving ribozymes. Here we show that the number of candidate viroid-like elements occurring in geographically and ecologically diverse niches is much higher than previously thought. We report that, amongst these circular genomes, fungal ambiviruses are viroid-like elements that undergo rolling circle replication and encode their own viral RdRp. Thus, ambiviruses are distinct infectious RNAs showing hybrid features of viroid-like RNAs and viruses. We also detected similar circular RNAs, containing active ribozymes and encoding RdRps, related to mitochondrial-like fungal viruses, highlighting fungi as an evolutionary hub for RNA viruses and viroid-like elements. Our findings point to a deep co-evolutionary history between RNA viruses and subviral elements and offer new perspectives in the origin and evolution of primordial infectious agents, and RNA life.
A recent study proposed five new RNA virus phyla, two of which, 'Taraviricota' and 'Arctiviricota', were stated to be 'dominant in the oceans'. However, the study's assignments classify 28,353 putative RdRp-containing contigs to known phyla but only 886 (2.8%) to the five proposed new phyla combined. I re-mapped the reads to the contigs, finding that known phyla also account for a large majority (93.8%) of reads according to the study's classifications, and that contigs originally assigned to 'Arctiviricota' accounted for only a tiny fraction (0.01%) of reads from Arctic Ocean samples. Performing my own virus identification and classifications, I found that 99.95 per cent of reads could be assigned to known phyla. The most abundant species was Beihai picorna-like virus 34 (15% of reads), and the most abundant order-like cluster was classified as Picornavirales (45% of reads). Sequences in the claimed new phylum 'Pomiviricota' were placed inside a phylogenetic tree for established order Durnavirales with 100 per cent confidence. Moreover, two contigs assigned to the proposed phylum 'Taraviricota' were found to have high-identity alignments to dinoflagellate proteins, tentatively identifying this group of RdRp-like sequences as deriving from non-viral transcripts. Together, these results comprehensively contradict the claim that new phyla dominate the data.
Multiple sequence alignments are widely used to infer evolutionary relationships, enabling inferences of structure, function, and phylogeny. Standard practice is to construct one alignment by some preferred method and use it in further analysis; however, undetected alignment bias can be problematic. I describe Muscle5, a novel algorithm which constructs an ensemble of high-accuracy alignment with diverse biases by perturbing a hidden Markov model and permuting its guide tree. Confidence in an inference is assessed as the fraction of the ensemble which supports it. Applied to phylogenetic tree estimation, I show that ensembles can confidently resolve topologies with low bootstrap according to standard methods, and conversely that some topologies with high bootstraps are incorrect. Applied to the phylogeny of RNA viruses, ensemble analysis shows that recently adopted taxonomic phyla are probably polyphyletic. Ensemble analysis can improve confidence assessment in any inference from an alignment.
Conventionally, hyperimmune globulin drugs manufactured from pooled immunoglobulins from vaccinated or convalescent donors have been used in treating infections where no treatment is available. This is especially important where multi-epitope neutralization is required to prevent the development of immune-evading viral mutants that can emerge upon treatment with monoclonal antibodies. Using microfluidics, flow sorting, and a targeted integration cell line, a first-in-class recombinant hyperimmune globulin therapeutic against SARS-CoV-2 (GIGA-2050) was generated. Using processes similar to conventional monoclonal antibody manufacturing, GIGA-2050, comprising 12,500 antibodies, was scaled-up for clinical manufacturing and multiple development/tox lots were assessed for consistency. Antibody sequence diversity, cell growth, productivity, and product quality were assessed across different manufacturing sites and production scales. GIGA-2050 was purified and tested for good laboratory procedures (GLP) toxicology, pharmacokinetics, and in vivo efficacy against natural SARS-CoV-2 infection in mice. The GIGA-2050 master cell bank was highly stable, producing material at consistent yield and product quality up to >70 generations. Good manufacturing practices (GMP) and development batches of GIGA-2050 showed consistent product quality, impurity clearance, potency, and protection in an in vivo efficacy model. Nonhuman primate toxicology and pharmacokinetics studies suggest that GIGA-2050 is safe and has a half-life similar to other recombinant human IgG1 antibodies. These results supported a successful investigational new drug application for GIGA-2050. This study demonstrates that a new class of drugs, recombinant hyperimmune globulins, can be manufactured consistently at the clinical scale and presents a new approach to treating infectious diseases that targets multiple epitopes of a virus.
RNA viruses encoding a polymerase gene (riboviruses) dominate the known eukaryotic virome. High-throughput sequencing is revealing a wealth of new riboviruses known only from sequence, precluding classification by traditional taxonomic methods. Sequence classification is often based on polymerase sequences, but standardised methods to support this approach are currently lacking. To address this need, we describe the polymerase palmprint, a segment of the palm sub-domain robustly delineated by well-conserved catalytic motifs. We present an algorithm, Palmscan, which identifies palmprints in nucleotide and amino acid sequences; PALMdb, a collection of palmprints derived from public sequence databases; and palmID, a public website implementing palmprint identification, search, and annotation. Together, these methods demonstrate a proof-of-concept workflow for high-throughput characterisation of RNA viruses, paving the path for the continued rapid growth in RNA virus discovery anticipated in the coming decade.
Earth’s life may have originated as self-replicating RNA. Some of the simplest current RNA replicators are RNA viruses, defined by linear RNA genomes encoding an RNA-dependent RNA polymerase (RdRP), and subviral agents with single-stranded, circular RNA genomes, such as viroids encoding paired self-cleaving ribozymes. Amongst a massive expansion of candidate viroid and viroid-like elements, we report that fungal pathogens, ambiviruses, are viroid-like elements which undergo rolling circle replication and encode their own viral RdRP, thus they are a distinct hybrid infectious agent. These findings point to a deep evolutionary history between modern RNA viruses and sub-viral elements and offer new perspectives on the evolution of primordial infectious agents, and RNA life. One-Sentence Summary Novel infectious agents resembling self-cleaving viroid-like RNAs whilst encoding a viral RNA-dependent RNA polymerase.
Public databases contain a planetary collection of nucleic acid sequences, but their systematic exploration has been inhibited by a lack of efficient methods for searching this corpus, which (at the time of writing) exceeds 20 petabases and is growing exponentially 1 . Here we developed a cloud computing infrastructure, Serratus, to enable ultra-high-throughput sequence alignment at the petabase scale. We searched 5.7 million biologically diverse samples (10.2 petabases) for the hallmark gene RNA-dependent RNA polymerase and identified well over 10 5 novel RNA viruses, thereby expanding the number of known species by roughly an order of magnitude. We characterized novel viruses related to coronaviruses, hepatitis delta virus and huge phages, respectively, and analysed their environmental reservoirs. To catalyse the ongoing revolution of viral discovery, we established a free and comprehensive database of these data and tools. Expanding the known sequence diversity of viruses can reveal the evolutionary origins of emerging pathogens and improve pathogen surveillance for the anticipation and mitigation of future pandemics.
Plasma-derived polyclonal antibody therapeutics, such as intravenous immunoglobulin, have multiple drawbacks, including low potency, impurities, insufficient supply and batch-to-batch variation. Here we describe a microfluidics and molecular genomics strategy for capturing diverse mammalian antibody repertoires to create recombinant multivalent hyperimmune globulins. Our method generates of diverse mixtures of thousands of recombinant antibodies, enriched for specificity and activity against therapeutic targets. Each hyperimmune globulin product comprised thousands to tens of thousands of antibodies derived from convalescent or vaccinated human donors or from immunized mice. Using this approach, we generated hyperimmune globulins with potent neutralizing activity against severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) in under 3 months, Fc-engineered hyperimmune globulins specific for Zika virus that lacked antibody-dependent enhancement of disease, and hyperimmune globulins specific for lung pathogens present in patients with primary immune deficiency. To address the limitations of rabbit-derived anti-thymocyte globulin, we generated a recombinant human version and demonstrated its efficacy in mice against graft-versus-host disease. Thousands of recombinant antibodies enriched for specificity against defined targets are assembled in multivalent mixtures with enhanced therapeutic activity.
Early detection of cancer has been a goal of cancer research in general and melanoma research in particular (Birnbaum et al., Lancet Glob Health 6:e885-e893, 2018; Alendar et al., Bosnian J Basic Med Sci 9:77-80, 2009). Early detection of metastasis has been targeted as pivotal to increasing survival rates (Menezes et al., Adv Cancer Res 132:1-44, 2016). Melanoma, though curable in its early stages, has a dramatic decrease in survival rates once metastasis has occurred (Sharma et al., Biotechnol Adv 36:1063-1078, 2018). The transition to metastasis is not well understood and is an area of increasing interest. Metastasis is always premeditated by the shedding of circulating tumor cells (CTCs) from the primary tumor. The ability to isolate rare CTCs from the bloodstream has led to a host of new targets and therapies for cancer (Micalizzi et al., Genes Dev 31:1827-1840, 2017). Detection of CTCs also allows for disease progression to be tracked in real time while eliminating the need to wait for additional tumors to grow. Using a photoacoustic flowmeter, in which we induce ultrasonic responses from circulating melanoma cells (CMCs), we identify and quantify these cells in order to track disease progression. Additionally, these CMCs are captured and isolated allowing for future analysis such as RNA-Seq or microarray analysis.