DNA polymerases are key players in DNA replication, repair, and maintenance. However, the overall abundance, diversity, and distribution of bacterial DNA polymerases have not been systematically explored. To close this knowledge gap, we computationally identified and characterized DNA polymerases and their homologs from A, B, C, X, and Y families in over 3000 representative bacterial species with complete genomes. We found that Y-family is the most abundant, followed by C and A families, whereas B and X families are rare. All species have replicative C-family polymerases, 96% have A-family polymerases, and 88% have Y-family members. In each family, we identified and annotated distinct groups, proofreading nucleases, and interaction motifs. Based on conserved associations for DnaE2 and Y-family groups, we identified 11 types of putative multimeric error-prone DNA polymerases supported by AlphaFold modeling. Approximately 90% of the complexes belong to four major types, exemplified by Meiothermus silvanus PolY-RecA complex, Mycobacterium tuberculosis ImuA-ImuB-DnaE2, Escherichia coli Pol V (UmuC-UmuD ' 2-RecA), and Bacillus subtilis YqjW-YqjX-RecA. We found that distribution patterns of distinct polymerase groups and multimeric complexes are shaped by bacterial lineages, replication-system types, and environmental factors. Our results thus provide a comprehensive picture of DNA polymerase diversity and distribution across the bacterial domain.
Translesion DNA synthesis (TLS) is typically performed by inherently error-prone Y-family DNA polymerases. Extensively studied Escherichia coli Pol V mutasome, composed of UmuC, an UmuD' dimer and RecA is an example of a multimeric Y-family TLS polymerase. Less commonly TLS is performed by DNA polymerases of other families. One of the most intriguing such cases in B-family is represented by archaeal PolB2 and its bacterial homologs. Previously thought to be catalytically inactive, PolB2 was recently shown to be absolutely required for targeted mutagenesis in Sulfolobus islandicus. However, the composition and structure of the PolB2 holoenzyme remain unknown. We used highly accurate AlphaFold structural models, coupled with protein sequence and genome context analysis to comprehensively characterize PolB2 and its associated proteins, PPB2, a small helical protein, and iRadA, a catalytically inactive Rad51 homolog. We showed that these three proteins can form a heteropentameric PolB2 complex featuring high confidence modeling scores. Unexpectedly, we found that PolB2 binds iRadA through a structural motif reminiscent of RadA/Rad51 oligomerization motif. In some mutasomes we identified clamp binding motifs, present in either iRadA or PolB2, but rarely in both. We also used AlphaFold to derive a three-dimensional structure of Pol V, for which the experimental structure remains unsolved thus precluding comprehensive understanding of its molecular mechanism. Our analysis showed that the structural features of Pol V explain many of the puzzling previous experimental results. Even though models of PolB2 and Pol V mutasomes are structurally different, we found striking similarities in their architectural organization and interactions.
Structural data on protein-DNA and protein-RNA interactions are indispensable in molecular biology research. In this article, we review available databases and other web-based resources devoted to 3D structures of protein-nucleic acid complexes. First, we describe the core databases that collect and disseminate experimental data. We then review derivative databases focused specifically on structural data on protein-nucleic acid interactions. Finally, we provide an overview of several useful web servers for structure prediction, analysis and comparison. Tools for investigating protein-nucleic acid complexes are relatively scarce. This is primarily because the methods that integrate structural information from both proteins and nucleic acids are in short supply. However, the emerging AI-driven techniques for structure prediction are expected to boost the development of such methods.
ABSTRACTFTDMP is a software framework for biomolecular docking and scoring. It can perform docking of subunits containing one or more protein, DNA, or RNA chains, followed by subsequent scoring of the resulting models. FTDMP can also be used for the ranking of user‐provided models of biomolecular complexes, generated by any structure prediction method. FTDMP evaluates models according to the consensus‐based method VoroIF‐jury, which combines individual scores derived from the Voronoi tessellation of biomolecular structures. In addition to the default scoring mode, FTDMP can easily adopt additional scores; thus, it may be used as a tool to assess newly developed scoring functions. FTDMP was evaluated during blind testing in recent CAPRI experiments and using protein–protein, protein–DNA, and protein–RNA docking benchmarks. It proved to be a useful tool for different research tasks, related to modeling biomolecular interactions. The software, cleaned docking benchmarks, and benchmarking results are available at https://bioinformatics.lt/software/ftdmp/.
The known diversity of CRISPR-Cas systems continues to expand. To encompass new discoveries, here we present an updated evolutionary classification of CRISPR-Cas systems. The updated CRISPR-Cas classification includes 2 classes, 7 types and 46 subtypes, compared with the 6 types and 33 subtypes in our previous survey 5 years ago. In addition, a classification of the cyclic oligoadenylate-dependent signalling pathway in type III systems is presented. We also discuss recently characterized alternative CRISPR-Cas functionalities, notably, type IV variants that cleave the target DNA and type V variants that inhibit the target replication without cleavage. Analysis of the abundance of CRISPR-Cas variants in genomes and metagenomes shows that the previously defined systems are relatively common, whereas the more recently characterized variants are comparatively rare. These low abundance variants comprise the long tail of the CRISPR-Cas distribution in prokaryotes and their viruses, and remain to be characterized experimentally.
Many recent lines of evidence from the human microbiome and other fields indicate bacterial involvement in various types of cancer. Helicobacter pylori has been recognized as the major cause of stomach cancer (gastric cancer), but the mechanism by which it destabilizes the human genome to cause cancer remains unclear. Our recent studies have identified a unique family of toxic restriction enzymes that excise a base (A: adenine) from their recognition sequence (5'-GTAC). At the resulting abasic sites (5'-GT_C), its inherent endonuclease activity or that of a separate endonuclease may yield atypical strand breaks that resist repair by ligation. Here, we present evidence demonstrating involvement of its H. pylori member, HpPabI, in stomach carcinogenesis: (i) Association of intact HpPabI gene with gastric cancer in the global H. pylori Genome Project and the open genomes; (ii) Frequent mutations at A in 5'-GTAC in the gastric cancer genomes as well as in H. pylori genomes; (iii) Its induction of chromosomal double-strand breaks in infected human cells and of mutagenesis in bacterial test systems. In addition, its unique regions that interact with DNA exhibit signs of diversifying selection. Our further analysis revealed similar oncogenic bacterium-restriction-enzyme pairs for other types of cancer. These results set another stage for cancer research and medicine around oncogenic restriction enzymes.
Present in all three domains of life, Argonaute proteins use short oligonucleotides as guides to recognize complementary nucleic acid targets. In eukaryotes, Argonautes are involved in RNA silencing, whereas in prokaryotes, they function in host defense against invading DNA. Here, we show that SPARDA (short prokaryotic Argonaute, DNase associated) systems from Xanthobacter autotrophicus (Xau) and Enhydrobacter aerosaccus (Eae) function in anti-plasmid defense. Upon activation, SPARDA nonspecifically degrades both invader and genomic DNA, causing host death, thereby preventing further spread of the invader in the population. X-ray structures of the apo Xau and EaeSPARDA complexes show that they are dimers, unlike other apo short pAgo systems, which are monomers. We show that dimerization in the apo state is essential for inhibition of XauSPARDA activity. We demonstrate by cryo-EM that activated XauSPARDA forms a filament. Upon activation, the recognition signal of the bound guide/target duplex is relayed to other functional XauSPARDA sites through a structural region that we termed the “beta-relay”. Owing to dramatic conformational changes associated with guide/target binding, XauSPARDA undergoes a “dimer–monomer–filament” transition as the apo dimer dissociates into the guide/target-loaded monomers that subsequently assemble into the filament. Within the activated filament, the DREN nuclease domains form tetramers that are poised to cleave dsDNA. We show that other SPARDAs also form filaments during activation. Furthermore, we identify the presence of the beta-relay in pAgo from all clades, providing new insights into the structural mechanisms of pAgo proteins. Taken together, these findings reveal the detailed structural mechanism of SPARDA and highlight the importance of the beta-relay mechanism in signal transduction in Argonautes.
The mycobacterial mutasome-comprising ImuA', ImuB, and DnaE2-has been implicated in DNA damage-induced mutagenesis in Mycobacterium tuberculosis. ImuB, which is predicted to enable mutasome function via its interaction with the β clamp, is a catalytically inactive Y-family DNA polymerase. Like some other members of the Y-family, ImuB features a recently identified amino acid motif with homology to the RecA N terminus (RecA-NT). Given the role of RecA-NT in RecA oligomerization, we hypothesized that ImuB RecA-NT mediates the interaction with ImuA', an RecA homolog of unknown function. Here, we constructed a panel of imuB alleles in which the RecA-NT was removed or mutated. Our results indicate that RecA-NT is critical for the interaction of ImuB with ImuA'. A region downstream of RecA-NT, ImuB-C, appears to stabilize the ImuB-ImuA' interaction, but its removal does not prevent complex formation. In contrast, replacing two hydrophobic residues of RecA-NT, L378 and V383, disrupts the ImuA'-ImuB interaction. To our knowledge, this is the first experimental evidence suggesting a role for RecA-NT in mediating the interaction between a Y-family member and an RecA homolog.
Structure-resolved protein interactions with other proteins, peptides and nucleic acids are key for understanding molecular mechanisms. The PPI3D web server enables researchers to query preprocessed and clustered structural data, analyze the results and make homology-based inferences for protein interactions. PPI3D offers three interaction exploration modes: (i) all interactions for proteins homologous to the query, (ii) interactions between two proteins or their homologs and (iii) interactions within a specific PDB entry. The server allows interactive analysis of the identified interactions in both summarized and detailed manner. This includes protein annotations, structures, the interface residues and the corresponding contact surface areas. In addition, users can make inferences about residues at the interaction interface for the query protein(s) from the sequence alignments and homology models. The weekly updated PPI3D database includes all the interaction interfaces and binding sites from PDB, clustered based on both protein sequence and structural similarity, yielding non-redundant datasets without loss of alternative interaction modes. Consequently, the PPI3D users avoid being flooded with redundant information, a typical situation for intensely studied proteins. Furthermore, PPI3D provides a possibility to download user-defined sets of interaction interfaces and analyze them locally. The PPI3D web server is available at https://bioinformatics.lt/ppi3d.
Argonaute (Ago) proteins are present in all three domains of life (bacteria, archaea and eukaryotes). They use small (15-30 nucleotides) oligonucleotide guides to bind complementary nucleic acid targets and are responsible for gene expression regulation, mobile genome element silencing, and defence against viruses or plasmids. According to their domain organization, Agos are divided into long and short Agos. Long Agos found in prokaryotes (long-A and long-B pAgos) and eukaryotes (eAgos) comprise four major functional domains (N, PAZ, MID and PIWI) and two structural linker domains L1 and L2. The majority (∼60%) of pAgos are short pAgos, containing only the MID and inactive PIWI domains. Here we focus on the prokaryotic Argonaute AfAgo from Archaeoglobus fulgidus DSM4304. Although phylogenetically classified as a long-B pAgo, AfAgo contains only MID and catalytically inactive PIWI domains, akin to short pAgos. We show that AfAgo forms a heterodimeric complex with a protein encoded upstream in the same operon, which is a structural equivalent of the N-L1-L2 domains of long pAgos. This complex, structurally equivalent to a long PAZ-less pAgo, outperforms standalone AfAgo in guide RNA-mediated target DNA binding. Our findings provide a missing piece to one of the first and the most studied pAgos.
Proteins often function as part of permanent or transient multimeric complexes, and understanding function of these assemblies requires knowledge of their three-dimensional structures. While the ability of AlphaFold to predict structures of individual proteins with unprecedented accuracy has revolutionized structural biology, modeling structures of protein assemblies remains challenging. To address this challenge, we developed a protocol for predicting structures of protein complexes involving model sampling followed by scoring focused on the subunit-subunit interaction interface. In this protocol, we diversified AlphaFold models by varying construction and pairing of multiple sequence alignments as well as increasing the number of recycles. In cases when AlphaFold failed to assemble a full protein complex or produced unreliable results, additional diverse models were constructed by docking of monomers or subcomplexes. All the models were then scored using a newly developed method, VoroIF-jury, which relies only on structural information. Notably, VoroIF-jury is independent of AlphaFold self-assessment scores and therefore can be used to rank models originating from different structure prediction methods. We tested our protocol in CASP15 and obtained top results, significantly outperforming the standard AlphaFold-Multimer pipeline. Analysis of our results showed that the accuracy of our assembly models was capped mainly by structure sampling rather than model scoring. This observation suggests that better sampling, especially for the antibody-antigen complexes, may lead to further improvement. Our protocol is expected to be useful for modeling and/or scoring protein assemblies.
EDITORIAL article Front. Bioinform., 22 May 2023Sec. Protein Bioinformatics Volume 3 - 2023 | https://doi.org/10.3389/fbinf.2023.1215141
We present the results for CAPRI Round 54, the 5th joint CASP-CAPRI protein assembly prediction challenge. The Round offered 37 targets, including 14 homo-dimers, 3 homo-trimers, 13 hetero-dimers including 3 antibody-antigen complexes, and 7 large assemblies. On average ~70 CASP and CAPRI predictor groups, including more than 20 automatics servers, submitted models for each target. A total of 21941 models submitted by these groups and by 15 CAPRI scorer groups were evaluated using the CAPRI model quality measures and the DockQ score consolidating these measures. The prediction performance was quantified by a weighted score based on the number of models of acceptable quality or higher submitted by each group among their 5 best models. Results show substantial progress achieved across a significant fraction of the 60+ participating groups. High-quality models were produced for about 40% for the targets compared to 8% two years earlier, a remarkable improvement resulting from the wide use of the AlphaFold2 and AlphaFold-Multimer software. Creative use was made of the deep learning inference engines affording the sampling of a much larger number of models and enriching the multiple sequence alignments with sequences from various sources. Wide use was also made of the AlphaFold confidence metrics to rank models, permitting top performing groups to exceed the results of the public AlphaFold-Multimer version used as a yard stick. This notwithstanding, performance remained poor for complexes with antibodies and nanobodies, where evolutionary relationships between the binding partners are lacking, and for complexes featuring conformational flexibility, clearly indicating that the prediction of protein complexes remains a challenging problem.
We present VoroIF-GNN, a novel single-model method for assessing inter-subunit interfaces in protein-protein complexes. Given a multimeric protein structural model, we derive interface contacts from the Voronoi tessellation of atomic balls, construct a graph of those contacts, and predict accuracy of every contact using an attention-based graph neural network. The contact-level predictions are then summarized to produce whole interface-level scores. VoroIF-GNN was blindly tested for its ability to estimate accuracy of protein complexes during CASP15 and showed strong performance in selecting the best multimeric model out of many. The method implementation is freely available at https://klimentolechnovic.github.io/voronota/expansion_js/ .
The widespread TnpB proteins of IS200/IS605 transposon family have recently emerged as the smallest RNA-guided nucleases capable of targeted genome editing in eukaryotic cells 1 , 2 . Bioinformatic analysis identified TnpB proteins as the likely predecessors of Cas12 nucleases 3 – 5 , which along with Cas9 are widely used for targeted genome manipulation. Whereas Cas12 family nucleases are well characterized both biochemically and structurally 6 , the molecular mechanism of TnpB remains unknown. Here we present the cryogenic-electron microscopy structures of the Deinococcus radiodurans TnpB–reRNA (right-end transposon element-derived RNA) complex in DNA-bound and -free forms. The structures reveal the basic architecture of TnpB nuclease and the molecular mechanism for DNA target recognition and cleavage that is supported by biochemical experiments. Collectively, these results demonstrate that TnpB represents the minimal structural and functional core of the Cas12 protein family and provide a framework for developing TnpB-based genome editing tools.
A DNA damage-inducible mutagenic gene cassette has been implicated in the emergence of drug resistance in Mycobacterium tuberculosis during anti-tuberculosis (TB) chemotherapy. However, the molecular composition and operation of the encoded 'mycobacterial mutasome' - minimally comprising DnaE2 polymerase and ImuA' and ImuB accessory proteins - remain elusive. Following exposure of mycobacteria to DNA damaging agents, we observe that DnaE2 and ImuB co-localize with the DNA polymerase III β subunit (β clamp) in distinct intracellular foci. Notably, genetic inactivation of the mutasome in an imuBAAAAGG mutant containing a disrupted β clamp-binding motif abolishes ImuB-β clamp focus formation, a phenotype recapitulated pharmacologically by treating bacilli with griselimycin and in biochemical assays in which this β clamp-binding antibiotic collapses pre-formed ImuB-β clamp complexes. These observations establish the essentiality of the ImuB-β clamp interaction for mutagenic DNA repair in mycobacteria, identifying the mutasome as target for adjunctive therapeutics designed to protect anti-TB drugs against emerging resistance.
Reliably scoring and ranking candidate models of protein complexes and assigning their oligomeric state from the structure of the crystal lattice represent outstanding challenges. A community-wide effort was launched to tackle these challenges. The latest resources on protein complexes and interfaces were exploited to derive a benchmark dataset consisting of 1677 homodimer protein crystal structures, including a balanced mix of physiological and non-physiological complexes. The non-physiological complexes in the benchmark were selected to bury a similar or larger interface area than their physiological counterparts, making it more difficult for scoring functions to differentiate between them. Next, 252 functions for scoring protein-protein interfaces previously developed by 13 groups were collected and evaluated for their ability to discriminate between physiological and non-physiological complexes. A simple consensus score generated using the best performing score of each of the 13 groups, and a cross-validated Random Forest (RF) classifier were created. Both approaches showed excellent performance, with an area under the Receiver Operating Characteristic (ROC) curve of 0.93 and 0.94, respectively, outperforming individual scores developed by different groups. Additionally, AlphaFold2 engines recalled the physiological dimers with significantly higher accuracy than the non-physiological set, lending support to the reliability of our benchmark dataset annotations. Optimizing the combined power of interface scoring functions and evaluating it on challenging benchmark datasets appears to be a promising strategy.
The mycobacterial mutasome - minimally comprising ImuA', ImuB, and DnaE2 proteins - has been implicated in DNA damage-induced mutagenesis in Mycobacterium tuberculosis. ImuB, predicted to enable mutasome function via its interaction with the β clamp, is a catalytically inactive member of the Y-family of DNA polymerases. Like other members of the Y family, ImuB features a recently identified amino acid motif with homology to the RecA-N-terminus (RecA-NT). In RecA, the motif mediates oligomerization of RecA monomers into RecA filaments. Given the role of ImuB in the mycobacterial mutasome, we hypothesized that the ImuB RecA-NT motif might mediate its interaction with ImuA', a RecA homolog of unknown function. To investigate this possibility, we constructed a panel of imuB alleles in which RecA-NT was removed, or mutated. Results from microbiological and biochemical assays indicate that RecA-NT is critical for the interaction of ImuB with ImuA'. A region downstream of RecA-NT (ImuB-C) also appears to stabilize the ImuB-ImuA' interaction, but its removal does not prevent complex formation. In contrast, replacing two key hydrophobic residues of RecA-NT, L378 and V383, is sufficient to disrupt ImuA'-ImuB interaction. To our knowledge, this constitutes the first experimental evidence showing the role of the RecA-NT motif in mediating the interaction between a Y-family member and a RecA homolog.