MOTIVATION:Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS:We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core Cα RMSDs of 0.62 Å for MHC-I and 0.33 Å for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63 817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION:PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.
With the recent breakthrough of highly accurate structure prediction methods, there has been a rapid growth of available protein structures. Efficient methods are needed to infer structural similarity within these datasets. We present an end-to-end alignment method, called SoftAlign, which takes as input the 3D coordinates of a protein pair and outputs a structural alignment. In addition to the traditional Smith-Waterman alignment method, we introduce a modified softmax alignment that shows very promising results for structure similarity detection. We demonstrate that the SoftAlign model is able to recapitulate TM-align alignments while running faster, and it is more accurate than Foldseek on alignment and classification tasks. Although SoftAlign is not the fastest method available, it is highly precise and can be used effectively with other prefilters. In addition to developing an end-to-end structural aligner, our main contribution is the introduction and analysis of a pseudo-alignment method based on softmax, which can be used with other architectures, even those not based on structural information. The code for SoftAlign is available at . ### Competing Interest Statement The authors have declared no competing interest.
Advances in computational structure prediction will vastly augment the hundreds of thousands of currently available protein complex structures. Translating these into discoveries requires aligning them, which is computationally prohibitive. Foldseek-Multimer computes complex alignments from compatible chain-to-chain alignments, identified by efficiently clustering their superposition vectors. Foldseek-Multimer is 3–4 orders of magnitudes faster than the gold standard, while producing comparable alignments; this allows it to compare billions of complex pairs in 11 h. Foldseek-Multimer is open-source software available at GitHub via https://github.com/steineggerlab/foldseek/ , https://search.foldseek.com/search/ and the BFMD database. Foldseek-Multimer offers a fast strategy for complex-to-complex alignment to quickly identify compatible sets of chain-to-chain alignments by their superpositions. It can compare billions of complex pairs in 11 h.
Most popular tools for reconstructing phylogenetic trees from multiple sequence alignments use a model of molecular evolution in which a single substitution matrix or a small set of fixed matrices are shared between all columns. Models with column-specific rate matrices can in principle be fit by automatic differentiation methods, but in practice the heavy computational burden associated with computing the gradients of the many matrix exponentials has hindered exploration of such models. Here, we present a highly efficient approach for reverse-mode differentiation of the log likelihood computed with Felsenstein’s algorithm under any time-reversible substitution model. PhyloGrad is implemented in Rust and has Python bindings to easily combine it with automatic differentiation tools. Depending on the tree size, PhyloGrad is 30-100 times faster than automatic differentiation in Pytorch and uses 10-100 times less memory. Even in the task of fitting one global model it is still at least 10 times faster than IQ-TREE3. PhyloGrad accelerates current model optimizations and enables the field to easily explore and implement novel site-specific models.
De novo assembly of ancient metagenomic datasets is a challenging task. Ultra-short fragment size and characteristic postmortem damage patterns of sequenced ancient DNA molecules leave current tools ill-equipped for ideal assembly. We present CarpeDeam, a novel damage-aware de novo assembler designed specifically for ancient metagenomic samples. Utilizing maximum-likelihood frameworks that integrate sample-specific damage patterns, CarpeDeam demonstrates improved recovery of longer continuous sequences and protein sequences in many simulated and empirical datasets compared to existing assemblers. As a pioneering ancient metagenome assembler, CarpeDeam opens the door for new opportunities in functional and taxonomic analyses of ancient microbial communities.
The ubiquitous availability of protein structures permits replacing sequence alignment with more accurate and sensitive structure alignment algorithms. LoL-align maximizes a lo cal log-odds score for proteins to be homologous, given their intra-protein C α – C α distances. LoL-align is markedly more sensitive in detecting remote homologs than TMalign and DALI on single- and multi-domain benchmarks and achieves better alignment quality at 5-20× their speed. LoL-align is fully integrated into Foldseek.
Several recent deep learning methods for metagenome binning claim improvements in the recovery of high-quality metagenome-assembled genomes. These methods differ in their approaches to learn the contig embeddings and to cluster them. Rapid advances in binning require rigorous benchmarking to evaluate the effectiveness of new methods. We have benchmarked newly developed state-of-the-art deep learning binners on CAMI2 and real metagenomic datasets. The results show that SemiBin2 and COMEBin give the best binning performance, although not always the best embedding accuracy. Interestingly, post-binning reassembly consistently improves the quality of low-coverage bins. We find that binning coassembled contigs with multi-sample coverage is effective for low-coverage dataset, while binning sample-wise assembled contigs with multi-sample coverage (multi-sample) is effective for high-coverage samples. In multi-sample binning, splitting the embedding space by sample before clustering showed enhanced performance compared with the standard approach of splitting final clusters by sample. Deep-learning binners using contrastive models emerged as the top-performing tools overall, with MetaBAT2 and GenomeFace demonstrating superior speed. To facilitate future development, we provide workflows for standardized benchmarking of metagenome binners.
Metagenomics has revolutionized environmental and human-associated microbiome studies. However, the limited fraction of proteins with known biological processes and molecular functions presents a major bottleneck. In prokaryotes and viruses, evolution favors keeping genes participating in the same biological processes colocalized as conserved gene clusters. Conversely, conservation of gene neighborhood indicates functional association. Here we present Spacedust, a tool for systematic, de novo discovery of conserved gene clusters. To find homologous protein matches, Spacedust uses fast and sensitive structure comparison with Foldseek. Partially conserved clusters are detected using novel clustering and order conservation P values. We demonstrate Spacedust's sensitivity with an all-versus-all analysis of 1,308 bacterial genomes, identifying 72,843 conserved gene clusters containing 58% of the 4.2 million genes. It recovered 95% of antiviral defense system clusters annotated by the specialized tool PADLOC. Spacedust's high sensitivity and speed will facilitate the annotation of large numbers of sequenced bacterial, archaeal and viral genomes.
As structure prediction methods are generating millions of publicly available protein structures, searching these databases is becoming a bottleneck. Foldseek aligns the structure of a query protein against a database by describing tertiary amino acid interactions within proteins as sequences over a structural alphabet. Foldseek decreases computation times by four to five orders of magnitude with 86%, 88% and 133% of the sensitivities of Dali, TM-align and CE, respectively.
The DescribePROT database of amino acid-level descriptors of protein structures and functions was substantially expanded since its release in 2020. This expansion includes substantial increase in the size, scope, and quality of the underlying data, the addition of experimental structural information, the inclusion of new data download options, and an upgraded graphical interface. DescribePROT currently covers 19 structural and functional descriptors for proteins in 273 reference proteomes generated by 11 accurate and complementary predictive tools. Users can search our resource in multiple ways, interact with the data using the graphical interface, and download data at various scales including individual proteins, entire proteomes, and whole database. The annotations in DescribePROT are useful for a broad spectrum of studies that include investigations of protein structure and function, development and validation of predictive tools, and to support efforts in understanding molecular underpinnings of diseases and development of therapeutics. DescribePROT can be freely accessed at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/.
Summary:The annotation of deeply sequenced, de novo assembled transcriptomes continues to be a challenge as some of the state-of-the-art tools are slow, difficult to install, and hard to use. We have tackled these issues with TransAnnot, a fast, automated transcriptome annotation pipeline that is easy to install and use. Leveraging the fast sequence searches provided by the MMseqs2 suite, TransAnnot offers one-step annotation of homologs from Swiss-Prot, gene ontology terms and orthogroups from eggNOG, and functional domains from Pfam. Users also have the option to annotate against custom databases. TransAnnot accepts sequencing reads (short and long), nucleotide sequences, or amino acid sequences as input for annotation. When benchmarked with test data sets of amino acid sequences, TransAnnot was 333, 284, and 18 times faster than comparable tools such as EnTAP, Trinotate, and eggNOG-mapper respectively. Availability and implementation:TransAnnot is free to use, open sourced under GPLv3, and is implemented in C++ and Bash. Source code, documentation, and pre-compiled binaries are available at https://github.com/soedinglab/transannot. TransAnnot is also available via bioconda (https://anaconda.org/bioconda/transannot).
The study of the deluge of metagenomic and genomic sequences is challenging due to the severe lack of function information. Predicting operons, groups of functionally related genes in prokaryotic genomes, is critical for bridging this gap. However, existing methods for operon prediction heavily rely on experimental data, functional annotations, or extensive characterization of homologous genes, making it difficult to accurately predict operons in newly sequenced or poorly characterized genomes. Here, we introduce UniOP, an unsupervised approach that uses a statistical model to predict operons from intergenic distances directly derived from the target genomic sequence. UniOP not only outperforms alternative approaches on ten complete genomes but also shows superior results on 3269 metagenome-assembled genomes across 13 bacterial and 2 archaeal phyla. Furthermore, we explored enhancing UniOP by incorporating the conservation of gene neighborhood and strandedness in respective genomes and examined the influence of Pfam annotations and motif searching on its performance. ### Competing Interest Statement The authors have declared no competing interest.
Metagenomics is a powerful approach to study environmental and human-associated microbial communities and, in particular, the role of viruses in shaping them. Viral genomes are challenging to assemble from metagenomic samples due to their genomic diversity caused by high mutation rates. In the standard de Bruijn graph assemblers, this genomic diversity leads to complex k-mer assembly graphs with a plethora of loops and bulges that are challenging to resolve into strains or haplotypes because variants more than the k-mer size apart cannot be phased. In contrast, overlap assemblers can phase variants as long as they are covered by a single read. Here, we present PenguiN, a software for strain resolved assembly of viral DNA and RNA genomes and bacterial 16S rRNA from shotgun metagenomics. Its exhaustive detection of all read overlaps in linear time combined with a Bayesian model to select strain-resolved extensions allow it to assemble longer viral contigs, less fragmented genomes, and more strains than existing assembly tools, on both real and simulated datasets. We show a 3–40-fold increase in complete viral genomes and a 6-fold increase in bacterial 16S rRNA genes. PenguiN is the first overlap-based assembler for viral genome and 16S rRNA assembly from large and complex metagenomic datasets, which we hope will facilitate studying the key roles of viruses in microbial communities.
Zooplankton are important eukaryotic constituents of marine ecosystems characterized by limited motility in the water. These metazoans predominantly occupy intermediate trophic levels and energetically link primary producers to higher trophic levels. Through processes including diel vertical migration (DVM) and production of sinking pellets they also contribute to the biological carbon pump which regulates atmospheric CO2 levels. Despite their prominent role in marine ecosystems, and perhaps, because of their staggering diversity, much remains to be discovered about zooplankton biology. In particular, the circadian clock, which is known to affect important processes such as DVM has been characterized only in a handful of zooplankton species. We present annotated de novo assembled transcriptomes from a diverse, representative cohort of 17 marine zooplankton representing six phyla and eight classes. These transcriptomes represent the first sequencing data for a number of these species. Subsequently, using translated proteomes derived from this data, we demonstrate in silico the presence of orthologs to most core circadian clock proteins from model metazoans in all sequenced species. Our findings, bolstered by sequence searches against publicly available data, indicate that the molecular machinery underpinning endogenous circadian clocks is widespread and potentially well conserved across marine zooplankton taxa.
Numerous cellular processes rely on the binding of proteins with high affinity to specific sets of RNAs. Yet most RNA-binding domains display low specificity and affinity in comparison to DNA-binding domains. The best binding motif is typically only enriched by less than a factor 10 in high-throughput RNA SELEX or RNA bind-n-seq measurements. Here, we provide insight into how cooperative binding of multiple domains in RNA-binding proteins (RBPs) can boost their effective affinity and specificity orders of magnitude higher than their individual domains. We present a thermodynamic model to calculate the effective binding affinity (avidity) for idealized, sequence-specific RBPs with any number of RBDs given the affinities of their isolated domains. For seven proteins in which affinities for individual domains have been measured, the model predictions are in good agreement with measurements. The model also explains how a two-fold difference in binding site density on RNA can increase protein occupancy 10-fold. It is therefore rationalized that local clusters of binding motifs are the physiological binding targets of multi-domain RBPs.
Genetic CLN5 variants are associated with childhood neurodegeneration and Alzheimer’s disease; however, the molecular function of ceroid lipofuscinosis neuronal protein 5 (Cln5) is unknown. We solved the Cln5 crystal structure and identified a region homologous to the catalytic domain of members of the N1pC/P60 superfamily of papain-like enzymes. However, we observed no protease activity for Cln5; and instead, we discovered that Cln5 and structurally related PPPDE1 and PPPDE2 have efficient cysteine palmitoyl thioesterase (S-depalmitoylation) activity using fluorescent substrates. Mutational analysis revealed that the predicted catalytic residues histidine-166 and cysteine-280 are critical for Cln5 thioesterase activity, uncovering a new cysteine-based catalytic mechanism for S-depalmitoylation enzymes. Last, we found that Cln5-deficient neuronal progenitor cells showed reduced thioesterase activity, confirming live cell function of Cln5 in setting S-depalmitoylation levels. Our results provide new insight into the function of Cln5, emphasize the importance of S-depalmitoylation in neuronal homeostasis, and disclose a new, unexpected enzymatic function for the N1pC/P60 superfamily of proteins.
SUMMARY The core promoter, the region immediately surrounding the transcription start site, plays a central role in setting metazoan gene expression levels, but how exactly it ‘computes’ expression remains poorly understood. To dissect core promoter function, we carried out a comprehensive structure-function analysis to measure synthetic promoters’ activities, with and without an external stimulus (hormonal activation). By using robotics and a dual-luciferase reporter assay, we tested ∼3000 mutational variants representing 19 different Drosophila melanogaster promoter architectures. We explored the impact of different types of mutations, including knockout of individual sequence motifs and motif combinations, variations of motif strength, positioning, and flanking sequences. We observe strong effects of the mutations on activity, and a linear combination of the individual motif features can largely account for the combinatorial effects on core promoter activity. Our findings shed new light on the quantitative assessment of gene expression, a fundamental process in all metazoans.
SummarySpacePHARER (CRISPR Spacer Phage-Host Pair Finder) is a sensitive and fast tool forde novoprediction of phage-host relationships via identifying phage genomes that match CRISPR spacers in genomic or metagenomic data. SpacePHARER gains sensitivity by comparing spacers and phages at the protein level, optimizing its scores for matching very short sequences, and combining evidence from multiple matches, while controlling for false positives. We demonstrate SpacePHARER by searching a comprehensive spacer list against all complete phage genomes.