MOTIVATION:Protein kinases are key regulators of cellular signaling and are frequently implicated in human diseases. Although kinase domains are structurally conserved, predicting the effects of amino acid substitutions remains challenging as mutations often introduce subtle structural perturbations that are not captured by sequence-based or evolutionary methods. Existing supervised approaches further rely on pathogenicity annotations that are inconsistent across databases, thereby motivating the development of structure-based, label-independent frameworks for mutation effect prediction. RESULTS:We present a structure-based method using SE(3)-transformers to learn residue compatibility with the local structural environment from experimentally resolved kinase 3D structures. Proteins are represented as atom-level graphs with physicochemical descriptors derived from the CHARMM force field and spatial connectivity. The model is trained on two self-supervised tasks given local structural context: masked residue atom reconstruction and masked residue classification. This formulation enables learning of geometric and physicochemical constraints without relying on pathogenicity labels. Evaluation using reconstruction loss, residue prediction accuracy, and comparison with BLOSUM substitution patterns indicate that the model captures biologically meaningful relationships between residue identity and 3D structural context. We interpret the scores assigned to alternative amino acids as measures of structural fitness, where low-scoring residues are hypothesized to be less compatible with the local environment and more likely to induce deleterious effects on protein structure and activity. AVAILABILITY AND IMPLEMENTATION:https://zenodo.org/records/20393799.
Abstract Motivation Molecular docking is a pillar of structure-based drug design and shows advantages in structure prediction of small-molecule ligand–protein complexes over co-folding methods for novel ligands and novel binding pockets. Here, we describe substantial improvements of our physics-based docking algorithm Attracting Cavities, which is widely used through the SwissDock webserver. Results AC 3.0 includes enhanced sampling features, new functionalities, and technical improvements. These lead to better sampling at lower execution times and higher versatility. Comparison with AutoDock Vina demonstrates better docking results on multiple test sets. Availability AC 3.0 will be made available free of charge through the SwissDock webserver ( www.swissdock.ch ).
The recent development of highly accurate protein structure prediction tools has led to a rapid expansion in the scope of computational structural biology, enabling a much wider range of modelling studies than ever before. These new in silico opportunities help life science researchers understand how proteins interact with their environment and support design of new molecules with desired properties. Ultimately, they have broad applications, e.g. in medicine, drug discovery or engineering. To ensure reproducibility and to facilitate data exchange and reuse, predicted structures or computed structure models can be stored using ModelCIF, a rich data representation designed to include the atomic coordinates/metadata. The previously published version of ModelCIF (1.4.4; 2022-12-21) mainly covered protein structure predictions generated by homology and ab initio modelling. In this work, we present an extension of the ModelCIF (https://github.com/ihmwg/ModelCIF) data standard and its associated tools. This extension supports important new use cases, including modelling protein-ligand and protein-protein interactions, sampling multiple conformational states and designing proteins de novo. We define guidelines for storage and validation of modelling results for those use cases by applying new and existing ModelCIF categories to capture protocols, inputs and outputs. Additionally, we outline updates to the software tools and resources that implement these new standards and provide functionality for model generation, validation, archiving, and visualisation. By enabling consistent metadata capture across different modelling workflows, this framework aims to support the FAIR dissemination of computational models, thereby promoting reproducibility and reusability in downstream applications.
The binding of antibodies (Abs) to antigens (Ags) is a fundamental mechanism by which the immune system defends the body against harmful invaders. Understanding and predicting this interaction is critical for developing effective therapeutic Abs. In this study, we performed an extensive structural analysis of the largest available non-redundant three-dimensional (3D) database of Ab-Ag complexes. By examining the 3D structures of these complexes, we identified key amino acids (AAs) and their positions within the paratope-epitope interface. We stressed the significance of examining AAs at each complementarity-determining regions (CDR) position, as their interaction frequencies at specific sites can vary greatly from their overall tendency to appear in CDRs. Importantly, we standardized all the paratope-epitope interaction patterns identified, converting them into a format that can be easily utilized by researchers outside the structural biology field, especially those working on antibody development using machine learning. Our findings provide valuable insights for optimizing Abs, particularly in therapeutic applications, where refining the CDR loops to improve Ag binding is essential for creating effective Abs. This research may also support the development of more precise and reliable Ab-Ag prediction tools.
Human protein kinases constitute a large superfamily of about 500 genes, historically classified into subfamilies based on phylogenetic relationship. However, many kinases remain unclassified. Phylogeny is typically based on multiple sequence alignments, and neglects the physico-chemical properties of residues at each position of the sequence. By incorporating these properties, we can gain deeper insights beyond basic alignments. Here we use, for the first time, a detailed physico-chemical description of kinases to identify class-specific structural regions, supporting an unsupervised classification method capable of classifying previously unlabeled kinases. This novel approach aligns with existing phylogeny-based classifications while offering refinements and enhanced accuracy. Ultimately, we use machine learning techniques to classify unlabeled kinases, validated by analyzing class-specific structural regions. This new classification approach goes beyond current rankings and can be applied to any type of protein, such as immunoglobulins and G protein-coupled receptors.
PURPOSE:The rapid expansion of precision oncology has led to a marked increase in the number of identified oncodriver genes and associated variants. This surge has elevated the frequency of mutations with unknown functional impact, highlighting the growing need for molecular modeling tools such as Swiss-PO to support variant interpretation for both clinical and research contexts. MATERIALS AND METHODS:Swiss-PO was expanded and redesigned to integrate large-scale oncogenomic data with structure-based and sequence-based analytical tools. The platform combines data of oncodriver genes, multiple sequence alignment (MSA) for conservation analysis across orthologs and gene families, experimental and predicted three-dimensional (3D) structures prediction visualization to assess impacts of mutations. Additional modules include a BRAF kinase mutation classification model and a dedicated webpage for ligands targeting proteins, and unified integration with external databases to enable multidimensional analyses. RESULTS:The updated Swiss-PO platform integrates data for nearly 1,500 oncodriver genes, encompassing more than 3 million mutations and post-translational modification annotations, over 26,000 experimental and predicted 3D protein structures, more than 4,000 MSAs , and information on over 200,000 protein ligands. In addition, the platform includes a BRAF kinase mutation class predictor to support therapeutic decision-making, establishing Swiss-PO as one of the most comprehensive publicly available resources for precision oncology. CONCLUSION:With these enhancements, Swiss-PO strengthens its role as a powerful and versatile resource for oncologists, bioinformaticians, and molecular biologists engaged in the interpretation of cancer-associated mutations and the advancement of precision medicine. The website is available at Swiss-PO.
Outcomes after CD19 CAR T-cell therapy in patients with diffuse large B-cell lymphoma are variable, partly influenced by the germline CD19 SNP rs2904880 (V174/L174). Using structural modeling, lymphoma cell lines, and patient data, we demonstrate that CAR hinge choice (CD8 versus CD28) differentially modulates avidity and cytotoxicity against CD19 V174 and V174L variants and correlates with distinct clinical responses to axi-cel and tisa-cel. This suggests that CD19 genotyping may guide CAR T selection.
Type 1 diabetes (T1D) is marked by the overexpression of class I major histocompatibility complex (MHC) antigens in pancreatic islets, which are targeted by islet-specific CD8+ T cells. Here, we aimed to improve regulatory T cell (Treg) infiltration into pancreatic islets by redirecting their specificity toward class I-restricted islet antigens. We functionally validated two islet-specific HLA-A2 (∗02:01)-restricted T cell receptors (TCRs), one specific for ZnT8186-194 (clone D222D) and the second for IGRP265-273 (clone 32), by dual-locus (TRAC/CD4) homology-directed editing. Clone D222D was peptide specific and CD8αβ dependent, while clone 32 exhibited antigen promiscuity and showed CD8α dependency. Engineered CD4to8 TCR Tregs maintained stable phenotypes, in vitro suppressive function compared to their polyclonal counterparts, and showed co-receptor-dependent migration in vivo. This approach demonstrates that TCR specificity, reflected by its functional activity, is crucial for tissue-specific trafficking, paving the way to improve the efficacy of Treg therapies for T1D.
MOTIVATION:Molecular docking is a pillar of structure-based drug design and shows advantages over co-folding by providing physical insights into small-molecule ligand-target interactions, independent of ligand and pocket novelty. Here, we describe substantial improvements of our docking algorithm Attracting Cavities, which is widely used through the SwissDock webserver. RESULTS:AC 3.0 includes enhanced sampling features, new functionalities, and technical improvements. These lead to better sampling at lower execution times and higher versatility. Comparison with AutoDock Vina demonstrates better docking results on multiple test sets. AVAILABILITY AND IMPLEMENTATION:AC 3.0 will be made freely available through the SwissDock webserver (swissdock.ch).
IntroductionThe development of cancer immunotherapy has accelerated in recent years. Understanding the specificity of T cell receptors (TCR) for peptides presented by the major histocompatibility complex (pMHC) is a critical step towards improving immunotherapy approaches, such as adoptive cell transfer and peptide vaccination. Despite notable computational advances, the unambiguous pairing of TCR with pMHC, from pools of thousands of candidates and unseen pMHC, remains elusive.MethodsTo meet this challenge and showcase the potential of using physics-based structure-based methods without being hindered by their computational cost, we developed a novel approach, TCRfp. This method transforms the 3D structure of TCRs into one-dimensional structural fingerprints (FPs) using the electroshape 5D (ES5D) technique.ResultsWe have modelled more than 15’000 3D structures of paired TCR alpha and beta chains with known sequences and pMHC specificity and encoded them into 1D TCRfp. Anticipating future clinical applications, we have translated the TCR modelling process into a fast pipeline. Similarity measures between TCR FPs correlate with their ability to recognize similar or identical epitopes within both the training set and in the external validation sets.DiscussionTCRfp constitutes a rapid approach for high-throughput TCR comparison and repertoire analysis based on molecular 3D structures. When tested on a private dataset and combined with a basic sequence-based method via logistic regression, TCRfp surpassed existing approaches in predicting TCR specificities. TCRfp represents a structurally informed complement to sequence-based approaches and could enhance our ability to decode immune recognition.
Deep-learning based co-folding methods predict the structures of proteins interacting with metal ions, small molecules, nucleic acids, peptides, and other proteins. One of their main objectives is their application for drug design, predicting the structure of small-molecule ligand/protein complexes. It has been shown that at present these models memorize ligand poses from the training data but do not generalize effectively to novel complexes and lack in the adherence to physical and chemical principles. Here, we use the recently introduced Runs N' Poses benchmark set of 2,600 protein-ligand systems annotated by their similarity to the training data, to show that the physics-based docking algorithms Attracting Cavities and AutoDock Vina outperform co-folding methods for novel ligands and novel binding pockets. In addition to predicting ligand poses and at variance with co-folding methods, they provide a physical rationale on why a ligand binds (or does not bind) and insight into experimental structural model deficiencies. ### Competing Interest Statement The authors have declared no competing interest.
AlphaFold (AF), a deep-learning based protein modelling approach, has revolutionized structural biology by generally achieving near-experimental accuracy in protein structure prediction. An impactful application of AF is the modelling of antibodies (Abs) and T-cell receptors (TCRs), key mediators of cellular immunity, whose structural specificity underlies responses in cancer, infection, and autoimmune diseases. In this work, we analyse AF3 performance by systematically examining how MSA composition and the number of inference phases affect prediction accuracy and computational efficiency. Using reduced UniRef90 subsets, we present a ∼45-fold accelerated and highly accurate variant of the AF3 workflow specifically optimized for the modelling of the Abs and TCRs receptor domains, enabling rapid, reliable structural predictions at a scale suitable for high-throughput immunological studies. Our findings provide a foundation for faster therapeutic discovery and deeper molecular mechanism understanding of immune recognition.
Type 1 diabetes (T1D) is marked by the overexpression of class I major histocompatibility complex (MHC) antigens in pancreatic islets, which are targeted by islet-specific CD8+ T cells. Here, we aimed to improve regulatory T cell (Treg) infiltration into pancreatic islets by redirecting their specificity toward class I-restricted islet antigens. We functionally validated two public islet specific HLA-A2 (*02:01) restricted TCRs, one specific for ZnT8186-194 (clone D222D), the second for IGRP265-273 (clone 32) by dual locus (TRAC/CD4) homology-directed editing. Clone D222D was peptide-specific and CD8β dependent while clone 32 exhibited antigen promiscuity and showed CD8α dependency. Engineered CD4to8 TCR Tregs maintained stable phenotypes, suppressed significantly better than their polyclonal counterpart, and showed co-receptor-dependent migration in vivo. This approach demonstrates that TCR specificity, reflected by its functional activity, is crucial for tissue-specific trafficking, paving the way to improve the efficacy of Treg therapies for T1D. ### Competing Interest Statement QT is a co-founder and scientific advisor of Sonoma Biotherapeutics.
While cancer immunotherapy has primarily focused on CD8 T cells, CD4 T cells are increasingly recognized for their role in antitumor immunity. The HLA-DRB3*02:02 allele is found in 50% of Caucasians. In this study, we screened HLA-DRB3*02:02 patients with melanoma for tumor-specific CD4 T cells and identified robust New York esophageal squamous cell carcinoma 1 (NY-ESO-1)123-137/HLA-DRB3*02:02 CD4 T cell activity in both peripheral blood and tumor tissue. By analyzing NY-ESO-1123-137/HLA-DRB3*02:02-restricted CD4 T cell clones, we uncovered an unexpectedly high cytotoxicity, strong T helper 1 polarization, and recurrent αβ T cell receptor (TCRαβ) usage across patients and anatomical sites. These responses were also present in other NY-ESO-1-expressing cancers. TCRs from these clones, when transduced into primary CD4 T cells, showed direct antitumor efficacy both in vitro and in vivo. Our findings suggest that these TCRs are promising for adoptive T cell transfer therapy, enabling broader targeting of NY-ESO-1-expressing adult and pediatric cancers in clinical settings.
Interactions between T-Cell Receptors (TCRs) and antigenic peptides presented on Major Histocompatibility Complex (MHC) molecules are central to the immune recognition of infected and malignant cells. The complexity of the TCR sequence space, the flexibility of the TCR-epitope interface and the lack of a standardized framework to visualize TCR specificity have hindered a comprehensive understanding of the principles governing TCR-epitope recognition across a broad range of epitopes. Here, we introduce a fully interpretable probabilistic framework, termed TCR specificity profiles (TSPs), which captures fundamental properties distinguishing epitope-specific from baseline TCR repertoires. We demonstrate that TSPs unravel key determinants of TCR-epitope recognition specificity. By identifying and analyzing TCRs recognizing dozens of epitope variants, we show that TSPs accurately predict cross-reactivity and reveal how TCR specificity evolves with epitope sequence, binding mode and MHC restriction. TSPs further enable interpretation of machine learning tools and reveal how AlphaFold3 can be used to decipher key determinants of TCR-epitope recognition specificity. ### Competing Interest Statement The authors have declared no competing interest. Swiss National Science Foundation, CRSII5_193749, 320030-231333
As (deep) generative chemistry rapidly enters the landscape of drug discovery, evaluating the models and their structural output relevance, tractability, and innovation potential becomes both a scientific and practical challenge. Existing metrics such as QED and SA Score, though widely adopted, are rooted in biased historical datasets and often fail to capture what medicinal chemists truly seek: compounds with meaningful pharmacological potential and realistic tractability. In this work, we introduce the OiiSTER-map - a simple, intuitive, and interpretable two-dimensional heatmap that classifies molecules based on bioactivity-informed usuality and structural elaboration derived from molecular fingerprints. Unlike traditional filters, the OiiSTER-map helps identify not only the drug discovery chemical 'sweet spot'. (Regular/Balanced), but also underappreciated territories such as Minimal/Unusual or Over-elaborated/Trivial regions, offering actionable insights into compound quality, relevance, and diversity. We hope that a bioactivity-informed, structurally aware, and easily interpretable tool like the OiiSTER-map, employed in combination with other well implemented metrics, can be decisive to go beyond current limitations of assessing (deep) generative models and to ensure a more mechanistically relevant, nuanced and useful evaluation of the thus designed virtual molecules. ### Competing Interest Statement The authors have declared no competing interest.
Estimating the cell line targets of cytotoxic small molecules is important for drug discovery and central for targeted therapy in oncology. Accurate prediction of sensitive cell lines enables early identification of efficacy and toxicity, optimization of drug selectivity, and can foster drug repurposing. While most bioactive compounds interact with multiple macromolecular targets, the cytotoxicity encompasses diverse complex biological and chemical mechanisms that could even not all be related to binding to macromolecules, making the prediction of cytotoxicity specificity particularly challenging. To address early-phase prediction of cancer cell line targets of cytotoxic compounds, we developed a method combining a machine-learning classification model with a ligand-based reverse screening procedure able to rank cell-line from the most probable to the least probable target of any cytotoxic molecule. The development focused on addressing the challenges related to the scarcity of available experimental data on non-cytotoxic compounds. A knowledge-guided generation of realistic alleged inactives allowed to train several binary logistic regression models. The most robust classification model was trained on 164,134 cytotoxic compounds extracted from ChEMBL to generate a score of predicted sensitivity of cell lines for any new cytotoxic molecule. The method demonstrated strong predictive ability, recovering at least one experimental target within the 15 most probable cell-lines for 71% of nearly 11,000 external cytotoxic compounds tested across 1018 cancer cell lines.