Determining the binding pose of a ligand to a protein, known as molecular docking, is a fundamental task in drug discovery. Generative approaches promise faster, improved, and more diverse pose sampling than physics-based methods, but are often hindered by chemically implausible outputs, poor generalisability, and high computational cost. To address these challenges, we introduce a novel fragmentation scheme, leveraging inductive biases from structural chemistry, to decompose ligands into rigid-body fragments. Building on this decomposition, we present SigmaDock, an SE(3) Riemannian diffusion model that generates poses by learning to reassemble these rigid bodies within the binding pocket. By operating at the level of fragments in SE(3), SigmaDock exploits well-established geometric priors while avoiding overly complex diffusion processes and unstable training dynamics. Experimentally, we show SigmaDock achieves state-of-the-art performance, reaching Top-1 success rates (RMSD<2 & PB-valid) above 79.9% on the PoseBusters set, compared to 12.7-30.8% reported by recent deep learning approaches, whilst demonstrating consistent generalisation to unseen proteins. SigmaDock is the first deep learning approach to surpass classical physics-based docking under the PB train-test split, marking a significant leap forward in the reliability and feasibility of deep learning for molecular modelling.
Fast, unconditional 3D generative models can now produce high-quality molecules, but adapting them for specific design tasks often requires costly retraining. To address this, we introduce Interpolate-Integrate and Replacement Guidance, two training-free, inference-time conditioning strategies that provide control over E(3)-equivariant flow-matching models. Our methods generate bioisosteric 3D molecules by conditioning on seed ligands or fragment sets to preserve key determinants like shape and pharmacophore patterns, without requiring the original fragment atoms to be present. We demonstrate their effectiveness on three drug-relevant tasks: natural product ligand hopping, bioisosteric fragment merging, and pharmacophore merging.
While conventional Transformers generally operate on sequence data, they can be used in conjunction with structure models, typically SE(3)-invariant or equivariant graph neural networks (GNNs), for 3D applications such as protein structure modelling. These hybrids typically involve either (1) preprocessing/tokenizing structural features as input for Transformers or (2) taking Transformer embeddings and processing them within a structural representation. However, there is evidence that Transformers can learn to process structural information on their own, such as the AlphaFold3 structural diffusion model. In this work we show that Transformers can function independently as structure models when passed linear embeddings of coordinates. We first provide a theoretical explanation for how Transformers can learn to filter attention as a 3D Gaussian with learned variance. We then validate this theory using both simulated 3D points and in the context of masked token prediction for proteins. Finally, we show that pre-training protein Transformer encoders with structure improves performance on a downstream task, yielding better performance than custom structural models. Together, this work provides a basis for using standard Transformers as hybrid structure-language models.
The flexibility of protein side chains is an essential contributor of conformational entropy and affects processes such as folding, stability and molecular interactions. Structure determination experiments and prediction tools such as AlphaFold generally fail to capture or represent the conformational heterogeneity of proteins in solution. Experiments can be used to study side-chain flexibility, but cannot be applied at scale, and most prediction methods focus on reconstructing the minimum free energy state rather than an ensemble representing side-chain configurations. Here, we use AlphaFold2 and its internal side-chain representations to develop AF2χ that predicts side-chain χ-angle distributions and generates structural ensembles. We extensively benchmark AF2χ predictions using experimental NMR 3 J -couplings and s 2 order parameters, as well as dihedral angle distributions derived from collections of experimental structures, demonstrating the accuracy of AF2χ in generating accurate side-chain ensembles. We also compare the accuracy of AF2χ with molecular dynamics simulations and recent machine learning models aimed to generate conformational ensembles and show that AF2χ provides state-of-the-art accuracy orders of magnitude faster than molecular simulations. With its speed and accuracy, AF2χ offers a strong complementary option to simulations and rotamer library approaches, making it particularly valuable for applications such as protein design, ligand docking and interpretation of biophysical experiments. ### Competing Interest Statement C.M.D. discloses membership of the Scientific Advisory Board of Fusion Antibodies and AI Proteins as well as being a founder of Dalton. K.L.-L. holds stock options in and is a consultant for Peptone Ltd. All other authors declare no conflict of interest.
MOTIVATION:Machine learning-based scoring functions (MLBSFs) have been found to exhibit inconsistent performance on different benchmarks and be prone to learning dataset bias. For the field to develop MLBSFs that learn a generalizable understanding of physics, a more rigorous understanding of how they perform is required. RESULTS:In this work, we compared the performance of a diverse set of popular MLBSFs (RFScore, SIGN, OnionNet-2, Pafnucy, and PointVS) to our proposed baseline models that can only learn dataset biases on a range of benchmarks. We found that these baseline models were competitive in accuracy to these MLBSFs in almost all proposed benchmarks, indicating these models only learn dataset biases. Our tests and provided platform, ToolBoxSF, will enable researchers to robustly interrogate MLBSF performance and determine the effect of dataset biases on their predictions. AVAILABILITY AND IMPLEMENTATION:https://github.com/guydurant/toolboxsf.
Anesthetisia is an important surgical and explorative tool in the study of consciousness. Much work has been done to connect the deeply anesthetized condition with decreased complexity. However, anesthesia-induced unconsciousness is also a dynamic condition in which functional activity and complexity may fluctuate, being perturbed by internal or external (e.g., noxious) stimuli. We use fMRI data from a cohort undergoing deep propofol anesthesia to investigate resting state dynamics using dynamic brain state models and spatiotemporal network analysis. We focus our analysis on group-level dynamics of brain state temporal complexity, functional activity, connectivity, and spatiotemporal modularization in deep anesthesia and wakefulness. We find that in contrast to dynamics in the wakeful condition, anesthesia dynamics are dominated by a handful of sink states that act as low-complexity attractors to which subjects repeatedly return. On a subject level, our analysis provides tentative evidence that these low-complexity attractor states appear to depend on subject-specific age and anesthesia susceptibility factors. Finally, our spatiotemporal analysis, including a novel spatiotemporal clustering of graphs representing hidden Markov models, suggests that dynamic functional organization in anesthesia can be characterized by mostly unchanging, isolated regional subnetworks that share some similarities with the brain's underlying structural connectivity, as determined from normative tractography data.
Therapeutic antibodies are manufactured, stored and administered in the free state; this makes understanding the unbound form key to designing and improving development pipelines. Prediction of unbound antibodies is challenging, specifically modelling of the CDRH3 loop, where inaccuracies are potentially worse due to a bias in structural data towards antibody-antigen complexes. This class imbalance provides a challenge for deep learning models trained on this data, potentially limiting generalisation to unbound forms. Here we discuss the importance of unbound structures in antibody development pipelines. We explore how the latest generation of structure predictors can provide new insights and assess how conformational heterogeneity may influence binding kinetics. We hypothesise that generative models may address some of these issues. While prediction of antibodies in complex is essential, we should not ignore the need for progress in modelling the unbound form.
Machine learning offers great promise for fast and accurate binding affinity predictions. However, current models lack robust evaluation and fail on tasks encountered in (hit-to-) lead optimisation, such as ranking the binding affinity of a congeneric series of ligands, thereby limiting their application in drug discovery. Here, we address these issues by first introducing a novel attention-based graph neural network model called AEV-PLIG (atomic environment vector-protein ligand interaction graph). Second, we introduce a new and more realistic out-of-distribution test set called the OOD Test. We benchmark our model on this set, CASF-2016, and a test set used for free energy perturbation (FEP) calculations, that not only highlights the competitive performance of AEV-PLIG, but provides a realistic assessment of machine learning models with rigorous physics-based approaches. Moreover, we demonstrate how leveraging augmented data (generated using template-based modelling or molecular docking) can significantly improve binding affinity prediction correlation and ranking on the FEP benchmark (weighted mean PCC and Kendall's tau increases from 0.41 and 0.26 to 0.59 and 0.42). These strategies together are closing the performance gap with FEP calculations (FEP+ achieves weighted mean PCC and Kendall's tau of 0.68 and 0.49 on the FEP benchmark) while being similar to 400,000 times faster.
Antigen receptor numbering allows the rapid delineation of the antigen-binding regions of antibody and T cell receptor (TCR) sequences, from sequence alone. It also allows the comparison of the vast diversity of antigen receptors in a consistent frame of reference. Numbering of antigen receptors is currently achieved by aligning sequences to a reference set. This approach may result in different numbering, depending on the reference set used or may fail to number query sequences derived from new species or rare sequence types. To address this problem, we have built a new numbering method (ANARCII) which requires no alignment step and is based on a Seq2Seq language model. Our results show that ANARCII can deal with the complexity that arises in experimentally collected sequencing data and generalise to sequences which are highly dissimilar to those in training. In test sets designed to contain challenging and ambiguous sequence patterns ANARCII numbering was identical to existing methods for over 99.99% of conserved residues and over 99.94% for complete CDR regions. The lightweight architecture allows numbering of over 90,000 sequences per minute on a single A100 GPU. Furthermore, the ANARCII package can be conditioned to fit rare sequence types and provide new training data for fine-tuning. We demonstrate that fine-tuned versions of ANARCII can correctly number other immunoglobulin domains such as TCRs and VNARs. Our model is freely available as a web tool (), as well as a package for high throughput numbering of next generation sequencing data (). ### Competing Interest Statement C.D. discloses membership of the Scientific Advisory Board of Fusion Antibodies and AI proteins, as well as a founder of Dalton. All other authors declare no conflict of interest.
Many proteins are highly flexible and their ability to adapt their shape can be fundamental to their functional properties. For example, the flexibility of antibody complementarity-determining region (CDR) loops influences binding affinity and specificity, making it a key factor in understanding and designing antigen interactions. With methods such as AlphaFold, it is possible to computationally predict a single, static protein structure with high accuracy. However, the reliable prediction of structural flexibility has not yet been achieved. A major factor limiting such predictions is the scarcity of suitable training data. Here we focus on predicting the structural flexibility of functionally important antibody and T cell receptor CDR3 loops. To this end, we constructed ALL-conformations by extracting CDR3s and CDR3-like loop motifs from all structures deposited in the Protein Data Bank. This dataset comprises 1.2 million loop structures representing more than 100,000 unique sequences and captures all experimentally observed conformations of these motifs. Using this dataset, we develop ITsFlexible, a deep learning tool with graph neural network architecture. We trained the model to binary classify CDR loops as 'rigid' or 'flexible' from inputs of antibody structures. ITsFlexible outperforms all alternative approaches on our crystal structure datasets and successfully generalizes to molecular dynamics simulations. We also used ITsFlexible to predict the flexibility of three CDRH3 loops with no solved structures and experimentally determined their conformations using cryogenic electron microscopy.
Generative models have emerged as potentially powerful methods for molecular design, yet challenges persist in generating molecules that effectively bind to the intended target. The ability to control the design process and incorporate prior knowledge would be highly beneficial for better tailoring molecules to fit specific binding sites. In this paper, we introduce MolSnapper, a novel tool that is able to condition diffusion models for structure-based drug design by seamlessly integrating expert knowledge in the form of 3D pharmacophores. We demonstrate through comprehensive testing on both the CrossDocked and Binding MOAD data sets that our method generates molecules better tailored to fit a given binding site, achieving high structural and chemical similarity to the original molecules. Additionally, MolSnapper yields approximately twice as many valid molecules as alternative methods.
Cryo-Electron Microscopy (cryo-EM) is a pivotal tool for determining the 3D structures of biological macromolecules. Current cryo-EM workflows, while effective, are computationally demanding and require manual intervention, creating bottlenecks for use in high-throughput scenarios such as structure-based drug discovery. Often in structure-based drug discovery, one can assume that all instances of a protein are equivalent at the resolutions needed for alignment and it therefore should be possible to harness information about particle poses from previous refinements. Current methods, however, do not leverage this form of prior knowledge, instead aligning each dataset from scratch. We present cryoPARES, a deep learning pose estimation method trained on pre-aligned datasets. Our method not only provides accurate angular predictions significantly faster than traditional approaches but also introduces automated particle pruning capabilities that eliminate manual intervention. These features, together with its single-pass operation, can enable real-time reconstructions that provide feedback during data acquisition. We demonstrate cryoPARES's effectiveness through the rapid structural determination of six ligand-bound complexes across three distinct protein targets and release three new fragment-bound cryo-EM datasets. ### Competing Interest Statement Alex Berndt, Amir Apelbaum, Judith Reeks, Pamela A Williams, Carl Poelking, and Michael Saur are employees of Astex Pharmaceuticals. Ruben Sanchez-Garcia is a Sustaining Innovation Postdoctoral Research Associate at Astex Pharmaceuticals.
Developing therapeutic antibodies is a challenging endeavor, often requiring large-scale screening to produce initial binders, that still often require optimization for developability. We present a computational pipeline for the discovery and design of therapeutic antibody candidates, which incorporates physics- and AI-based methods for the generation, assessment, and validation of candidate antibodies with improved developability against diverse epitopes, via efficient few-shot experimental screens. We demonstrate that these orthogonal methods can lead to promising designs. We evaluated our approach by experimentally testing a small number of candidates against multiple SARS-CoV-2 variants in three different tasks: (i) traversing sequence landscapes of binders, we identify highly sequence dissimilar antibodies that retain binding to the Wuhan strain, (ii) rescuing binding from escape mutations, we show up to 54% of designs gain binding affinity to a new subvariant and (iii) improving developability characteristics of antibodies while retaining binding properties. These results together demonstrate an end-to-end antibody design pipeline with applicability across a wide range of antibody design tasks. We experimentally characterized binding against different antigen targets, developability profiles, and cryo-EM structures of designed antibodies. Our work demonstrates how combined AI and physics computational methods improve productivity and viability of antibody designs.
Computational epitope profiling methods group antibodies that bind to the same epitope. They can be used to predict epitopes and reduce the number of antibodies that need to be characterized experimentally. Conventionally, computational epitope profiling is achieved by clustering antibodies by sequence similarity. While sequence-similar antibodies are likely to share a common function, such methods neglect that antibodies with highly diverse sequences can exhibit similar binding site geometries and engage common epitopes. The SPACE2 algorithm described here is an epitope profiling method that clusters antibodies based on the structural similarity of models predicted with a state-of-the-art protein structure prediction tool. SPACE2 accurately clusters antibodies that engage the same epitope and exhibits far greater data coverage than conventional methods. Furthermore, SPACE2 detects signals of functional convergence and, unlike sequence-based methods, is able to link antibodies diverse in sequence, genetic lineage, and species origin. These results reiterate that structural data provides orthogonal information to sequence and improves our ability to study antibodies and their epitopes.
Many proteins are highly flexible and their ability to adapt their shape can be fundamental to their functional properties. We can now computationally predict a single, static protein structure with high accuracy. However, we are not yet able to reliably predict structural flexibility. A major factor limiting such predictions is the scarcity of suitable training data. Here, we focus on predicting the structural flexibility of the functionally important antibody and T-cell receptor CDR3 loops. We extracted a dataset of CDR3 like loop motifs from the PDB to create ALL-conformations, a dataset containing 1.2 million structures and more than 100,000 unique sequences. Using this dataset, we develop ITsFlexible a method classifying CDR3 flexibility, which outperforms all alternative approaches on our crystal structure datasets and successfully generalises to MD simulations. We also used ITsFlexible to predict the flexibility of three completely novel CDRH3 loops and experimentally determined their conformations using cryo-EM. ### Competing Interest Statement The authors have declared no competing interest.
Immunomodulatory imide drugs (IMiDs), including thalidomide, lenalidomide, and pomalidomide, can be used to induce degradation of a protein of interest that is fused to a short degron motif, which often comprises a zinc finger (ZF). These IMiDs, however, also induce the degradation of endogenous ZF-containing neosubstrates, including IKZF1, IKZF3, and SALL4. To improve degradation selectivity, we took a bump-and-hole approach to design and screen bumped IMiD analogues against 8380 ZF mutants. This yielded a bumped IMiD analogue that induces efficient degradation of a mutant ZF degron, while not affecting other cellular proteins, including IKZF1, IKZF3, and SALL4. In proof-of-concept studies, this system was applied to induce degradation of the optimum degron fused to CDK9, HPRT1, NanoLuc, or TRIM28. We anticipate that this system will be a valuable addition to the current arsenal of degron systems for use in target validation.
T-cell receptor (TCR) structures are currently under-utilised in early-stage drug discovery and repertoire-scale informatics. Here, we leverage a large dataset of solved TCR structures from Immunocore to evaluate the current state-of-the-art for TCR structure prediction, and identify which regions of the TCR remain challenging to model. Through clustering analyses and the training of a TCR-specific model capable of large-scale structure prediction, we find that the alpha chain VJ-recombined loop (CDR3α) is as structurally diverse and correspondingly difficult to predict as the beta chain VDJ-recombined loop (CDR3β). This differentiates TCR variable domain loops from the genetically analogous antibody loops and supports the conjecture that both TCR alpha and beta chains are deterministic of antigen specificity. We hypothesise that the larger number of alpha chain joining genes compared to beta chain joining genes compensates for the lack of a diversity gene segment. We also provide over 1.5M predicted TCR structures to enable repertoire structural analysis and elucidate strategies towards improving the accuracy of future TCR structure predictors. Our observations reinforce the importance of paired TCR sequence information and capture the current state-of-the-art for TCR structure prediction, while our model and 1.5M structure predictions enable the use of structural TCR information at an unprecedented scale.
Large Language Models are versatile, general-purpose tools with a wide range of applications. Recently, the advent of "reasoning models" has led to substantial improvements in their abilities in advanced problem-solving domains such as mathematics and software engineering. In this work, we assessed the ability of reasoning models to perform chemistry tasks directly, without any assistance from external tools. We created a novel benchmark, called ChemIQ, consisting of 816 questions assessing core concepts in organic chemistry, focused on molecular comprehension and chemical reasoning. Unlike previous benchmarks, which primarily use multiple choice formats, our approach requires models to construct short-answer responses, more closely reflecting real-world applications. The reasoning models, OpenAI's o3-mini, Google's Gemini 2.5 Pro, and DeepSeek R1, answered 50%-57% of questions correctly in their highest reasoning modes, with higher reasoning levels significantly increasing performance on all tasks. These models substantially outperformed the nonreasoning models which achieved only 3%-7% accuracy. We found that Large Language Models can now convert SMILES strings to IUPAC names, a task earlier models were unable to perform. Additionally, we show that the latest reasoning models can elucidate structures from 1D and 2D 1H and 13C NMR data, with Gemini 2.5 Pro correctly generating SMILES strings for around 90% of molecules containing up to 10 heavy atoms, and in one case solving a structure comprising 25 heavy atoms. For each task, we found evidence that the reasoning process mirrors that of a human chemist. Our results demonstrate that the latest reasoning models are becoming increasingly capable of performing advanced chemical reasoning.