The third Critical Assessment of Computational Hit-finding Experiments (CACHE) challenged computational teams to identify chemically novel ligands targeting the macrodomain 1 of SARS-CoV-2 Nsp3, a promising coronavirus drug target. Twenty-three groups deployed diverse design strategies to collectively select 1739 ligand candidates. While over 85% of the designed molecules were chemically novel, the best experimentally confirmed hits were structurally similar to previously published compounds. Confirming a trend observed in CACHE #1 and #2, two of the best-performing workflows used compounds selected by physics-based computational screening methods to train machine learning models able to rapidly screen large chemical libraries, while four others used exclusively physics-based approaches. Three pharmacophore searches and one fragment growing strategy were also part of the seven winning workflows. While active molecules discovered by CACHE #3 participants largely mimicked the adenine ring of the endogenous substrate, ADP-ribose, preserving the canonical chemotype commonly observed in previously reported Nsp3-Mac1 ligands, they still provide novel structure-activity relationship insights that may inform the development of future antivirals. Collectively, these results show that multiple molecular design strategies can efficiently converge on similar potent molecules.
In many enzymes, movement of domains from open to closed state forms the envi ronment required for catalysis. We have studied ligand-induced domain motion in 82 enzymes by generating ensembles of AlphaFold 3 (AF3) models both with and without the presence of ligands that are known to trigger such motion. It was found that the results heavily depend on the number of apo and holo structures of each enzyme in the Protein Data Bank (PDB). For enzymes with more apo than holo structures, 64.8% of models generated without ligand are closer to the open apo than to the closed holo state. In contrast, for enzymes that have more holo than apo structures in the PDB, 75.5% of AF3 models without any ligand are in the holo conformation, revealing strong mem orization. In both cases, adding the ligand has only a moderate impact. However, the impact of ligand is substantial for proteins that have only a few structures in the training set. Ligands are placed with higher accuracy if there are more holo structures with dif ferent ligands in the PDB. We have found that nonbinder ligands also generate similar domain motion, and the distributions of the predicted enzyme conformations remain close to those obtained with the native trigger ligands, but with lower ligand pLDDT values. For enzymes with more holo than apo structures in the PDB, AlphaFold2 also generates the majority of models close to the holo state, suggesting the same memori zation effects seen for AF3.
AlphaFold-like models have transformed monomer structure prediction yet reliable, generalizable prediction of protein-protein interactions (PPIs) remains challenging, particularly for antibody-antigen docking. These limitations often stem from reliance on pattern memorization under severe data scarcity. Here, we review the emerging transition toward physics-integrated machine learning to address these gaps. We categorize recent efforts to improve generalization into three complementary approaches: (I) enriching inputs with physics-based sampling (e.g., molecular dynamics/fast Fourier transform ensembles); (II) designing architectures with strict geometric inductive biases (e.g., SE(3)-equivariance); and (III) constraining the generative process using physical energy functions or potentials. While standard models often struggle on out-of-distribution targets, these hybrid strategies aim to enforce physical plausibility at different stages of the prediction pipeline, offering a path from memorization to true physical generalization.
Affinity-controlled release provides a versatile approach for the delivery of proteins from hydrogel systems by harnessing noncovalent interactions between a molecule of interest and a binding ligand. We present a strategy for the controlled release of native antibodies by leveraging affinity interactions with peptide ligands specific to the fragment crystallizable (Fc) region. Two Fc-binding ligands (FcLs) were engineered using distinct spacers, yielding different degrees of equilibrium dissociation constants (KD) for the Fc region of human IgG1: 2.54 ± 0.03 × 10-8 M (HWRGWV-GAKSKG; FcL1) and 3.01 ± 0.09 × 10-7 M (HWRGWV-K(PEG); FcLPEG). These ligands were immobilized within a chemically cross-linked hyaluronan-oxime hydrogel, where controlled release of bioactive bevacizumab was observed with FcL1 but not with the lower-affinity FcLPEG. To further explore the versatility of this approach, FcL1 was incorporated into a physically cross-linked hyaluronan-methylcellulose hydrogel, demonstrating tunable release of multiple IgG1 antibodies, including bevacizumab and adalimumab, each over a 7-day period. Together, this work demonstrates a broadly applicable strategy to tune antibody release.
Therapeutic antibodies are a leading class of biologics, yet their unique architecture poses challenges for computational modeling. Each antibody comprises paired heavy and light variable domains with conserved framework regions that maintain structure and hypervariable complementarity-determining regions (CDRs) that directly contact antigens. This functional asymmetry, where CDRs determine binding specificity while frameworks provide scaffolding, suggests that region-aware training strategies could yield superior representations. Existing protein language models treat all regions uniformly, potentially missing critical features present in CDRs. We developed a region-aware pretraining strategy for paired variable-domain sequences using two protein language models: a 3-billion-parameter model (ESM2) and a compact 600-million-parameter model (ESM C). We compared three masking approaches: uniform whole-chain masking, CDR-focused masking, and a hybrid strategy. Final models were trained on over 1.6 million paired antibody sequences and evaluated on binding affinity datasets with over 90,000 antibody variants across six antigens, including single-mutant panels and combinatorial libraries. Here we show that CDR-focused training produces embeddings with superior predictive performance for antibody–antigen binding. Our approach achieves up to 27% improvements in binding affinity prediction compared to benchmarked antibody models. Remarkably, training exclusively on paired sequences proves sufficient; pretraining on billions of unpaired sequences provides no measurable benefit. Our compact model matches or exceeds larger antibody-specific baselines. These findings establish that prioritizing paired sequences with CDR-aware supervision over scale and complex training schemes achieves both computational efficiency and predictive accuracy, providing a practical framework for next-generation antibody language models.
Immunoglobulins (IGs) made by chronic lymphocytic leukemia (CLL) B cells are unique in that they bind themselves (homo-dimerize). This interaction leads to signal transduction with functional consequences that depend on the affinity of homo-dimerization. We have studied the antigen-binding properties of the IGs from a subset of patients with CLL (Subset #4) that homo-dimerize at high affinity. Previously, we had found that subset #4 IGs bound viable lymphocytes. Our new studies, probing an array of >8,000 antigens, indicate that these IGs also bind influenza virus. Because of the IGs high-affinity homo-dimerization, we asked if the defined foreign- and self-antigenic interactions were mediated by conventional B-cell receptor (BCR) domains or a non-conventional receptor created by homo-dimerization. The studies indicated the latter since abrogation of homo-dimerization eliminated binding to influenza virus and its hemagglutinin and to viable lymphocytes. Using these findings, we modeled a developmental path whereby a naive IgM+ B cell with subset #4 heavy and light chain variable domains used the conventional BCR to interact with auto- and foreign antigens and acquire homo-dimerization capacity to create the non-conventional antigen-receptor when transitioning to a leukemic cell. Future studies will determine if this process is an idiosyncratic occurrence or a physiologic principle.
Knowledge of protein-metabolite interactions can enhance mechanistic understanding and chemical probing of biochemical processes, but the discovery of endogenous ligands remains challenging. Here, we combined rapid affinity purification with precision mass spectrometry and high-resolution molecular docking to precisely map the physical associations of 296 chemically diverse small-molecule metabolite ligands with 69 distinct essential enzymes and 45 transcription factors in the gram-negative bacterium Escherichia coli. We then conducted systematic metabolic pathway integration, pan-microbial evolutionary projections, and independent in-depth biophysical characterization experiments to define the functional significance of ligand interfaces. This effort revealed principles governing functional crosstalk on a network level, divergent patterns of binding pocket conservation, and scaffolds for designing selective chemical probes. This structurally resolved ligand interactome mapping pipeline can be scaled to illuminate the native small-molecule networks of complete cells and potentially entire multi-cellular communities.
In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein-protein and protein-ligand complex structures. For protein-protein docking, we combined physics-based sampling-using ClusPro FFT docking and molecular dynamics-with AlphaFold (AF)-based sampling, followed by AF-based refinement. Our method produced numerous high-accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics-based sampling alongside deep learning-based refinement. For protein-ligand docking, we integrated the ClusPro LigTBM template-based approach with a machine learning-based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics-based sampling and a diffusion model. Our template-based strategy achieved a mean lDDT-PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics-based modeling with AI-driven refinement can significantly enhance the accuracy of both protein-protein and protein-ligand structure predictions.
Optimization exercises strive toward increasing the efficacy and selectivity of small molecules toward the target of interest while simultaneously phasing out design elements that lead to off-target interactions. Given the nonequilibrium nature of biological systems, greater reliance should be placed on engineering kinetic selectivity in addition to equilibrium thermodynamic selectivity; however, the rational design of kinetic selectivity is a challenging endeavor. This study presents a systematic knowledge-based approach to the design of inhibitors that vary in their binding kinetics for Bruton's tyrosine kinase (BTK), a target for treating B-cell malignancies and autoimmune diseases. A detailed kinetic assessment was performed on existing BTK inhibitors, which, together with structural studies, provided critical insights into BTK-inhibitor interactions that control the kinetics of enzyme inhibition. Subsequently, a series of pyrazolopyrimidines was designed with the objective of modifying interactions between the inhibitor and the regulatory (R) spine in the kinase back pocket, which were hypothesized to modulate the stability of the transition state on the binding reaction coordinate. This resulted in the development of BTK inhibitors with extended residence time in which the variation in kon and koff was uncoupled from equilibrium thermodynamic affinity.
We have investigated the impact of conformational diversity on the prediction of druggability to see whether structural ensembles of protein targets offer more precise insights. The study is based on binding hot spot analyses performed on 37 binding sites from 33 proteins adapted from well-known druggability benchmark sets. Binding hot spots are regions on proteins that significantly influence binding free energy and hence can be used to predict druggability. Using fast Fourier transform-based algorithms, small organic probe molecules are docked to pinpoint these hot spots with denser probe clusters indicating higher binding affinity. The binding sites are mapped across the structural ensemble of protein crystal structures from the Protein Data Bank with 90% sequence identity. Druggability is analyzed according to the hot spot strength, connectivity, compactness, and maximum dimension. Our results show that a protein's druggability depends on consensus across the structural ensemble, requiring approximately 70% of structures to have a strong hot spot at the binding site of interest and approximately 50% of structures to meet all three druggability criteria. The ability to occasionally access a rare druggable conformation is not sufficient for a protein to be druggable in practice. Hot spot strength proves to be crucial for druggability. However, failing to meet the secondary druggability criteria of connectivity/compactness and maximum dimension for 30% or more structures indicates the inhomogeneity of the ensemble and signals the need for a more detailed analysis of mapping results. Such inhomogeneity may occur due to conformational differences caused by the binding of charged ligands or by the substantial flexibility of the binding site. The hot spots can be determined by the public server FTMove, and the codes for processing the server output for druggability analysis are available on GitHub.
Antibodies are a leading class of biologics, yet their architecture with conserved framework regions and hypervariable complementarity-determining regions (CDRs) poses unique challenges for computational modeling. We present a region-aware pretraining strategy for paired heavy (VH) and light (VL) sequences in variable domains using ESM2-3B and ESM C (600M) protein language models. We compare three masking strategies: whole-chain, CDR-focused, and a hybrid approach. Through evaluation on binding affinity datasets spanning single-mutant panels and combinatorial mutants, we demonstrate that CDR-focused training produces superior embeddings for functional prediction. Notably, training only on VH-VL pairs proves sufficient, eliminating the need for massive unpaired pretraining that provides no measurable downstream benefit. Our compact 600M ESM C model achieves state-of-the-art performance, matching or exceeding larger antibody-specific baselines. These findings establish a principled framework for antibody language models: prioritize paired sequences with CDR-aware supervision over scale and complex training curricula to achieve both computational efficiency and predictive accuracy. ### Competing Interest Statement The authors have declared no competing interest. Merck Research Laboratories
To investigate the molecular basis of homeostatic synaptic plasticity, we adapted a photo-proximity labeling-based functional proteomics workflow to identify protein-protein interactions involving the GluA1 subunit of α-amino-3-hydroxy-5-methyl-4-isoxazolepropionic acid receptor (AMPAR) in live primary rat neurons. Using antibodies conjugated to a photoactivatable flavin-based catalyst, we demonstrated target selective biotinylation and recovery of AMPAR along with both well-described and previously unreported auxiliary proteins associated with neurotransmission. This resulted in the identification of the calcium sensor neuronal calcium sensor 1 (NCS1), which we validated and functionally characterized as a key regulator of homeostatic plasticity initiated via engagement with the calcium-permeable AMPARs.
Targeted protein degradation (TPD) is a rapidly emerging and potentially transformative therapeutic modality. However, the large majority of >600 known ubiquitin ligases have yet to be exploited as TPD effectors by proteolysis-targeting chimeras (PROTACs) or molecular glue degraders (MGDs). We report here a chemical-genetic platform, Site-specific Ligand Incorporation-induced Proximity (SLIP), to identify actionable ("PROTACable") sites on any potential effector protein in intact cells. SLIP uses genetic code expansion to encode copper-free "click" ligation at a specific effector site in intact cells, enabling the in situ formation of a covalent PROTAC-effector conjugate against a target protein of interest. Modification at actionable effector sites drives degradation of the targeted protein, establishing the potential of these sites for TPD. Using SLIP, we systematically screened dozens of sites across E3 ligases and E2 enzymes from diverse classes, identifying multiple novel potentially PROTACable effector sites which are competent for TPD. SLIP adds a powerful approach to the proximity-induced pharmacology (PIP) toolbox, enabling future effector ligand discovery to fully enable TPD and other emerging PIP modalities.
Cancer cells, including the most aggressive brain cancer, glioblastoma, can exhibit a high degree of phenotypic plasticity. However, it is not well understood whether and how distinct phenotypic states might be associated with or driven by specific signaling processes. Here, we use a novel approach to the identification of phenotype-specific signaling networks in cells undergoing the Go-or-Grow switch in populations of GBM cell lines. We find that the transition to invasive spread may be associated with the onset of DNA damage response, cell cycle arrest, and activation of the AKT-mTOR signaling. We reconstructed the large-scale signaling networks, revealing integration of diverse pathways and possible feedback interactions mediating phenotypic stability. We further show that phosphorylation outcomes may be preferentially mediated by 14-3-3 proteins, previously implicated in stress response and control of the cell cycle. Finally, we demonstrate that the phenotype-specific phospho-site signatures have predictive power for disease-free patient survival. This study paves the way for a more comprehensive understanding of phenotypic plasticity in GBM and other aggressive cancers. ### Competing Interest Statement The authors have declared no competing interest.
In fragment-based drug design (FBDD), libraries of low molecular weight compounds are screened against a receptor. Due to their size, fragment hits typically bind with weak affinities by forming a handful of highly efficient interactions with the receptor. Such fragment hits must be expanded into more potent lead compounds in order to achieve higher binding affinities. Approaches for expanding fragments to leads include growing—the iterative expansion of the scaffold, and merging—the linking of two fragment hits. In both cases the design can be facilitated by information on the ligand binding preferences of the target protein. Here we describe a protocol for fragment expansion using E-FTMap, an automated web server that identifies important pharmacophore binding regions within a binding site of proteins using the receptor structure alone. E-FTMap distributes 119 small organic probes across a binding site, identifies energy minima in which similar probes bind, and clusters probes by their atom types to identify regions which preferably bind specific atom types. Unless a priori known, the binding site for this analysis can be identified by our FTMap server that uses only 16 probes to find binding hot spots that are generally preferable for ligand binding, whereas the subsequent use of E-FTMap provides atom-specific information. The utility of E-FTMap as a tool for guiding the expansion of fragments into higher affinity binders is demonstrated by its application to 17 proteins that have been targeted by FBDD. The E-FTMap webserver is publicly accessible at https://eftmap.bu.edu/.
The neural network-based program AlphaFold2 (AF2) provides high accuracy structure prediction for a large fraction of globular proteins. An important question is whether these models are accurate enough for reliably docking small ligands. Several recent papers and the results of CASP15 reveal that local conformational errors reduce the success rates of direct ligand docking. Here, we focus on the ability of the models to conserve the location of binding hot spots, regions on the protein surface that significantly contribute to the binding free energy of the protein-ligand interaction. Clusters of hot spots predict the location and even the druggability of binding sites, and hence are important for computational drug discovery. The hot spots are determined by protein mapping that is based on the distribution of small fragment-sized probes on the protein surface and is less sensitive to local conformation than docking. Mapping models taken from the AlphaFold Protein Structure Database show that identifying binding sites is more reliable than docking, but the success rates are still 5% to 10% lower than based on mapping X-ray structures. The drop in accuracy is particularly large for models of multidomain proteins. However, both the model binding sites and the mapping results can be substantially improved by generating AF2 models for the ligand binding domains of interest rather than the entire proteins and even more if using forced sampling with multiple initial seeds. The mapping of such models tends to reach the accuracy of results obtained by mapping the X-ray structures.
Cryptic sites can expand the space of druggable proteins, but the potential usefulness of such sites needs to be investigated before any major effort. Given that the binding pockets are not formed, the druggability of such sites is not well understood. The analysis of proteins and their ligands shows that cryptic sites that are formed primarily by the motion of side chains moving out of the pocket to enable ligand binding generally do not bind drug-sized molecules with sufficient potency. By contrast, sites that are formed by loop or hinge motion are potentially valuable drug targets. Arguments are provided to explain the underlying causes in terms of classical enzyme inhibition theory and the kinetics of side chain motion and ligand binding.
In computational biology, accurate prediction of phosphopeptide-protein complex structures is essential for understanding cellular functions and advancing drug discovery and personalized medicine. While AlphaFold has significantly improved protein structure prediction, it faces accuracy challenges in predicting structures of complexes involving phosphopeptides possibly due to structural variations introduced by phosphorylation in the peptide component. Our study addresses this limitation by refining AlphaFold to improve its accuracy in modeling these complex structures. We employed weighted metrics for a comprehensive evaluation across various protein families. The enhanced model notably outperforms the original AlphaFold, showing a substantial increase in the weighted average local distance difference test (lDDT) scores for peptides: from 52.74 to 76.51 in the Top 1 model and from 56.32 to 77.91 in the Top 5 model. These advancements not only deepen our understanding of the role of phosphorylation in cellular signaling but also have extensive implications for biological research and the development of innovative therapies.### Competing Interest StatementThe authors have declared no competing interest.
The knowledge of ligand binding hot spots and of the important interactions within such hot spots is crucial for the design of lead compounds in the early stages of structure-based drug discovery. The computational solvent mapping server FTMap can reliably identify binding hot spots as consensus clusters, free energy minima that bind a variety of organic probe molecules. However, in its current implementation, FTMap provides limited information on regions within the hot spots that tend to interact with specific pharmacophoric features of potential ligands. E-FTMap is a new server that expands on the original FTMap protocol. E-FTMap uses 119 organic probes, rather than the 16 in the original FTMap, to exhaustively map binding sites, and identifies pharmacophore features as atomic consensus sites where similar chemical groups bind. We validate E-FTMap against a set of 109 experimentally derived structures of fragment-lead pairs, finding that highly ranked pharmacophore features overlap with the corresponding atoms in both fragments and lead compounds. Additionally, comparisons of mapping results to ensembles of bound ligands reveal that pharmacophores generated with E-FTMap tend to sample highly conserved protein-ligand interactions. E-FTMap is available as a web server at https://eftmap.bu.edu.