Protein-ligand interactions underpin biological regulation and drug action, with both binding affinity and binding kinetics shaping functional outcomes. By analysing kinetic data for 4,311 protein-small-molecule pairs, we find that when association occurs below the diffusion-controlled limit, the rates of ligand association (k on) and dissociation (k off) are primarily determined by how the initial encounter complex reorganizes into the final bound state, and that this reorganization is governed chiefly by the intrinsic dynamic properties of the protein rather than by structural features of the ligand. Counterintuitively, therefore, k off exhibits minimal dependence on ligand structure, so that dissociation proceeds through protein-gated conformational transitions rather than through direct rupture of protein-ligand contacts. This mechanistic behaviour stands in marked contrast to that for protein-protein complexes, based on an analysis of 1,561 interactions. Together, these findings challenge prevailing assumptions regarding the molecular determinants of small-molecule binding kinetics, and have broad implications for rationally modulating protein-ligand interactions and drug-target residence times.
Flavoenzymes, which use a flavin adenine dinucleotide (FAD) cofactor, feature a wide range of activities and substrate specificities. This includes the ability to act as oxidases, relying on molecular oxygen (O2) as a co-substrate, or as dehydrogenases wherein the FAD is redox-cycled not by O2, but by other redox-active proteins or small-molecule substrates. Nicotine oxidoreductase (NicA2) is a flavin-dependent enzyme that provides a useful model system for interrogating structure-function relationships: first characterized as an oxidase, NicA2 has been proven to be a true dehydrogenase, with a native cytochrome c redox partner that greatly stimulates activity. Out of the 9,000 members of the flavin amine oxidase (FAO) superfamily, NicA2 and its downstream homolog pseudooxynicotine amine oxidase (Pnao) are the only members experimentally characterized as using a cytochrome c as an electron acceptor. This finding raises the question of whether other members may function similarly, and thus might in fact be unrecognized dehydrogenases. Additionally, NicA2 is structurally similar to known oxidases, for example Corynebacterium ammoniagenes monoamine oxidase (caMAO) (RMSD value of 1.89), therefore the molecular switch to convert an enzyme from a dehydrogenase to an oxidase may lie in a few key residues. Our bioinformatics analysis identifies dozens of genomes in which a cytochrome c is immediately downstream or upstream of a NicA2 homolog and one case of a larger molecular weight enzyme annotated as both a FAO and cytochrome c. Using the genome neighborhood network, many likely uncharacterized dehydrogenases were found that may have cytochrome c or cytochrome p450 cofactors. Upon further sequence and structural analysis, the dually annotated enzyme is a naturally occurring fusion: a NicA2 homolog fused to a cytochrome c at the enzyme's C-terminus, and its expression is currently in progress. Eight rationally designed variants based on conserved residues in known oxidases have been selected for mutagenesis and steady state kinetics. Additional information about the first half reaction with flavin has been elucidated by x-ray crystallography of a nicotine analog inhibitor complex. The structure, refined to 2.88 Å, shows good correspondence between positioning in the analog complex and the substrate complex structure which validates that the charge-transfer complex and spectral shift revealed in spectroscopic studies are mirroring steps along the reaction coordinate. Overall our study will elucidate how NicA2 and other cytochrome-dependent family members have evolved to interact with a cytochrome instead of O2, paving the way for the use of flavin-dependent enzymes in biosensing and bioremediation.
Machine learning (ML) is rapidly gaining traction in many areas of experimental molecular science for elucidating relationships and patterns in large or complex data sets. Historically, ML was largely the preserve of those with specialized training in fields such as statistics or cheminformatics. Increasingly, however, ML methodologies are becoming part of the standard toolkit for experimental scientists across a range of disciplines. For scientists without a significant background in computer science or statistics, lowering the barrier of entry to these ML techniques is important to broadening access to these powerful methods. Here we provide detailed, step‐by‐step protocols for performing four ML methods that are particularly useful for applications in biochemistry, cell biology, and drug discovery: hierarchical clustering, principal component analysis (PCA), partial least squares discriminant analysis (PLSDA), and partial least squares regression (PLSR). The protocols are written for the widely used software MATLAB, but no prior experience with MATLAB is required to use them. We include an explanation of each step, pitched at a level to be understood by investigators without any prior experience with ML, MATLAB, or any kind of coding. We also highlight the scientific issues pertaining to selecting and scaling the data to be analyzed. Throughout, we emphasize the relationship between the scientific question and how to choose data and methods that will allow it to be addressed in a meaningful way. Our aim is to provide a basic introduction that will equip experimental chemical biologists, chemists, and other biomedical scientists with the knowledge required to use ML to aid in the design of experiments, the formulation and data‐driven testing of hypotheses, and the analysis of experimental data. © 2025 Wiley Periodicals LLC. Basic Protocol 1 : Clustering Basic Protocol 2 : Principal component analysis Basic Protocol 3 : Partial least squares‐discriminant analysis Basic Protocol 4 : Partial least squares regression
We have investigated the impact of conformational diversity on the prediction of druggability to see whether structural ensembles of protein targets offer more precise insights. The study is based on binding hot spot analyses performed on 37 binding sites from 33 proteins adapted from well-known druggability benchmark sets. Binding hot spots are regions on proteins that significantly influence binding free energy and hence can be used to predict druggability. Using fast Fourier transform-based algorithms, small organic probe molecules are docked to pinpoint these hot spots with denser probe clusters indicating higher binding affinity. The binding sites are mapped across the structural ensemble of protein crystal structures from the Protein Data Bank with 90% sequence identity. Druggability is analyzed according to the hot spot strength, connectivity, compactness, and maximum dimension. Our results show that a protein's druggability depends on consensus across the structural ensemble, requiring approximately 70% of structures to have a strong hot spot at the binding site of interest and approximately 50% of structures to meet all three druggability criteria. The ability to occasionally access a rare druggable conformation is not sufficient for a protein to be druggable in practice. Hot spot strength proves to be crucial for druggability. However, failing to meet the secondary druggability criteria of connectivity/compactness and maximum dimension for 30% or more structures indicates the inhomogeneity of the ensemble and signals the need for a more detailed analysis of mapping results. Such inhomogeneity may occur due to conformational differences caused by the binding of charged ligands or by the substantial flexibility of the binding site. The hot spots can be determined by the public server FTMove, and the codes for processing the server output for druggability analysis are available on GitHub.
Protein arginine methyl transferase 5 (PRMT5) plays a global role in cell physiology and is an established therapeutic target in cancer. In approximately 10-15% of human cancers, deletion of the methylthioadenosine phosphorylase (MTAP) gene results in accumulation of methylthioadenosine (MTA), exposing a synthetic lethality and opportunity for precision medicine by selective targeting of PRMT5 in this context. Reported small molecule PRMT5 inhibitors engage either cosubstrate S-adenosyl methionine (SAM) or peptide-substrate pockets through diverse mechanisms. A subset of chemotypes demonstrate uncompetitive engagement with SAM or its inhibitory metabolic precursor, MTA. Although uncompetitive engagement can be evaluated in cell-free systems, no methods exist to directly assess this in cells. Here, we describe the development of a fluorescent probe that acts as a dynamic BRET biosensor of the intracellular SAM/MTA pool that overcomes the current limitations of competitive binding analyses. Using this biosensor, we evaluate a range of diverse PRMT5 inhibitors to mechanistically characterize and quantify uncompetitive target engagement as well as ternary complex formation at PRMT5-SAM and PRMT5-MTA complexes in live cells, enabling direct insights into drug mechanism-of-action and metabolite-dependent responses of inhibitors.
Macrocycles are emerging as a prominent modality in drug discovery, including for conventionally druggable targets for which simpler, acyclic ligands are readily discoverable. Given the additional synthetic challenges associated with macrocyclic chemotypes, we address what benefits macrocycles provided for these highly druggable targets. To do this, we examine the effects of macrocyclization on inhibitors of highly druggable kinase targets. For each example, we isolate closely matched acyclic/macrocyclic compound pairs, allowing us to pinpoint the effects of macrocyclization on binding affinity, selectivity, and ADME properties absent confounding factors. Our findings show that while the impact of macrocyclization on potency is variable, a profound effect on selectivity is common. Macrocyclization can also bring benefits for membrane permeability, efflux ratio, blood-brain barrier penetrance, and metabolic stability. These findings lead us to propose specific circumstances in which a drug discoverer targeting kinases or other conventionally druggable target classes should consider a macrocycle approach.
Enhancing protein thermal stability is important for biomedical and industrial applications as well as in the research laboratory. Here, we describe a simple machine-learning method which identifies amino acid substitutions that contribute to thermal stability based on comparison of the amino acid sequences of homologous proteins derived from bacteria that grow at different temperatures. A key feature of the method is that it compares the sequences based not simply on the amino acid identity, but rather on the structural and physicochemical properties of the side chain. The method accurately identified stabilizing substitutions in three well-studied systems and was validated prospectively by experimentally testing predicted stabilizing substitutions in a polyamine oxidase. In each case, the method outperformed the widely used bioinformatic consensus approach. The method can also provide insight into fundamental aspects of protein structure, for example, by identifying how many sequence positions in a given protein are relevant to temperature adaptation.
The design of PROteolysis-TArgeting Chimeras (PROTACs) requires bringing an E3 ligase into proximity with a target protein to modulate the concentration of the latter through its ubiquitination and degradation. Here, we present a method for generating high-accuracy structural models of E3 ligase-PROTAC-target protein ternary complexes. The method is dependent on two computational innovations: adding a "silent" convolution term to an efficient protein-protein docking program to eliminate protein poses that do not have acceptable linker conformations and clustering models of multiple PROTACs that use the same E3 ligase and target the same protein. Results show that the largest consensus clusters always have high predictive accuracy and that the ensemble of models can be used to predict the dissociation rate and cooperativity of the ternary complex that relate to the degrading activity of the PROTAC. The method is demonstrated by applications to known PROTAC structures and a blind test involving PROTACs against BRAF mutant V600E. The results confirm that PROTACs function by stabilizing a favorable interaction between the E3 ligase and the target protein but do not necessarily exploit the most energetically favorable geometry for interaction between the proteins.
Scaffold proteins help mediate interactions between protein partners, often to optimize intracellular signaling. Herein, we use comparative, biochemical, biophysical, molecular, and cellular approaches to investigate how the scaffold protein NEMO contributes to signaling in the NF-κB pathway. Comparison of NEMO and the related protein optineurin from a variety of evolutionarily distant organisms revealed that a central region of NEMO, called the Intervening Domain (IVD), is conserved between NEMO and optineurin. Previous studies have shown that this central core region of the IVD is required for cytokine-induced activation of IκB kinase (IKK). We show that the analogous region of optineurin can functionally replace the core region of the NEMO IVD. We also show that an intact IVD is required for the formation of disulfide-bonded dimers of NEMO. Moreover, inactivating mutations in this core region abrogate the ability of NEMO to form ubiquitin-induced liquid-liquid phase separation droplets in vitro and signal-induced puncta in vivo. Thermal and chemical denaturation studies of truncated NEMO variants indicate that the IVD, while not intrinsically destabilizing, can reduce the stability of surrounding regions of NEMO, due to the conflicting structural demands imparted on this region by flanking upstream and downstream domains. This conformational strain in the IVD mediates allosteric communication between N- and C-terminal regions of NEMO. Overall, these results support a model in which the IVD of NEMO participates in signal-induced activation of the IKK/NF-κB pathway by acting as a mediator of conformational changes in NEMO.
Macrocyclic compounds (MCs) can have complex conformational properties that affect pharmacologically important behaviors such as membrane permeability. We measured the passive permeability of 3600 diverse nonpeptidic MCs and used machine learning to analyze the results. Incorporating selected properties based on the three-dimensional (3D) conformation gave models that predicted permeability with Q2 = 0.81. A biased spatial distribution of polar versus nonpolar regions was particularly important for good permeability, consistent with a mechanism in which the initial insertion of nonpolar portions of a MC helps facilitate the subsequent membrane entry of more polar parts. We also examined effects on permeability of 800 substructural elements by comparing matched molecular pairs. Some substitutions were invariably beneficial or invariably deleterious to permeability, while the influence of others was highly contextual. Overall, the work provides insights into how the permeability of MCs is influenced by their 3D conformational properties and suggests design hypotheses for achieving macrocycles with high membrane permeability.
Noncovalent complexes of transforming growth factor-beta family growth/differentiation factors with their prodomains are classified as latent or active, depending on whether the complexes can bind their respective receptors. For the antiMullerian hormone (AMH), the hormone-prodomain complex is active, and the prodomain is displaced upon binding to its type II receptor, AMH receptor type-2 (AMHR2), on the cell surface. However, the mechanism by which this displacement occurs is unclear. Here, we used ELISA assays to measure the dependence of prodomain displacement on AMH concentration and analyzed results with respect to the behavior expected for reversible binding in combination with ligand-induced receptor dimerization. We found that, in solution, the prodomain has a high affinity for the growth factor (GF) (Kd = 0.4 pM). Binding of the AMH complex to a single AMHR2 molecule does not affect this Kd and does not induce prodomain displacement, indicating that the receptor binding site in the AMH complex is fully accessible to AMHR2. However, recruitment of a second AMHR2 molecule to bind the ligand bivalently leads to a 1000-fold increase in the Kd for the AMH complex, resulting in rapid release of the prodomain. Displacement occurs only if the AMHR2 is presented on a surface, indicating that prodomain displacement is caused by a conformational change in the GF induced by bivalent binding to AMHR2. In addition, we demonstrate that the bone morphogenetic protein 7 prodomain is displaced from the complex with its GF by a similar process, suggesting that this may represent a general mechanism for receptor-mediated prodomain displacement in this ligand family.
Macrocycles, including macrocyclic peptides, have shown promise for targeting challenging protein-protein interactions (PPIs). One PPI of high interest is between Kelch-like ECH-Associated Protein-1 (KEAP1) and Nuclear Factor (Erythroid-derived 2)-like 2 (Nrf2). Guided by X-ray crystallography, NMR, modeling, and machine learning, we show that the full 20 nM binding affinity of Nrf2 for KEAP1 can be recapitulated in a cyclic 7-mer peptide, c[(D)-β-homoAla-DPETGE]. This compound was identified from the Nrf2-derived linear peptide GDEETGE (KD = 4.3 μM) solely by optimizing the conformation of the cyclic compound, without changing any KEAP1 interacting residue. X-ray crystal structures were determined for each linear and cyclic peptide variant bound to KEAP1. Despite large variations in affinity, no obvious differences in the conformation of the peptide binding residues or in the interactions they made with KEAP1 were observed. However, analysis of the X-ray structures by machine learning showed that locations of strain in the bound ligand could be identified through patterns of subangstrom distortions from the geometry observed for unstrained linear peptides. We show that optimizing the cyclic peptide affinity was driven partly through conformational preorganization associated with a proline substitution at position 78 and with the geometry of the noninteracting residue Asp77 and partly by decreasing strain in the ETGE motif itself. This approach may have utility in dissecting the trade-off between conformational preorganization and strain in other ligand-receptor systems. We also identify a pair of conserved hydrophobic residues flanking the core DxETGE motif which play a conformational role in facilitating the high-affinity binding of Nrf2 to KEAP1.
Oxidative stress can cause extensive damage to DNA, proteins, and lipids. Thus, long-term oxidative stress has been linked to many diseases including cancer and Alzheimer's disease. The highly-charged binding interface between regulatory protein Kelch-like ECH-associated protein 1 (KEAP1) and transcription factor nuclear erythroid factor 2-like 2 (Nrf2) is a common drug target for diseases related to oxidative stress. In the absence of oxidative stress, KEAP1 binds with high affinity to the DxETGE motif of Nrf2. Oxidative stress causes KEAP1 to release Nrf2, allowing it to translocate to the nucleus and upregulate the transcription of antioxidant enzymes. Our current peptide inhibitors of this protein-protein interaction (PPI) are modeled after the DxETGE motif and have exhibited low cell permeability, prompting the bioisosteric replacement of charged residues. We previously identified Arg415 as the KEAP1 residue with the most significant contribution to binding energy at the KEAP1-Nrf2 interface. It has been observed that the R415A mutation abolishes Nrf2 binding, and so a more conservative KEAP1 R415K variant was characterized to better understand the interactions between this residue and residue Glu79 on Nrf2. The variant was characterized using X-ray crystallography, differential scanning fluorimetry, and the binding was assessed using a competitive fluorescence anisotropy binding assay. This mutation caused a 40-fold decrease in binding affinity, which suggested that retaining the charge alone was not sufficient to restore wild-type binding affinity. KEAP1 R415K was successfully crystallized in 8% w/v polyethylene glycol 20,000, 0.2 M imidazole, and 0.01 M nickel (II) chloride hexahydrate, and the crystals diffracted to a resolution of 2.43 Å. Structures of the unliganded R415K Kelch domain were determined using molecular replacement. Optimization of the crystals toward the structure determination of the R415K variant bound to peptide inhibitors is underway. These liganded structures can then be used to assess the binding interactions between KEAP1 Arg415 and Nrf2 Glu79 and guide future peptide inhibitor optimizations to improve their potency and bioavailability.
Macrocyclic compounds (MCs) are of growing interest for inhibition of challenging drug targets. We consider afresh what structural and physicochemical features could be relevant to the bioactivity of this compound class. Using these features, we performed Principal Component Analysis to map oral and non-oral macrocycle drugs and clinical candidates, and also commercially available synthetic MCs, in structure-property space. We find that oral MC drugs occupy defined regions that are distinct from those of the non-oral MC drugs. None of the oral MC regions are effectively sampled by the synthetic MCs. We identify 13 properties that can be used to design synthetic MCs that sample regions overlapping with oral MC drugs. The results advance our understanding of what molecular features are associated with bioactive and orally bioavailable MCs, and illustrate an approach by which synthetic chemists can better evaluate MC designs. We also identify underexplored regions of macrocycle chemical space.
Oxidative stress is involved in a wide range of human diseases such as cancer, diabetes, and Alzheimer’s disease. High levels of reactive oxygen species (ROS) cause oxidative damage to DNA and other molecules vital to cell function, such as membrane lipids and proteins. Kelch‐like ECH‐associated protein 1 (Keap1) is a primary regulator of the body’s natural defense system against oxidative stress. In the absence of oxidative stress, Keap1 forms a homodimer that binds a single molecule of transcription factor nuclear factor‐like 2 (Nrf2), causing Nrf2 ubiquitination and degradation. Oxidative stress causes Keap1 to release Nrf2, allowing it to translocate to the nucleus, bind to Maf protein, and activate the production of antioxidant enzymes by binding to the antioxidant response element (ARE). Blocking the protein‐protein interaction (PPI) between Keap1 and Nrf2 could have great therapeutic value. The antioxidant enzymes upregulated by free Nrf2 provide an efficient, catalytic detoxification of ROS. However, as PPIs frequently take place between large, flat interfaces on the two participating proteins, these interactions can be difficult to block. A previous study by our collaborators at UCSF identified 43 thiol‐containing compounds (monophores) that can be covalently tethered to Keap1 at a cysteine near the protein interface. The goal of this project is to characterize the interaction between Keap1 and the monophores using X‐ray crystallography to obtain structures of the protein‐monophore complexes, or adducts. This structural data will allow us to judge which compounds are suitable for scaffold optimization, a technique in which a larger molecule is built from an initial smaller compound to improve its potency and selectivity. Information about the binding modes of the monophores could also be incorporated into the design of other types of drugs that have proven effective in blocking PPIs, such as macrocycles. To identify adduct crystallization conditions, a DTNB assay, which measures the number of reactive sulfhydryl groups in a protein, was used to screen buffer conditions without dithiothreitol (DTT) present. A high‐throughput version of this assay was developed and utilized, in combination with differential scanning fluorimetry (DSF), to identify stabilizing conditions in which adduct can form. This assay was validated with a monophore known to have a stabilizing effect on Keap1, and a new buffer condition (50 mM EPPS, 50 mM NaCl, pH 8.0) was identified in which adduct crystallization attempts can be made.Support or Funding InformationThis work was supported by the Arnold and Mabel Beckman Foundation.
Binding hot spots are regions of proteins that, due to their potentially high contribution to the binding free energy, have high propensity to bind small molecules. We present benchmark sets for testing computational methods for the identification of binding hot spots with emphasis on fragment-based ligand discovery. Each protein structure in the set binds a fragment, which is extended into larger ligands in other structures without substantial change in its binding mode. Structures of the same proteins without any bound ligand are also collected to form an unbound benchmark. We also discuss a set developed by Astex Pharmaceuticals for the validation of hot and warm spots for fragment binding. The set is based on the assumption that a fragment that occurs in diverse ligands in the same subpocket identifies a binding hot spot. Since this set includes only ligand-bound proteins, we added a set with unbound structures. All four sets were tested using FTMap, a computational analogue of fragment screening experiments to form a baseline for testing other prediction methods, and differences among the sets are discussed.