Bruton's tyrosine kinase (BTK) is a nonreceptor tyrosine kinase clinically validated to impact B-cell development. Molecules designed to target BTK, through either covalent or reversible inhibition, have transformed the treatment of hematopoietic malignancies. Wen, T.; Wang, J.; Shi, Y.; Qian, H.; Liu, P. Inhibitors targeting Bruton's tyrosine kinase in cancers: drug development advances. Leukemia 2021, 35(2), 312-332.10.1038/s41375-020-01072-6. These advancements are paving the way for new therapeutics to treat nononcology indications, De Bondt, M.; Renders, J.; Struyf, S.; Hellings, N. Inhibitors of Bruton's tyrosine kinase as emerging therapeutic strategy in autoimmune diseases. Autoimmun. Rev. 2024, 23(5), 103532.10.1016/j.autrev.2024.103532 such as multiple sclerosis (MS), and provide benefits to patients with progressive disease. In this context, we describe the discovery of a highly selective, CNS-penetrant, reversible BTK inhibitor designed to sequester Tyr-551, the critical phosphorylation site, into an inactive conformation, thereby blocking B-cell receptor (BCR) signaling. While this class of molecules demonstrated excellent safety when administered at doses that fully inhibited B-cell activity in the periphery, increasing exposure to achieve similar efficacy in the CNS led to adverse findings. This raises the question of whether it was a molecule-specific off-target toxicity or a consequence of blocking BTK function in microglia.
Molecular similarity plays a central role in ligand-based drug discovery, such as virtual screening, analog searching, and goal-directed molecular generation. However, traditional similarity measures, ranging from fingerprint-based Tanimoto coefficients to 3D shape overlays, are often computationally expensive at scale or rely on hand-crafted molecular descriptors. Meanwhile, many deep learning approaches to similarity-aware design still depend on similarity-specific supervision or costly data curation, limiting their generality across targets. In this work, we propose pretrained embedding distance (PED) as an effective alternative, computed directly from pretrained molecular models without task-specific training. Experimental results show that PED exhibits distinct correlations with traditional similarity metrics, and performs effectively in both ranking molecules for virtual screening and guiding molecular generation via reward design. These findings suggest that pretrained molecular embeddings capture rich structural information and can serve as a promising and scalable similarity measurement for modern AI-aided drug discovery.
Goal-directed optimization is essential for steering molecular generators to propose candidates with desired properties. However, it is often implemented with policy-gradient reinforcement learning, which requires a generation-trajectory log-probability whose form depends on the model architecture and generation procedure. This makes an optimizer difficult to reuse across architectures and conditional generative designs. Supervised fine-tuning needs none of that machinery, but its update is driven by a fixed dataset, so the reward never enters the update. We introduce Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set. Ablations show that reward information is passed primarily through elite selection, rather than through continuous weighting within the selected set. Because the update consumes only scored molecules and the model's native loss, the same rule applies across autoregressive, masked-diffusion, and discrete-flow generators, and across de novo, motif-extension, and linker-design tasks. Under a fixed budget of 3D shape alignment oracle calls on two kinase reference compounds, EW-SFT consistently outperforms the corresponding native optimizers. It further improves goal-directed optimization under a 2D similarity oracle on four held-out references and achieves comparable performance on a sample-efficiency benchmark without a trajectory-level RL formulation. These results demonstrate that EW-SFT is a unified and effective optimizer across molecular generators, design constraints, references, and oracles.
Abstract Boltz-2 has enabled accurate binding affinity prediction by leveraging co-folded protein–ligand structures, but the absence of a public training recipe has limited its use in active drug discovery projects, where new experimental measurements and congeneric ligand series continually arrive during lead optimization. We present an open framework for affinity fine-tuning of Boltz-2, showing that adapting only the affinity prediction components with project-specific experimental data can make the model substantially more useful for lead optimization. We evaluate the approach in two internal studies: a multi-target retrospective benchmark against physics-based and machine learning baselines, and a large single-target study with up to 1,700 ligands. In both settings, fine-tuned Boltz-2 improves correlation over the off-the-shelf model, and in some cases reaches performance competitive with free energy perturbation (FEP) methods. By releasing this framework, we aim to enable the community to adapt Boltz-2 to the specific targets and data of their own drug discovery campaigns. The code is available at: https://github.com/molecularinformatics/Boltz2_affinity
We report the discovery of a chemical series that enhances ApoE secretion from human astrocytes through mechanisms independent of LXR agonism. Target deconvolution of hits from a phenotypic screen in astrocytoma cells employed chemoproteomics, photoaffinity probes, in vitro KINOMEscan analysis, and targeted siRNA knockdown experiments. Photoaffinity labeling coupled with quantitative chemical proteomics identified aryl hydrocarbon receptor (AhR), a transcription factor not previously associated with ApoE secretion, as the primary target. A diverse panel of AhR agonists and antagonists together with genetic knockdown confirmed that ApoE secretion increases when AhR activity is reduced. Using a luciferase reporter assay, we demonstrated that active series analogs exhibit AhR antagonism while inactive compounds do not. Since deletion of AhR has severe peripheral effects, chronic inhibition of AhR is not an attractive therapeutic approach for Alzheimer’s disease; nevertheless, these results position AhR as a modulator of ApoE secretion and a biological pathway worth exploring.
Covalent BTK-inhibitor drugs often contain reactive acrylamide warheads designed to irreversibly bind to their protein targets at free thiol cysteines in the kinase active site. This reactivity also makes covalent inhibitors susceptible to conjugation to endogenous tripeptide glutathione (GSH), leading to clearance. During lead optimization efforts for the drug discovery of covalent BTK inhibitor BIIB129, some expected GSH adducts resulted in an unexpected and highly abundant rearrangement fragment ion in LC-MS/MS. By examining more than 30 inhibitors, the rearrangements were found to be dependent on the presence of a cycloalkane linker that connects the warhead to the kinase hinge binder motif of drug molecules. The proposed mechanism includes the formation of a 16-membered macrocyclic intermediate between the γ-glutamic acid residue (Glu) of GSH and a methyl-cyclobutyl cation, resulting in a rearrangement fragment originating from two distant parts of the adduct molecule separated by the warhead conjugated with the cysteine residue in between. Rich sets of chemical analogues available during the lead optimization enabled confirmation of the macrocyclic rearrangement. Proposed macrocyclic rearrangement was verified using GSH derivatives: N-acetylation of the γ-Glu blocked the rearrangement, and esterification of the γ-Glu side chain resulted in an expected shift in the mass of rearranged fragment ion. Proposed rearranged ion structures were supported by MS3 and MS4 fragmentations. Comparisons of the ion fragmentation of GSH conjugates between cis and trans matched pairs suggest a concerted mechanism for the cyclobutane linker and a stepwise mechanism for the methylcyclobutane linker, respectively.
Molecular alignment and 3D similarity are crucial tasks in computational drug discovery, enabling applications such as virtual screening and pharmacophore modeling. ROSHAMBO, an open-source package for optimizing molecular alignment using Gaussian volume overlaps, demonstrated near-state-of-the-art performance and accuracy across multiple target classes. However, its computational efficiency has been a limiting factor in the virtual screening of ultralarge chemical libraries. To address this limitation, we introduce ROSHMABO2, an optimized version that achieves a greater than 200-fold improvement in performance over the original ROSHAMBO implementation through algorithmic innovations, GPU acceleration, and optimized memory handling. This performance establishes ROSHMABO2 as an ideal tool for high-throughput applications, such as virtual screening and chemical library design, enabling efficient exploration of large chemical spaces. In addition to its computational enhancements, the new version retains its modularity, accessibility, and compatibility with diverse workflows. These improvements position ROSHAMBO2 as a transformative tool for modern cheminformatics, addressing the growing demands for scalable molecular modeling. ROSHAMBO2 is accessible at https://github.com/molecularinformatics/roshambo2 and is available for use under the MIT license.
Multiple sclerosis (MS) is a chronic disease with an underlying pathology characterized by inflammation-driven neuronal loss, axonal injury, and demyelination. Bruton's tyrosine kinase (BTK), a nonreceptor tyrosine kinase and member of the TEC family of kinases, is involved in the regulation, migration, and functional activation of B cells and myeloid cells in the periphery and the central nervous system (CNS), cell types which are deemed central to the pathology contributing to disease progression in MS patients. Herein, we describe the discovery of BIIB129 (25), a structurally distinct and brain-penetrant targeted covalent inhibitor (TCI) of BTK with an unprecedented binding mode responsible for its high kinome selectivity. BIIB129 (25) demonstrated efficacy in disease-relevant preclinical in vivo models of B cell proliferation in the CNS, exhibits a favorable safety profile suitable for clinical development as an immunomodulating therapy for MS, and has a low projected total human daily dose.
In recent years, reinforcement learning (RL) has emerged as a valuable tool in drug design, offering the potential to propose and optimize molecules with desired properties. However, striking a balance between capabilities, flexibility, reliability, and efficiency remains challenging due to the complexity of advanced RL algorithms and the significant reliance on specialized code. In this work, we introduce ACEGEN, a comprehensive and streamlined toolkit tailored for generative drug design, built using TorchRL, a modern RL library that offers thoroughly tested reusable components. We validate ACEGEN by benchmarking against other published generative modeling algorithms and show comparable or improved performance. We also show examples of ACEGEN applied in multiple drug discovery case studies. ACEGEN is accessible at https://github.com/acellera/acegen-open and available for use under the MIT license.
Efficient virtual screening techniques are critical in drug discovery for identifying potential drug candidates. We present an open-source package for molecular alignment and 3D similarity calculations optimized for large-scale virtual screening of small molecules. This work parallels widely used proprietary tools and offers an approach complementary to structure-based virtual screening. Our package employs the PAPER software for optimizing molecular alignments based on Gaussian volume overlaps. GPU acceleration is utilized to significantly reduce computational time and resource requirements. After obtaining the optimal alignments between the target and the query molecules, both shape and color (based on pharmacophore features) scores are computed to assess molecular similarity, with aligned molecules optionally being output in sdf format. The package was benchmarked using the DUDE-Z public data sets. Results demonstrated the package's near-state-of-the-art performance and robustness across multiple target classes, with speed that enables many routine ligand-based drug discovery workflows. As an open-source and freely available resource (github.com/molecularinformatics/roshambo) with both a convenient Python API and command line interface, our package also addresses the need for accessible and efficient virtual screening tools in drug discovery.
The rank ordering of ligands remains one of the most attractive challenges in drug discovery. While physics-based in silico binding affinity methods dominate the field, they still have problems, which largely revolve around forcefield accuracy and sampling. Recent advances in machine learning have gained traction for protein–ligand binding affinity predictions in early drug discovery programs. In this article, we perform retrospective binding free energy evaluations for 172 compounds from our internal collection spread over four different protein targets and five congeneric ligand series. We compared multiple state-of-the-art free energy methods ranging from physics-based methods with different levels of complexity and conformational sampling to state-of-the-art machine-learning-based methods that were available to us. Overall, we found that physics-based methods behaved particularly well when the ligand perturbations were made in the solvation region, and they did not perform as well when accounting for large conformational changes in protein active sites. On the other end, machine-learning-based methods offer a good cost-effective alternative for binding free energy calculations, but the accuracy of their predictions is highly dependent on the experimental data available for training the model.
Virtual screening of large compound libraries to identify potential hit candidates is one of the earliest steps in drug discovery. As the size of commercially available compound collections grows exponentially to the scale of billions, active learning and Bayesian optimization have recently been proven as effective methods of narrowing down the search space. An essential component of those methods is a surrogate machine learning model that predicts the desired properties of compounds. An accurate model can achieve high sample efficiency by finding hits with only a fraction of the entire library being virtually screened. In this study, we examined the performance of a pretrained transformer-based language model and graph neural network in a Bayesian optimization active learning framework. The best pretrained model identifies 58.97% of the top-50,000 compounds after screening only 0.6% of an ultralarge library containing 99.5 million compounds, improving 8% over the previous state-of-the-art baseline. Through extensive benchmarks, we show that the superior performance of pretrained models persists in both structure-based and ligand-based drug discovery. Pretrained models can serve as a boost to the accuracy and sample efficiency of active learning-based virtual screening.
Phosphorothioates (PS) have proven their effectiveness in the area of therapeutic oligonucleotides with applications spanning from cancer treatment to neurodegenerative disorders. Initially, PS substitution was introduced for the antisense oligonucleotides (PS ASOs) because it confers an increased nuclease resistance meanwhile ameliorates cellular uptake and in-vivo bioavailability. Thus, PS oligonucleotides have been elevated to a fundamental asset in the realm of gene silencing therapeutic methodologies. But, despite their wide use, little is known on the possibly different structural changes PS-substitutions may provoke in DNA·RNA hybrids. Additionally, scarce information and significant controversy exists on the role of phosphorothioate chirality in modulating PS properties. Here, through comprehensive computational investigations and experimental measurements, we shed light on the impact of PS chirality in DNA-based antisense oligonucleotides; how the different phosphorothioate diastereomers impact DNA topology, stability and flexibility to ultimately disclose pro-Sp S and pro-Rp S roles at the catalytic core of DNA Exonuclease and Human Ribonuclease H; two major obstacles in ASOs-based therapies. Altogether, our results provide full-atom and mechanistic insights on the structural aberrations PS-substitutions provoke and explain the origin of nuclease resistance PS-linkages confer to DNA·RNA hybrids; crucial information to improve current ASOs-based therapies.
Deep generative models applied to the generation of novel compounds in small-molecule drug design have attracted a lot of attention in recent years. To design compounds that interact with specific target proteins, we propose a Generative Pre-Trained Transformer (GPT)-inspired model for de novo target-specific molecular design. By implementing different keys and values for the multi-head attention conditional on a specified target, the proposed method can generate drug-like compounds both with and without a specific target. The results show that our approach (cMolGPT) is capable of generating SMILES strings that correspond to both drug-like and active compounds. Moreover, the compounds generated from the conditional model closely match the chemical space of real target-specific molecules and cover a significant portion of novel compounds. Thus, the proposed Conditional Generative Pre-Trained Transformer (cMolGPT) is a valuable tool for de novo molecule design and has the potential to accelerate the molecular optimization cycle time.
Absorption, distribution, metabolism, and excretion (ADME), which collectively define the concentration profile of a drug at the site of action, are of critical importance to the success of a drug candidate. Recent advances in machine learning algorithms and the availability of larger proprietary as well as public ADME data sets have generated renewed interest within the academic and pharmaceutical science communities in predicting pharmacokinetic and physicochemical endpoints in early drug discovery. In this study, we collected 120 internal prospective data sets over 20 months across six ADME in vitro endpoints: human and rat liver microsomal stability, MDR1-MDCK efflux ratio, solubility, and human and rat plasma protein binding. A variety of machine learning algorithms in combination with different molecular representations were evaluated. Our results suggest that gradient boosting decision tree and deep learning models consistently outperformed random forest over time. We also observed better performance when models were retrained on a fixed schedule, and the more frequent retraining generally resulted in increased accuracy, while hyperparameters tuning only improved the prospective predictions marginally.
In the scope of drug discovery, the molecular design aims to identify novel compounds from the chemical space where the potential drug-like molecules are estimated to be in the order of 10^60 - 10^100. Since this search task is computationally intractable due to the unbounded search space, deep learning draws a lot of attention as a new way of generating unseen molecules. As we seek compounds with specific target proteins, we propose a Transformer-based deep model for de novo target-specific molecular design. The proposed method is capable of generating both drug-like compounds (without specified targets) and target-specific compounds. The latter are generated by enforcing different keys and values of the multi-head attention for each target. In this way, we allow the generation of SMILES strings to be conditional on the specified target. Experimental results demonstrate that our method is capable of generating both valid drug-like compounds and target-specific compounds. Moreover, the sampled compounds from conditional model largely occupy the real target-specific molecules' chemical space and also cover a significant fraction of novel compounds.
The availability of large chemical libraries containing hundreds of millions to billions of diverse drug-like molecules combined with an almost unlimited amount of compute power to achieve scientific calculations has led investors and researchers to have a renewed interest in virtual screening (VS) methods to identify biologically active compounds. The number of in silico screening tools and software which employ the knowledge of the protein target or known bioactive ligands is increasing at a rapid pace, creating a crowded computational landscape where it has become difficult to assess the real advantages and disadvantages in terms of accuracy and efficiency of each individual VS technology. In the current work, we evaluate the performance of several state-of-the-art commercial software for 3D ligand-based VS against well-known 2D methods using an internally curated benchmarking data set. Our results show that the best individual methods can differ significantly based on the data set, and that combining them using data fusion techniques results in improved enrichment in the top 1 % of retrieved hits. Although 2D methods alone can already provide a significant enrichment in the number of predicted active compounds, the combination of data-fused 2D results with just one out of the best 3D methods (ROCS, FLAP or Blaze) further improves early enrichment and the likelihood of identifying additional chemotypes.
PFRED a software application for the design, analysis, and visualization of antisense oligonucleotides and siRNA is described. The software provides an intuitive user-interface for scientists to design a library of siRNA or antisense oligonucleotides that target a specific gene of interest. Moreover, the tool facilitates the incorporation of various design criteria that have been shown to be important for stability and potency. PFRED has been made available as an open-source project so the code can be easily modified to address the future needs of the oligonucleotide research community. A compiled version is available for downloading at https://github.com/pfred/pfred-gui/releases/tag/v1.0 as a java Jar file. The source code and the links for downloading the precompiled version can be found at https://github.com/pfred .
Apolipoprotein E (apoE) is a 34 kDa key lipid transport protein in both the brain and the periphery. In the brain, it is produced predominantly by astrocytes and has been shown to play an important role in promoting neurite outgrowth and synaptogenesis, clearing amyloid β42, promoting cerebrovascular integrity and in neuroimmune modulation. Given the critical role of apoE in the brain, it is important to understand its regulation in the CNS. To that end, utilizing a phenotypic screen, we identified novel chemical matter that increased astrocytic apoE secretion in vitro. We designed a clickable photoaffinity probe and carried out quantitative chemical proteomics to identify liver x receptor β (LXRβ) as the target. Binding of the ligand stabilized LXRβ, as shown by CETSA. Our findings demonstrated that the lead chemical matter bound directly to LXRβ, and highlight the power of chemical proteomics to identify the target of a phenotypic screening hit.