Generative structure-based drug design (SBDD) models have shown great promise to accelerate our ability to discover novel drug candidates. However, these models have been criticized for producing compounds that are not very synthesizable, and therefore not practically applicable to drug design. In this work, we propose a way to circumvent the synthesizability issue by introducing a model-guided virtual screening (MGVS) pipeline which pairs SBDD models with efficient chemical similarity search methods to identify synthesizable analogs of generated compounds in existing ultra-large compound databases. Using this approach, we demonstrate that synthesizable analogs of generated compounds with equivalent or better docking scores and similar predicted binding poses can be reliably identified across a wide range of protein targets. We find that MGVS outperforms standard virtual ligand screening (VLS), consistently yielding at least a 25x improvement in screening efficiency across three different SBDD models. As drug-like chemical spaces continue to grow and standard VLS methods focused on exhaustive screening become increasingly impractical, approaches like MGVS that effectively narrow the search space will become critical for advancing drug discovery.
Generative structure-based drug design (SBDD) models have shown great promise for accelerating our ability to discover novel drug candidates. However, these models have been criticized for producing compounds that are not very synthesizable and therefore not practically applicable to drug design. In this work, we propose a way to circumvent the synthesizability issue by introducing a model-guided virtual screening (MGVS) pipeline that pairs SBDD models with efficient chemical similarity search methods to identify synthesizable analogues of generated compounds in existing ultralarge compound databases. Using this approach, we demonstrate that synthesizable analogues of generated compounds with equivalent or better docking scores and similar predicted binding poses can be reliably identified across a wide range of protein targets. We find that MGVS outperforms standard virtual ligand screening (VLS), consistently yielding at least a 25x improvement in docking-based screening efficiency across three different SBDD models. As drug-like chemical spaces continue to grow and standard VLS methods focused on exhaustive screening become increasingly impractical, approaches such as MGVS that effectively narrow the search space will become critical for advancing drug discovery.
Homeodomain transcription factors (TFs) recognize their DNA targets through both sequence-specific base contacts and readout of local DNA shape. Although intrinsic DNA structure is encoded by nucleotide sequence, it also undergoes protein-induced structural deformation upon binding. Yet, the interplay between intrinsic and protein-induced DNA shape remains unclear. Here, we dissect how these two readout modes determine binding specificity in a trimeric complex composed of the Drosophila Hox TF Sex combs reduced and its cofactors, Homothorax and Extradenticle. Guided by SELEX-seq data, we performed molecular dynamics simulations of this complex bound to sequences of varying binding affinities. We find that minor groove width reflects intrinsic DNA structure, whereas minor groove width fluctuations capture protein-induced stabilization and reshaping of the DNA. Within the trimeric complex, Homothorax reduces conformational fluctuations in an orientation- and sequence-dependent manner, with charged residues in its N-terminal arm playing key roles in DNA shape readout. This demonstrates that recognition involves a context-dependent balance between conformational selection and induced fit. We extend this analysis to two other homeodomain TFs, Distal-less and Engrailed, revealing that even closely related proteins produce distinct DNA shape signatures depending on sequence context. To translate these insights to novel sequences or mutant proteins, we evaluate whether AlphaFold 3 (AF3) can capture mutation-sensitive DNA shape readout. Although AF3 accurately reproduces wild-type structures, it struggles to predict how mutations or conformational dynamics alter DNA shape. To bridge this gap, we developed a hybrid pipeline integrating AlphaFold-based homology modeling, molecular dynamics simulations, and DeepPBS, a deep learning method for binding specificity prediction. This multiscale framework successfully captures the active modulation of DNA structure missed by AF3 alone, providing new insights into TF binding specificity and creating a roadmap for integrating deep learning and physics-based methods to study molecular mechanisms.
Multiple layers of molecular determinants and mechanisms affect binding specificity between transcription factors (TFs) and DNA. DNA sequence-based deep learning models using convolutional neural networks (CNNs) and self-attention (SA) transformers have improved modeling accuracy and advanced our understanding of TF-DNA binding specificity through network interpretation. However, the systematic evaluation of various strategies for handling DNA sequence orientations in deep learning models-and their interpretation-remains underexplored, especially in the context of learning low-affinity binding site specificity. Using SELEX-seq data for eight Exd-Hox heterodimers in Drosophila, we compared canonical models with data augmentation and reverse-complement weight-sharing models. We found that reverse-complement weight-sharing CNN models and SA models trained with augmented data with reverse complements outperformed other approaches in modeling binding specificity. In this work, we evaluated several interpretation methods, including Gradient*input, DeconvNet, DeepLIFT, and in silico mutagenesis (ISM). Compared to other interpretation methods, ISM was less sensitive to model hyperparameter settings. In this work, we identified Exd-Ubx binding at low-affinity sites and suggested possible biophysical mechanisms. The findings of this study will be relevant for studying the functional role of low-affinity TF binding in gene regulatory mechanisms with possible implications on TF-DNA binding specificity guided protein design.
Fine-tuning of gene expression by transcription factors (TFs) is essential for normal development, but how TF binding is converted into precise expression levels is poorly understood. We used a Drosophila enhancer responding to the TF Distal-less (Dll) to investigate this relationship. By combining computational structure analysis and quantifying relative TFBS affinity and enhancer activity, we describe how TF concentration is converted into precise expression across the developing Drosophila wing. We show that, although the regulatory function of the enhancer generally follows the classical Hill equation, several additional terms are needed to describe the effects of affinity, orientation, sequence, and wing region on the maximum achievable expression and position of transcriptional activity in a tissue. We also found that Dll relies on two distinct sets of amino acid-DNA contacts for binding to higher affinity sites, but contacts are more evenly distributed at lower affinities. Our work integrates information at molecular and tissue scales to describe how the modulation of a single TFBS determines the regulatory function of an enhancer in a complex in vivo environment.
MOTIVATION:DNA sequence and shape readout represent different modes of protein-DNA recognition. Current tools lack the functionality to simultaneously consider alterations in different readout modes caused by sequence mutations. DNAdesign is a web-based tool to compare and design mutations based on both DNA sequence and shape characteristics. Users input a wild-type sequence, select sites to introduce mutations and choose a set of DNA shape parameters for mutation design. RESULTS:DNAdesign utilizes Deep DNAshape to provide ultra-fast predictions of DNA shape based on extended k-mers and offers multiple encoding methods for nucleotide sequences, including the physicochemical encoding of DNA through their functional groups in the major and minor groove. DNAdesign provides all mutation candidates along the sequence and shape dimensions, with interactive visualization comparing each candidate with the wild-type DNA molecule. DNAdesign provides an approach to studying gene regulation and applications in synthetic biology, such as the design of synthetic enhancers and transcription factor binding sites. AVAILABILITY AND IMPLEMENTATION:The DNAdesign webserver and documentation are freely accessible at https://dnadesign.usc.edu.
Interaction between transcription factors (TFs) and DNA plays a key role in regulating gene expression. It is generally believed that these interactions are controlled through recognition of DNA core motifs by TFs. Nevertheless, several studies pointed out the limitation of this view, in particular, DNA sequence variants influencing TF binding are often located outside of core motifs. One possible explanation is that the physical properties of DNA may play a role in TF-DNA interactions. Recent studies have supported the importance of DNA shape features, especially in flanking regions of core motifs. Another important physical property of DNA is DNA breathing, the spontaneous opening of double-stranded DNA through thermal motions. But there have been few genomic studies of the role of DNA breathing in TF-DNA interactions. In this work, we analyzed in vitro TF-DNA binding data of three TFs and found that DNA breathing features inside or near core motifs are correlated with binding affinity. This suggests that these TFs may prefer locally and temporally melted DNA formed through breathing. We extended the analysis to 44 TFs with in vivo ChIP-seq binding data. We found that for a large proportion of TFs, their breathing features in or near core motifs are associated with binding, but the sign and magnitude of these associations vary substantially across TF families. Altogether, our study supports the hypothesis that DNA breathing features near binding motifs contribute to TF-DNA interactions.
Forkhead homologue 1 (Fkh1) is a yeast transcription factor that plays essential roles in cell-cycle dynamics. Here, we report the co-crystal structure of the DNA-binding domain (DBD) of the yeast Fkh1 protein in complex with a 19-base pair oligonucleotide containing the core binding site and flanking regions. The three-dimensional structure of the Fkh1-DBD reveals a previously unknown protein fold among all known Forkhead proteins. The winged-helix fold forms base-specific contacts of α-helix H3 with the major groove of the core binding site. Wing 1 and Wing 2 form DNA shape-mediated contacts with the minor groove of the binding site flanking regions. The conformation of Wing 2 is distinct from all known Forkhead proteins, with α-helices H5 and H6 wrapping back onto the protein core, creating a stable Wing 2 loop. Backbone interactions with β-strands S1 and S2 reveal a structural mechanism for previously observed flanking region preferences in SELEX-seq experiments. In vivo yeast experiments on Fkh1 mutants demonstrate that wing residues interacting with flanking regions are important for Fkh1 function. Molecular dynamics simulations relate Fkh1 function to conformational flexibility of wing residues. The novel Forkhead fold enables Fkh1 function with implications, such as structure-based protein design, for other DNA-binding proteins.
Cas12a is a class 2 type V CRISPR-associated nuclease that uses an effector complex comprised of a single protein activated by a CRISPR-encoded small RNA to cleave double-stranded DNA at specific sites. Cas12a processes unique features as compared to other CRISPR effector nucleases such as Cas9, and has been demonstrated as an effective tool for manipulating complex genomes. Prior studies have indicated that DNA flexibility at the region adjacent to the protospacer-adjacent-motif (PAM) contributes to Cas12a target recognition. Here, we adapted a SELEX-seq approach to further examine the connection between PAM-adjacent DNA flexibility and off-target binding by Cas12a. A DNA library containing DNA-DNA mismatches at PAM + 1 to + 6 positions was generated and subjected to binding in vitro with FnCas12a in the absence of pairing between the RNA guide and DNA target. The bound and unbound populations were sequenced to determine the propensity for off-target binding for each of the individual sequences. Analyzing the position and nucleotide dependency of the DNA-DNA mismatches showed that PAM-dependent Cas12a off-target binding requires unpairing of the protospacer at PAM + 1 and increases with unpairing at PAM + 2 and + 3. This revealed that PAM-adjacent DNA flexibility can tune Cas12a off-target binding. The work adds support to the notion that physical properties of the DNA modulate Cas12a target discrimination, and has implications on Cas12a-based applications.
We present RNAproDB (https://rnaprodb.usc.edu/), a new webserver, analysis pipeline, database, and highly interactive visualization tool, designed for protein-RNA complexes, and applicable to all forms of nucleic acid containing structures. RNAproDB computes several mapping schemes to place nucleic acid components and present protein-RNA interactions appropriately. Various structural annotations are computed including non-canonical base-pairing geometries, hydrogen bonds, and protein-RNA and RNA-RNA water-mediated interactions. This information is presented through integrated visualization and data tools. Subgraph selection facilitates studying smaller components of the interface. Molecular surface electrostatic potential can be visualized. RNAproDB enables analyzing and exploring experimentally determined, predicted, and designed protein-nucleic acid complexes. We present a quantitative analysis of pre-analyzed protein-RNA structures in RNAproDB revealing statistical patterns of molecular binding and recognition.
In recent years, computational methods and artificial intelligence approaches have proven uniquely suited for studying patterns in molecular biology. In this focus issue, we spoke with researchers about using these tools to address various biological questions and explore both current implications and future possibilities.
Predicting specificity in protein-DNA interactions is a challenging yet essential task for understanding gene regulation. Here, we present Deep Predictor of Binding Specificity (DeepPBS), a geometric deep-learning model designed to predict binding specificity across protein families based on protein-DNA structures. The DeepPBS architecture allows investigation of different family-specific recognition patterns. DeepPBS can be applied to predicted structures, and can aid in the modeling of protein-DNA complexes. DeepPBS is interpretable and can be used to calculate protein heavy atom-level importance scores, demonstrated as a case-study on p53-DNA interface. When aggregated at the protein residue level, these scores conform well with alanine scanning mutagenesis experimental data. The inference time for DeepPBS is sufficiently fast for analyzing simulation trajectories, as demonstrated on a molecular-dynamics simulation of a Drosophila Hox-DNA tertiary complex with its cofactor. DeepPBS and its corresponding data resources offer a foundation for machine-aided protein-DNA interaction studies, guiding experimental choices and complex design, as well as advancing our understanding of molecular interactions.
Recently, the remarkable growth of available crystal structure data and libraries of commercially available or readily synthesizable molecules have unlocked previously inaccessible regions of chemical space for drug development. Paired with improvements in virtual ligand screening methods, these expanded libraries are having a notable impact on early drug design efforts. Yet screening-based methods still face scalability limits, due to computational constraints and the sheer scale of drug-like space. Machine learning approaches are overcoming these limitations by learning the fundamental intra- and intermolecular relationships in drug-target systems from existing data. Here, we introduce DrugHIVE, a deep hierarchical variational autoencoder that outperforms state-of-the-art autoregressive and diffusion-based methods in both speed and performance on common generative benchmarks. DrugHIVE's hierarchical design enables improved control over molecular generation. Its capabilities include dramatically increasing virtual screening efficiency and accelerating a wide range of common drug design tasks, including de novo generation, molecular optimization, scaffold hopping, linker design, and high-throughput pattern replacement. Our highly scalable method can even be applied to receptors with high-confidence AlphaFold-predicted structures, extending the ability to generate high-quality drug-like molecules to a majority of the unsolved human proteome.
Sequence-dependent DNA shape plays an important role in understanding protein-DNA binding mechanisms. High-throughput prediction of DNA shape features has become a valuable tool in the field of protein-DNA recognition, transcription factor-DNA binding specificity, and gene regulation. However, our widely used webserver, DNAshape, relies on statistically summarized pentamer query tables to query DNA shape features. These query tables do not consider flanking regions longer than two base pairs, and acquiring a query table for hexamers or higher-order k-mers is currently still unrealistic due to limitations in achieving sufficient statistical coverage in molecular simulations or structural biology experiments. A recent deep-learning method, Deep DNAshape, can predict DNA shape features at the core of a DNA fragment considering flanking regions of up to seven base pairs, trained on limited simulation data. However, Deep DNAshape is rather complicated to install, and it must run locally compared to the pentamer-based DNAshape webserver, creating a barrier for users. Here, we present the Deep DNAshape webserver, which has the benefits of both methods while being accurate, fast, and accessible to all users. Additional improvements of the webserver include the detection of user input in real time, the ability of interactive visualization tools and different modes of analyses. URL: https://deepdnashape.usc.edu.
Circadian clock genes are emerging targets in many types of cancer, but their mechanistic contributions to tumor progression are still largely unknown. This makes it challenging to stratify patient populations and develop corresponding treatments. In this work, we show that in breast cancer, the disrupted expression of circadian genes has the potential to serve as biomarkers. We also show that the master circadian transcription factors (TFs) BMAL1 and CLOCK are required for the proliferation of metastatic mesenchymal stem-like (mMSL) triple-negative breast cancer (TNBC) cells. Using currently available small molecule modulators, we found that a stabilizer of cryptochrome 2 (CRY2), the direct repressor of BMAL1 and CLOCK transcriptional activity, synergizes with inhibitors of proteasome, which is required for BMAL1 and CLOCK function, to repress a transcriptional program comprising circadian cycling genes in mMSL TNBC cells. Omics analyses on drug-treated cells implied that this repression of transcription is mediated by the transcription factor binding sites (TFBSs) features in the cis-regulatory elements (CRE) of clock-controlled genes. Through a massive parallel reporter assay, we defined a set of CRE features that are potentially repressed by the specific drug combination. The identification of cis -element enrichment might serve as a new concept of defining and targeting tumor types through the modulation of cis -regulatory programs, and ultimately provide a new paradigm of therapy design for cancer types with unclear drivers like TNBC.
Protein-nucleic acid (PNA) binding plays critical roles in the transcription, translation, regulation, and three-dimensional organization of the genome. Structural models of proteins bound to nucleic acids (NA) provide insights into the chemical, electrostatic, and geometric properties of the protein structure that give rise to NA binding but are scarce relative to models of unbound proteins. We developed a deep learning approach for predicting PNA binding given the unbound structure of a protein that we call PNAbind. Our method utilizes graph neural networks to encode the spatial distribution of physicochemical and geometric properties of protein structures that are predictive of NA binding. Using global physicochemical encodings, our models predict the overall binding function of a protein, and using local encodings, they predict the location of individual NA binding residues. Our models can discriminate between specificity for DNA or RNA binding, and we show that predictions made on computationally derived protein structures can be used to gain mechanistic understanding of chemical and structural features that determine NA recognition. Binding site predictions were validated against benchmark datasets, achieving AUROC scores in the range of 0.92–0.95. We applied our models to the HIV-1 restriction factor APOBEC3G and showed that our model predictions are consistent with and help explain experimental RNA binding data.
DNAproDB (https://dnaprodb.usc.edu/) is a database, visualization tool, and processing pipeline for analyzing structural features of protein-DNA interactions. Here, we present a substantially updated version of the database through additional structural annotations, search, and user interface functionalities. The update expands the number of pre-analyzed protein-DNA structures, which are automatically updated weekly. The analysis pipeline identifies water-mediated hydrogen bonds that are incorporated into the visualizations of protein-DNA complexes. Tertiary structure-aware nucleotide layouts are now available. New file formats and external database annotations are supported. The website has been redesigned, and interacting with graphs and data is more intuitive. We also present a statistical analysis on the updated collection of structures revealing salient patterns in protein-DNA interactions.
Analyzing and visualizing the tertiary structure and complex interactions of RNA is essential for being able to mechanistically decipher their molecular functions in vivo. Secondary structure visualization software can portray many aspects of RNA; however, these layouts are often unable to preserve topological correspondence since they do not consider tertiary interactions between different regions of an RNA molecule. Likewise, quaternary interactions between two or more interacting RNA molecules are not considered in secondary structure visualization tools. The RNAscape webserver produces visualizations that can preserve topological correspondence while remaining both visually intuitive and structurally insightful. RNAscape achieves this by designing a mathematical structural mapping algorithm which prioritizes the helical segments, reflecting their tertiary organization. Non-helical segments are mapped in a way that minimizes structural clutter. RNAscape runs a plotting script that is designed to generate publication-quality images. RNAscape natively supports non-standard nucleotides, multiple base-pairing annotation styles and requires no programming experience. RNAscape can also be used to analyze RNA/DNA hybrid structures and DNA topologies, including G-quadruplexes. Users can upload their own three-dimensional structures or enter a Protein Data Bank (PDB) ID of an existing structure. The RNAscape webserver allows users to customize visualizations through various settings as desired. URL: https://rnascape.usc.edu/.
Development of the malaria parasite, Plasmodium falciparum, is regulated by a limited number of sequence-specific transcription factors (TFs). However, the mechanisms by which these TFs recognize genome-wide binding sites is largely unknown. To address TF specificity, we investigated the binding of two TF subsets that either bind CACACA or GTGCAC DNA sequence motifs and further characterized two additional ApiAP2 TFs, PfAP2-G and PfAP2-EXP, which bind unique DNA motifs (GTAC and TGCATGCA). We also interrogated the impact of DNA sequence and chromatin context on P. falciparum TF binding by integrating high-throughput in vitro and in vivo binding assays, DNA shape predictions, epigenetic post-translational modifications, and chromatin accessibility. We found that DNA sequence context minimally impacts binding site selection for paralogous CACACA-binding TFs, while chromatin accessibility, epigenetic patterns, co-factor recruitment, and dimerization correlate with differential binding. In contrast, GTGCAC-binding TFs prefer different DNA sequence context in addition to chromatin dynamics. Finally, we determined that TFs that preferentially bind divergent DNA motifs may bind overlapping genomic regions due to low-affinity binding to other sequence motifs. Our results demonstrate that TF binding site selection relies on a combination of DNA sequence and chromatin features, thereby contributing to the complexity of P. falciparum gene regulatory mechanisms.
Harmen J. Bussemaker合作论文数Department of Biological Sciences, Columbia University/Department of Systems Biology, Columbia University Irving Medical Center6