AlphaFold, a groundbreaking artificial intelligence model developed by DeepMind, has transformed the field of structural biology by predicting protein structures with unprecedented accuracy. Despite its widespread recognition and application across academia and industry, comprehensive reviews detailing AlphaFold's unexpected applications within the molecular sciences remain scarce. In this review, we critically examine AlphaFold's emerging roles across diverse molecular scientific disciplines. Specifically, we highlight its applications in enzyme engineering and drug development, nucleic acid modeling and vaccine design, the development of protein-based materials and targeted drug delivery systems, and modeling of complex systems and biological networks. To conclude, the review outlines potential future developments and enduring challenges within the application of AlphaFold to molecular sciences. Overall, this review aims to systematically analyze the most recent advances; explore novel interdisciplinary applications of AlphaFold within the realms of biology, chemistry, and materials science; and offer insights into future directions for research and application.
Gene expression involves a series of complex regulatory processes. It is known that self-feedback regulation plays an important role in it. However, self-feedback burst dynamics in single-cell sequencing data has not been extensively studied yet. Here we propose a single-cell stochastic burst model with coupled positive- and negative-feedback gene expression circuits and then analyze its dynamical behavior using mouse fibroblast scRNA-seq data. The burst dynamics parameters are inferred genomewide and the self-feedback regulation patterns are identified. The results show that positive feedback can restore a bimodal distribution of gene expression, while negative feedback and no feedback often lead to a unimodal distribution. Additionally, we analyze the effects of feedback regulation types on burst dynamics, namely, burst size, burst frequency, and noise. It is found that as the gene mean increases, the burst frequency and burst size of the three feedback types show an increasing trend, while noise shows a decreasing trend. On the other hand, the mean burst frequency of negative feedback is greater than that of positive feedback, while the mean burst size and noise of positive feedback are greater than those of negative feedback. This helps us understand the gene expression patterns, cell differentiation, and fate determination.
Imbalanced data, where certain classes are significantly underrepresented in a dataset, is a widespread machine learning (ML) challenge across various fields of chemistry, yet it remains inadequately addressed. This data imbalance can lead to biased ML or deep learning (DL) models, which fail to accurately predict the underrepresented classes, thus limiting the robustness and applicability of these models. With the rapid advancement of ML and DL algorithms, several promising solutions to this issue have emerged, prompting the need for a comprehensive review of current methodologies. In this review, we examine the prominent ML approaches used to tackle the imbalanced data challenge in different areas of chemistry, including resampling techniques, data augmentation techniques, algorithmic approaches, and feature engineering strategies. Each of these methods is evaluated in the context of its application across various aspects of chemistry, such as drug discovery, materials science, cheminformatics, and catalysis. We also explore future directions for overcoming the imbalanced data challenge and emphasize data augmentation via physical models, large language models (LLMs), and advanced mathematics. The benefit of balanced data in new material design and production and the persistent challenges are discussed. Overall, this review aims to elucidate the prevalent ML techniques applied to mitigate the impacts of imbalanced data within the field of chemistry and offer insights into future directions for research and application.
Understanding the temporal dynamics of gene expression within spatial contexts is essential for deciphering cellular differentiation. RNA velocity, which estimates the future state of gene expression by distinguishing spliced from unspliced mRNA, offers a powerful tool for studying these dynamics. However, current spatial transcriptomics technologies face limitations in simultaneously capturing both spliced and unspliced transcripts at high resolution. To address this challenge, a novel computational framework called KSRV (Kernel PCA–based Spatial RNA Velocity) that integrates single-cell RNA-seq with spatial transcriptomics using Kernel Principal Component Analysis. It enables accurately inference of RNA velocity in spatially resolved tissue at single-cell resolution. KSRV was validated by using 10x Visium data and MERFISH datasets. The results demonstrate its both accuracy and robustness comparing with the existed method such as SIRV and spVelo. Furthermore, KSRV successfully revealed spatial differentiation trajectories in the mouse brain and during mouse organogenesis, highlighting its potential for advancing our understanding of spatially dynamic biological processes.
Generative artificial intelligence (AI) models, a class of AI techniques that learn data distributions to synthesize novel samples, have emerged as impactful tools across scientific disciplines. In recent years, these models have found extensive applications in fields such as natural language processing and biomedical sciences. Despite their growing influence, comprehensive reviews on the application of generative models in biomolecular sciences remain limited. In this review, we provide a systematic overview of recent advances in generative models applied to biomolecular sciences. We discuss several prominent generative architectures, including variational autoencoders, generative adversarial networks, and diffusion models, highlighting their applications in molecular design and bioinformatics. Additionally, we examine how these models contribute to critical challenges such as molecular property prediction and molecular generation. Finally, we discuss key challenges that remain in this field, including model interpretability, scalability, and the need for high-quality molecular datasets. We highlight emerging research directions that aim to overcome these limitations and propose strategies for improving the reliability and applicability of generative models in biomolecular problems. Through this review, our objective is to provide researchers with a comprehensive understanding of the current landscape of generative modeling in biomolecular sciences and to inspire further advancements in this interdisciplinary area.
Cells must have the ability to respond to various external signals, but our understanding of this signal response is very limited. Here we analyze a simple cascade model of stochastic gene expression, which assumes that the upstream (input) gene is expressed in a constitutive manner whereas the downstream (output) gene is expressed in a bursty fashion via a two-state model, and that the former regulates the latter via transcription factors (products of the former). By large scale sampling in the space of model parameters, we find that the output distributions can exhibit diverse patterns such as unimodal, bimodal or trimodal modes, depending on the regulation way (positive or negative). These patterns, which are different from those in the case of no regulation, are independent of the choice of parameter values and hence qualitatively invariant. Diverse expression patterns resulting from signal response would be important for cell fate decisions and utilized by cells for a better survival in complex environments.
Chaos is omnipresent in nature, and its understanding provides enormous social and economic benefits. However, the unpredictability of chaotic systems is a textbook concept due to their sensitivity to initial conditions, aperiodic behaviour, fractal dimensions, nonlinearity and strange attractors. In this work, we introduce, for the first time, chaotic learning, a novel multiscale topological paradigm that enables accurate predictions from chaotic systems. We show that seemingly random and unpredictable chaotic dynamics counterintuitively offer unprecedented quantitative predictions. Specifically, we devise multiscale topological Laplacians to embed real-world data into a family of interactive chaotic dynamical systems, modulate their dynamical behaviours and enable the accurate prediction of the input data. As a proof of concept, we consider 28 datasets from four categories of realistic problems: 10 brain waves, four benchmark protein datasets, 13 single-cell RNA sequencing datasets and an image dataset, as well as two distinct chaotic dynamical systems, namely the Lorenz and Rossler attractors. We demonstrate chaotic learning predictions of the physical properties from chaos. Our new chaotic learning paradigm profoundly changes the textbook perception of chaos and bridges topology, chaos and learning for the first time.
It has been experimentally demonstrated that orthogonal small molecules can control both the average level and variability of gene products through selective modulation of gene expression kinetics, but we lack clear understanding of how the mean and variability are independently controlled. Here, we address this interesting issue by analyzing a genetic switch model. First, we demonstrate that the protein distribution exhibits three representative modes capturing essential features of expression noise in this model. Second, employing a binomial moment approach, we derive the exact analytical expressions for both the mean protein abundance and the coefficient of variation (CV). Third, we show that different combinations of regulating reaction rates in the presence and absence of feedback can achieve distinct CVs while keeping the same mean. While our investigation implies that the controlled modulation of cell-to-cell heterogeneity in gene expression can be achieved within an isogenic population, our results may provide theoretical guidelines for designing gene functional circuits with tunable noise characteristics.
Anesthetics are crucial in surgical procedures and therapeutic interventions, but they come with side effects and varying levels of effectiveness, calling for novel anesthetic agents that offer more precise and controllable effects. Targeting Gamma-aminobutyric acid (GABA) receptors, the primary inhibitory receptors in the central nervous system, could enhance their inhibitory action, potentially reducing side effects while improving the potency of anesthetics. In this study, we introduce a proteomic learning of GABA receptor-mediated anesthesia based on 24 GABA receptor subtypes by considering over 4000 proteins in protein-protein interaction (PPI) networks and over 1.5 millions known binding compounds. We develop a corresponding drug-target interaction network to identify potential lead compounds for novel anesthetic design. To ensure robust proteomic learning predictions, we curated a dataset comprising 136 targets from a pool of 980 targets within the PPI networks. We employed three machine learning algorithms, integrating advanced natural language processing (NLP) models such as pretrained transformer and autoencoder embeddings. Through a comprehensive screening process, we evaluated the side effects and repurposing potential of over 180,000 drug candidates targeting the GABRA5 receptor. Additionally, we assessed the ADMET (absorption, distribution, metabolism, excretion, and toxicity) properties of these candidates to identify those with near-optimal characteristics. This approach also involved optimizing the structures of existing anesthetics. Our work presents an innovative strategy for the development of new anesthetic drugs, optimization of anesthetic use, and deeper understanding of potential anesthesia-related side effects.
The ongoing opioid crisis highlights the urgent need for novel therapeutic strategies that can be rapidly deployed. This study presents a novel approach to identify potential repurposable drugs for the treatment of opioid addiction, aiming to bridge the gap between transcriptomic data analysis and drug discovery. Specifically, we perform a meta-analysis of seven transcriptomic data sets related to opioid addiction by differential gene expression (DGE) analysis and propose a novel multiscale topological differentiation to identify key genes from a protein-protein interaction (PPI) network derived from DEGs. This method uses persistent Laplacians to accurately single out important nodes within the PPI network through a multiscale manner to ensure high reliability. Subsequent functional validation by pathway enrichment and rigorous data curation yields 1,865 high-confidence targets implicated in opioid addiction, which are cross-referenced with DrugBank to compile a repurposing candidate list. To evaluate drug-target interactions, we construct predictive models utilizing two natural language processing-derived molecular embeddings and a conventional molecular fingerprint. Based on these models, we prioritize compounds with favorable binding affinity profiles, and select candidates that are further assessed through molecular docking simulations to elucidate their receptor-level interactions. Additionally, pharmacokinetic and toxicological evaluations are performed via ADMET (absorption, distribution, metabolism, excretion, and toxicity) profiling, providing a multidimensional assessment of druggability and safety. This study offers a generalizable approach for drug repurposing in other complex diseases beyond opioid addiction.
In prokaryotic and eukaryotic cells, most genes are transcribed in a bursty fashion on one hand and complex gene regulations may lead to complex promoter structure on the other hand. This raises an unsolved issue: how does promoter structure shape transcriptional bursting kinetics characterized by burst size and frequency? Here we analyze stochastic models of gene transcription, which consider complex regulatory mechanisms. Notably, we develop an efficient method to derive exact burst-size distributions. The analytical results show that if the promoter of a gene contains only one active state, the burst size indeed follows a geometric distribution, in agreement with the previous result derived under certain limiting conditions. However, if it contains a multitude of active states, the burst size in general obeys a non-geometric distribution, which is a linearly weighted sum of geometric distributions. This superposition principle reveals the essential feature of bursting kinetics in complex cases of transcriptional regulation although it seems that there has been no direct experimental confirmation. The derived burst-size distributions not only highlight the importance of promoter structure in regulating bursting kinetics, but can be also used in the exact inference of this kinetics based on experimental data.
Transformer models have emerged as pivotal tools within the realm of drug discovery, distinguished by their unique architectural features and exceptional performance in managing intricate data landscapes. Leveraging the innate capabilities of transformer architectures to comprehend intricate hierarchical dependencies inherent in sequential data, these models showcase remarkable efficacy across various tasks, including new drug design and drug target identification. The adaptability of pre-trained transformer-based models renders them indispensable assets for driving data-centric advancements in drug discovery, chemistry, and biology, furnishing a robust framework that expedites innovation and discovery within these domains. Beyond their technical prowess, the success of transformer-based models in drug discovery, chemistry, and biology extends to their interdisciplinary potential, seamlessly combining biological, physical, chemical, and pharmacological insights to bridge gaps across diverse disciplines. This integrative approach not only enhances the depth and breadth of research endeavors but also fosters synergistic collaborations and exchange of ideas among disparate fields. In our review, we elucidate the myriad applications of transformers in drug discovery, as well as chemistry and biology, spanning from protein design and protein engineering, to molecular dynamics (MD), drug target identification, transformer-enabled drug virtual screening (VS), drug lead optimization, drug addiction, small data set challenges, chemical and biological image analysis, chemical language understanding, and single cell data. Finally, we conclude the survey by deliberating on promising trends in transformer models within the context of drug discovery and other sciences.
Genetically identical cell populations growing in the same environment show a large degree of heterogeneity in gene expression profiles, resulting in significant phenotypic consequences. One of the important sources of this heterogeneity is transcriptional bursting. In recent years, single cell transcriptome sequencing (scRNA-seq) data have been widely used to infer the kinetics of transcriptional bursting. Moreover, the increasing evidence showed that genes are jointly regulated by several competitive pathways when activated by external signals, and there is molecular memory in the process of gene state switching. Therefore, a gene expression model considering both gene activation pathway and molecular memory can better reflect the nature of transcriptional bursting. In this paper, we used an approximate Bayesian computation algorithm called ABC-PRC combined with a gene expression model that considered crosstalk and memory, to infer the kinetics of transcriptional bursting. We analyze the internal relationship between transcriptional bursting and gene expression using scRNA-seq data obtained from mouse embryonic stem cells and mouse embryonic fibroblasts. Both model analysis and data simulations demonstrate that the bimodal coefficient increases with the increase of burst size, and the noise intensity decreases with the increase of burst frequency. Additionally, we show that some previously proposed models can be simplified as special cases of the gene expression model and inference algorithm proposed here under certain circumstances.
Complex molecular details of transcriptional regulation can be coarse -grained by assuming that reaction waiting times for promoter -state transitions, the mRNA synthesis, and the mRNA degradation follow general distributions. However, how such a generalized two -state model is analytically solved is a long-standing issue. Here we first present analytical formulas of burst -size distributions for this model. Then, we derive an iterative equation for the mRNA moment -generating function, by which mRNA raw and binomial moments of any order can be conveniently calculated. The analytical results obtained in the special cases of phase -type waiting -time distributions not only provide insights into the mechanisms of complex transcriptional regulations but also bring conveniences for experimental data -based statistical inferences.
Global exponential periodicity of nonlinear neural networks with multiple time-varying delays is investigated. Such neural networks cannot be written in the vector-matrix form because of the existence of the multiple delays. It is noted that although the neural network with multiple time-varying delays has been investigated by Lyapunov-Krasovskii functional method in the literature, the sufficient conditions in the linear matrix inequality form have not been obtained. Two sets of sufficient conditions in the linear matrix inequality form are established by Lyapunov-Krasovskii functional and linear matrix inequality to ensure that two arbitrary solutions of the neural network with multiple delays attract each other exponentially. This is a key prerequisite to prove the existence, uniqueness, and global exponential stability of periodic solutions. Some examples are provided to demonstrate the effectiveness of the established results. We compare the established theoretical results with the previous results and show that the previous results are not applicable to the systems in these examples.
Intracellular biochemical networks often display large fluctuations in the molecule numbers or the concentrations of reactive species, making molecular approaches necessary for system descriptions. For Markovian reaction networks, the fluctuation-dissipation theorem (FDT) has been well established and extensively used in fast evaluation of fluctuations in reactive species. For non-Markovian reaction networks, however, the similar FDT has not been established so far. Here, we present a generalized FDT (gFDT) for a large class of non-Markovian reaction networks where general intrinsic-event waiting-time distributions account for the effect of intrinsic noise and general stochastic reaction delays represent the impact of extrinsic noise from environmental perturbations. The starting point is a generalized chemical master equation (gCME), which describes the probabilistic behavior of an equivalent Markovian reaction network and identifies the structure of the original non-Markovian reaction network in terms of stoichiometries and effective transition rates (extensions of common reaction propensity functions). From this formulation follows directly the solution of the linear noise approximation of the stationary gCME for all the components in the non-Markovian reaction network. While the gFDT can quickly trace noisy sources in non-Markovian reaction networks, example analysis verifies its effectiveness.
Transcription involves gene activation, nuclear RNA export (NRE) and RNA nuclear retention (RNR). All these processes are multistep and biochemical. A multistep reaction process can create memories between reaction events, leading to non-Markovian kinetics. This raises an unsolved issue: how does molecular memory affect stochastic transcription in the case that NRE and RNR are simultaneously considered? To address this issue, we analyze a non-Markov model, which considers multistep activation, multistep NRE and multistep RNR can interpret many experimental phenomena. In order to solve this model, we introduce an effective transition rate for each reaction. These effective transition rates, which explicitly decode the effect of molecular memory, can transform the original non-Markov issue into an equivalent Markov one. Based on this technique, we derive analytical results, showing that molecular memory can significantly affect the nuclear and cytoplasmic mRNA mean and noise. In addition to the results providing insights into the role of molecular memory in gene expression, our modeling and analysis provide a paradigm for studying more complex stochastic transcription processes.
Recurrent neural networks are widely used in time series prediction and classification. However, they have problems such as insufficient memory ability and difficulty in gradient back propagation. To solve these problems, this paper proposes a new algorithm called SS-RNN, which directly uses multiple historical information to predict the current time information. It can enhance the long-term memory ability. At the same time, for the time direction, it can improve the correlation of states at different moments. To include the historical information, we design two different processing methods for the SS-RNN in continuous and discontinuous ways, respectively. For each method, there are two ways for historical information addition: 1) direct addition and 2) adding weight weighting and function mapping to activation function. It provides six pathways so as to fully and deeply explore the effect and influence of historical information on the RNNs. By comparing the average accuracy of real datasets with long short-term memory, Bi-LSTM, gated recurrent units, and MCNN and calculating the main indexes (Accuracy, Precision, Recall, and F1-score), it can be observed that our method can improve the average accuracy and optimize the structure of the recurrent neural network and effectively solve the problems of exploding and vanishing gradients.
Apart from intrinsic stochastic variability, gene expression also involves stochastic reaction delay arising from heterogeneity and fluctuation processes, which can affect the efficiency of reactants (e.g., mRNA or protein) in exploring their environments. In contrast to the former that has been extensively investigated, the impact of the latter on gene expression remains not fully understood. Here, we analyze a non-Markovian model of bursty gene expression with general delay distribution. We analytically find that the effect of stochastic reaction delay is equivalent to the introduction of negative feedback, and stationary protein distribution only depends on the mean of the delay and is independent of its distribution. We numerically show that the stochastic reaction delay always slightly amplifies the mean protein level but remarkably reduces the protein noise (quantified by the ratio of the variance over the squared average). Our analysis indicates that stochastic reaction delay is an important factor affecting gene expression.
An important task in the post-gene era is to understand the role of stochasticity in gene regulation. Here, we analyze a cascade model of stochastic gene expression, where the upstream gene stochastically generates proteins that regulate, as transcription factors, stochastic synthesis of the downstream output. We find that in contrast to fast input fluctuations that do not change the behavior of the downstream system qualitatively, slow input fluctuations can induce different modes of the distribution of downstream output and even stochastic focusing or defocusing of the downstream output level, although the regulatory protein follows the same distribution in both cases. This finding is counterintuitive but can have broad biological implications, e.g., slow input rather than fast fluctuations may both increase the survival probability of cells and enhance the sensitivity of intracellular regulation. In addition, we find that input fluctuations can minimize the output noise.