Complex diseases often emerge from coordinated, system-level cellular state changes that are difficult to address with target centric drug discovery. Phenotypic drug discovery offers a principled alternative but remains constrained by the cost and scalability of pathologically relevant assays. Here we present GEMGen, a large language modelc-based framework that performs in silico phenotypic drug discovery by generating small molecules directly from transcriptomic representations of cellular states. GEMGen encodes desired phenotypic transitions as text-based representations of up- and down-regulated gene sets, enabling transferable modeling across experimental platforms and data modalities. Trained on large scale chemical perturbation data, GEMGen robustly identifies phenotype-oriented compounds and mechanistically related but structurally distinct candidates across multiple benchmarks. Applied to signatures induced by genetic perturbations, GEMGen produces small molecules that phenocopy gene knockdown effects and identifies chemically novel inhibitors, including previously unreported KEAP1 inhibitors that activate NRF2 signaling. Extending this approach to a disease relevant model of fibrosis, GEMGen generates compounds that reverse profibrotic transcriptional programs and cellular phenotypes. These results establish a scalable framework for translating transcriptomic phenotypes into candidate therapeutic molecules, enabling systems-level exploration of vast chemical space and offering a complementary in silico counterpart to physical phenotypic drug screens. ### Competing Interest Statement The authors have declared no competing interest.
Abstract Single-cell multi-omics technologies have recently advanced to enable the profiling of epigenomic, transcriptomic, and proteomic layers within individual cells, offering new opportunities to characterize cellular states as integrated biological systems. However, developing a unified framework that can seamlessly integrate diverse omics modalities and remain robust to heterogeneous modality missingness remains challenging. Existing methods are often designed for specific modalities or modality pairs, relying on dataset-specific training or paired measurements. Here we present HoloCell, to our knowledge the first generative foundation model for joint representation learning and generative modeling across all three major single-cell omics modalities, i.e., epigenomics, transcriptomics, and proteomics. HoloCell contains over 860 million parameters and is pretrained on the Human-Multi-Omics-Corpus, which comprises approximately 468 million single-cell profiles across these three omics layers, corresponding to over 425 billion tokens. HoloCell introduces a a simple yet biologically motivated hierarchical tokenization strategy that encodes cis-regulatory elements, genes, and proteins as structured tokens within a shared modeling framework. We evaluated HoloCell across single-omics representation learning, paired multi-omics integration, unpaired multi-omics alignment, and cross-modal generation via iterative diffusion and remasking, demonstrating its superior performance and flexibility across diverse omics tasks. From a representation perspective, HoloCell provides a unified digital mapping of cellular states across multiple omics layers, capturing cell heterogeneity as an integrated system. From a generation perspective, its iterative diffusion and remasking frame-work permits flexible generation orders beyond fixed left-to-right causality, enabling in silico simulation of multi-omics information flow. Together, these capabilities position HoloCell as a versatile foundation model toward the emerging concept of a virtual cell, offering both systematic characterization and generative simulation of cellular systems within a unified framework.
Accurate prediction of RNA secondary structure underpins transcriptome annotation, mechanistic analysis of non-coding RNAs, and RNA therapeutic design. Recent gains from deep learning and RNA foundation models are difficult to interpret because current benchmarks may overestimate generalization across RNA families. We present the Comprehensive Hierarchical Annotation of Non-coding RNA Groups (CHANRG), a benchmark of 170,083 structurally non-redundant RNAs curated from more than 10 million sequences in Rfam 15.0 using structure-aware deduplication, genome-aware split design and multiscale structural evaluation. Across 29 predictors, foundation-model methods achieved the highest held-out accuracy but lost most of that advantage out of distribution, whereas structured decoders and direct neural predictors remained markedly more robust. This gap persisted after controlling for sequence length and reflected both loss of structural coverage and incorrect higher-order wiring. Together, CHANRG and a padding-free, symmetry-aware evaluation stack provide a stricter and batch-invariant framework for developing RNA structure predictors with demonstrable out-of-distribution robustness.
Protein engineering holds substantial promise for designing proteins with customized functions, yet the vast landscape of potential mutations versus limited laboratory capacity constrains the discovery of optimal sequences. Here, to address this, we present the μProtein framework, which accelerates protein engineering by combining μFormer, a deep learning model for accurate mutational effect prediction, with μSearch, a reinforcement learning algorithm designed to efficiently navigate the protein fitness landscape using μFormer as an oracle. μProtein leverages single-mutation data to predict optimal sequences with complex, multi-amino-acid mutations through its modelling of epistatic interactions and a multi-step search strategy. In addition to strong performance on benchmark datasets, μProtein identified high-gain-of-function multi-point mutants for the enzyme β-lactamase, surpassing one of the highest-known activity levels, in wet laboratory, trained solely on single-mutation data. These results demonstrate μProtein’s capability to discover impactful mutations across the vast protein sequence space, offering a robust, efficient approach for protein optimization. μProtein, combining deep learning and reinforcement learning, is developed to design high-function proteins. This framework, trained only on single-mutation data, discovers multi-site β-lactamase mutants with up to 2,000× growth rates.
Advances in natural language processing and large language models have sparked growing interest in modeling DNA, often referred to as the "language of life". However, DNA modeling poses unique challenges. First, it requires the ability to process ultra-long DNA sequences while preserving single-nucleotide resolution, as individual nucleotides play a critical role in DNA function. Second, success in this domain requires excelling at both generative and understanding tasks: generative tasks hold potential for therapeutic and industrial applications, while understanding tasks provide crucial insights into biological mechanisms and diseases. To address these challenges, we propose HybriDNA, a decoder-only DNA language model that incorporates a hybrid Transformer-Mamba2 architecture, seamlessly integrating the strengths of attention mechanisms with selective state-space models. This hybrid design enables HybriDNA to efficiently process DNA sequences up to 131kb in length with single-nucleotide resolution. HybriDNA achieves state-of-the-art performance across 33 DNA understanding datasets curated from the BEND, GUE, and LRB benchmarks, and demonstrates exceptional capability in generating synthetic cis-regulatory elements (CREs) with desired properties. Furthermore, we show that HybriDNA adheres to expected scaling laws, with performance improving consistently as the model scales from 300M to 3B and 7B parameters. These findings underscore HybriDNA's versatility and its potential to advance DNA research and applications, paving the way for innovations in understanding and engineering the "language of life".
Foundation models have revolutionized natural language processing and artificial intelligence, significantly enhancing how machines comprehend and generate human languages. Inspired by the success of these foundation models, researchers have developed foundation models for individual scientific domains, including small molecules, materials, proteins, DNA, RNA and even cells. However, these models are typically trained in isolation, lacking the ability to integrate across different scientific domains. Recognizing that entities within these domains can all be represented as sequences, which together form the "language of nature", we introduce Nature Language Model (NatureLM), a sequence-based science foundation model designed for scientific discovery. Pre-trained with data from multiple scientific domains, NatureLM offers a unified, versatile model that enables various applications including: (i) generating and optimizing small molecules, proteins, RNA, and materials using text instructions; (ii) cross-domain generation/design, such as protein-to-molecule and protein-to-RNA generation; and (iii) top performance across different domains, matching or surpassing state-of-the-art specialist models. NatureLM offers a promising generalist approach for various scientific tasks, including drug discovery (hit generation/optimization, ADMET optimization, synthesis), novel material design, and the development of therapeutic proteins or nucleotides. We have developed NatureLM models in different sizes (1 billion, 8 billion, and 46.7 billion parameters) and observed a clear improvement in performance as the model size increases.
A select few genes act as pivotal drivers in the process of cell state transitions. However, finding key genes involved in different transitions is challenging. Here, to address this problem, we present CellNavi, a deep learning-based framework designed to predict genes that drive cell state transitions. CellNavi builds a driver gene predictor upon a cell state manifold, which captures the intrinsic features of cells by learning from large-scale, high-dimensional transcriptomics data and integrating gene graphs with directional connections. Our analysis shows that CellNavi can accurately predict driver genes for transitions induced by genetic, chemical and cytokine perturbations across diverse cell types, conditions and studies. By leveraging a biologically meaningful cell state manifold, it is proficient in tasks involving critical transitions such as cellular differentiation, disease progression and drug response. CellNavi represents a substantial advancement in driver gene prediction and cell state manipulation, opening new avenues in disease biology and therapeutic discovery.
A select few genes act as pivotal drivers in the process of cell state transitions. However, finding key genes involved in different transitions is challenging. To address this problem, we present CellNavi, a deep learning-based framework designed to predict genes that drive cell state transitions. CellNavi builds a driver gene predictor upon a cell state manifold, which captures the intrinsic features of cells by learning from large-scale, high-dimensional transcriptomics data and integrating gene graphs with causal connections. Our analysis shows that CellNavi can accurately predict driver genes for transitions induced by genetic modifications and chemical treatments across diverse cell types, conditions, and studies. It is proficient in tasks involving critical transitions such as cellular differentiation, disease progression, and drug response. CellNavi represents a substantial advancement in the methodology for predicting driver genes and manipulating cell states, opening up new research opportunities in disease biology and therapeutic innovation. ### Competing Interest Statement P.D., F.J., S.Z., C.L., Y.M., H.X., G.L., H.L., and T.L. are paid employees of Microsoft Research. The remaining authors declare no competing interests.
AbstractGenerative drug design facilitates the creation of compounds effective against pathogenic target proteins. This opens up the potential to discover novel compounds within the vast chemical space and fosters the development of innovative therapeutic strategies. However, the practicality of generated molecules is often limited, as many designs focus on a narrow set of drug-related properties, failing to improve the success rate of subsequent drug discovery process. To overcome these challenges, we develop TamGen, a method that employs a GPT-like chemical language model and enables target-aware molecule generation and compound refinement. We demonstrate that the compounds generated by TamGen have improved molecular quality and viability. Additionally, we have integrated TamGen into a drug discovery pipeline and identified 14 compounds showing compelling inhibitory activity against the Tuberculosis ClpP protease, with the most effective compound exhibiting a half maximal inhibitory concentration (IC50) of 1.9 μM. Our findings underscore the practical potential and real-world applicability of generative drug design approaches, paving the way for future advancements in the field.
Linking cis -regulatory sequences to target genes has been a long-standing challenge. In this study, we introduce CREaTor, an attention-based deep neural network designed to model cis -regulatory patterns for genomic elements up to 2 Mb from target genes. Coupled with a training strategy that predicts gene expression from flanking candidate cis -regulatory elements (cCREs), CREaTor can model cell type-specific cis -regulatory patterns in new cell types without prior knowledge of cCRE-gene interactions or additional training. The zero-shot modeling capability, combined with the use of only RNA-seq and ChIP-seq data, allows for the ready generalization of CREaTor to a broad range of cell types.
Structure-based drug design is drawing growing attentions in computer-aided drug discovery. Compared with the virtual screening approach where a pre-defined library of compounds are computationally screened, de novo drug design based on the structure of a target protein can provide novel drug candidates. In this paper, we present a generative solution named TamGent (Target-aware molecule generator with Transformer) that can directly generate candidate drugs from scratch for a given target, overcoming the limits imposed by existing compound libraries. Following the Transformer framework (a state-of-the-art framework in deep learning), we design a variant of Transformer encoder to process 3D geometric information of targets and pre-train the Transformer decoder on 10 million compounds from PubChem for candidate drug generation. Systematical evaluation on candidate compounds generated for targets from DrugBank shows that both binding affinity and drugability are largely improved. TamGent outperforms previous baselines in terms of both effectiveness and efficiency. The method is further verified by generating candidate compounds for the SARS-CoV-2 main protease and the oncogenic mutant KRAS G12C. The results show that our method not only re-discovers previously verified drug molecules , but also generates novel molecules with better docking scores, expanding the compound pool and potentially leading to the discovery of novel drugs.
Understanding protein sequences is vital and urgent for biology, healthcare, and medicine. Labeling approaches are expensive yet time-consuming, while the amount of unlabeled data is increasing quite faster than that of the labeled data due to low-cost, high-throughput sequencing methods. In order to extract knowledge from these unlabeled data, representation learning is of significant value for protein-related tasks and has great potential for helping us learn more about protein functions and structures. The key problem in the protein sequence representation learning is to capture the co-evolutionary information reflected by the inter-residue co-variation in the sequences. Instead of leveraging multiple sequence alignment as is usually done, we propose a novel method to capture this information directly by pre-training via a dedicated language model, i.e., Pairwise Masked Language Model (PMLM). In a conventional masked language model, the masked tokens are modeled by conditioning on the unmasked tokens only, but processed independently to each other. However, our proposed PMLM takes the dependency among masked tokens into consideration, i.e., the probability of a token pair is not equal to the product of the probability of the two tokens. By applying this model, the pre-trained encoder is able to generate a better representation for protein sequences. Our result shows that the proposed method can effectively capture the inter-residue correlations and improves the performance of contact prediction by up to 9% compared to the MLM baseline under the same setting. The proposed model also significantly outperforms the MSA baseline by more than 7% on the TAPE contact prediction benchmark when pre-trained on a subset of the sequence database which the MSA is generated from, revealing the potential of the sequence pre-training method to surpass MSA based methods in general.
BackgroundA universally applicable approach that provides standard HALE measurements for different regions has yet to be developed because of the difficulties of health information collection. In this study, we developed a natural language processing (NLP) based HALE estimation approach by using individual-level electronic medical records (EMRs), which made it possible to calculate HALE timely in different temporal or spatial granularities.MethodsWe performed diagnostic concept extraction and normalisation on 13•99 million EMRs with NLP to estimate the prevalence of 254 diseases in WHO Global Burden of Disease Study (GBD). Then, we calculated HALE in Chongqing, 2017, by using the life table technique and Sullivan's method, and analysed the contribution of diseases to the expected years "lost" due to disability (DLE).FindingsOur method identified a life expectancy at birth (LE0) of 77•9 years and health-adjusted life expectancy at birth (HALE0) of 71•7 years for the general Chongqing population of 2017. In particular, the male LE0 and HALE0 were 76•3 years and 68•9 years, respectively, while the female LE0 and HALE0 were 80•0 years and 74•4 years, respectively. Cerebrovascular diseases, cancers, and injuries were the top three deterioration factors, which reduced HALE by 2•67, 2•15, and 1•19 years, respectively.InterpretationThe results demonstrated the feasibility and effectiveness of EMRs-based HALE estimation. Moreover, the method allowed for a potentially transferable framework that facilitated a more convenient comparison of cross-sectional and longitudinal studies on HALE between regions. In summary, this study provided insightful solutions to the global ageing and health problems that the world is facing.FundingNational Key R and D Program of China ( 2018YFC2000400 ).
Background: Early detection of influenza activity followed by timely response is a critical component of preparedness for seasonal influenza epidemic and influenza pandemic. However, most relevant studies were conducted at the regional or national level with regular seasonal influenza trends. There are few feasible strategies to forecast influenza activity at the local level with irregular trends. Methods: Multi-source electronic data, including historical percentage of influenza-like illness (ILI%), weather data, Baidu search index and Sina Weibo data of Chongqing, China, were collected and integrated into an innovative Self-adaptive AI Model (SAAIM), which was constructed by integrating Seasonal Autoregressive Integrated Moving Average model and XGBoost model using a self-adaptive weight adjustment mechanism. SAAIM was applied to ILI% forecast in Chongqing from 2017 to 2018, of which the performance was compared with three previously available models on forecasting. Findings: ILI% showed an irregular seasonal trend from 2012 to 2018 in Chongqing. Compared with three reference models, SAAIM achieved the best performance on forecasting ILI% of Chongqing with the mean absolute percentage error (MAPE) of 11.9%, 7.5%, and 11.9% during the periods of the year 2014-2016, 2017, and 2018 respectively. Among the three categories of source data, historical influenza activity contributed the most to the forecast accuracy by decreasing the MAPE by 19.6%, 43.1%, and 11.1%, followed by weather information (MAPE reduced by 3.3%, 17.1%, and 2.2%), and Internet-related public sentiment data (MAPE reduced by 1.1%, 0.9%, and 1.3%). Interpretation: Accurate influenza forecast in areas with irregular seasonal influenza trends can be made by SAAIM with multi-source electronic data. (c) 2019 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Mitochondria generate most cellular energy and are targeted by multiple pathogens during infection. In turn, metazoans employ surveillance mechanisms such as the mitochondrial unfolded protein response (UPRmt) to detect and respond to mitochondrial dysfunction as an indicator of infection. The UPRmt is an adaptive transcriptional program regulated by the transcription factor ATFS-1, which induces genes that promote mitochondrial recovery and innate immunity. The bacterial pathogen Pseudomonas aeruginosa produces toxins that disrupt oxidative phosphorylation (OXPHOS), resulting in UPRmt activation. Here, we demonstrate that Pseudomonas aeruginosa exploits an intrinsic negative regulatory mechanism mediated by the Caenorhabditis elegans bZIP protein ZIP-3 to repress UPRmt activation. Strikingly, worms lacking zip-3 were impervious to Pseudomonas aeruginosa-mediated UPRmt repression and resistant to infection. Pathogen-secreted phenazines perturbed mitochondrial function and were the primary cause of UPRmt activation, consistent with these molecules being electron shuttles and virulence determinants. Surprisingly, Pseudomonas aeruginosa unable to produce phenazines and thus elicit UPRmt activation were hypertoxic in zip-3-deletion worms. These data emphasize the significance of virulence-mediated UPRmt repression and the potency of the UPRmt as an antibacterial response.
Different representations of the same concept could often be seen in scientific reports and publications. Entity normalization (or entity linking) is the task to match the different representations to their standard concepts. In this paper, we present a two-step ensemble CNN method that normalizes microbiology-related entities in free text to concepts in standard dictionaries. The method is capable of linking entities when only a small microbiology-related biomedical corpus is available for training, and achieved reasonable performance in the online test of the BioNLP-OST19 shared task Bacteria Biotope.
Mitochondria form a cellular network of organelles, or cellular compartments, that efficiently couple nutrients to energy production in the form of ATP. As cancer cells rely heavily on glycolysis, historically mitochondria and the cellular pathways in place to maintain mitochondrial activities were thought to be more relevant to diseases observed in non-dividing cells such as muscles and neurons. However, more recently it has become clear that cancers rely heavily on mitochondrial activities including lipid, nucleotide and amino acid synthesis, suppression of mitochondria-mediated apoptosis as well as oxidative phosphorylation (OXPHOS) for growth and survival. Considering the variety of conditions and stresses that cancer cell mitochondria may incur such as hypoxia, reactive oxygen species and mitochondrial genome mutagenesis, we examine potential roles for a mitochondrial-protective transcriptional response known as the mitochondrial unfolded protein response (UPRmt) in cancer cell biology.