Histopathology refers to the microscopic examination of diseased tissues and routinely guides treatment decisions for cancer and other diseases. Currently, this analysis focuses on morphological features but rarely considers gene expression information, which can add an important molecular dimension. Here, we introduce SpotWhisperer, an AI method that links histopathological images to spatial gene expression profiles and their text annotations, enabling molecularly grounded histopathology analysis through natural language. Our method outperforms pathology vision-language models on a newly curated benchmark dataset, dedicated to spatially resolved H&E annotation. Integrated into a web interface, SpotWhisperer enables interactive exploration of cell types and disease mechanisms using free-text queries with access to inferred spatial gene expression profiles. In summary, SpotWhisperer analyzes cost-effective pathology images with spatial gene expression and natural-language AI, demonstrating a path for routine integration of microscopic molecular information into histopathology. ### Competing Interest Statement V.H.K. reports being an invited speaker for Sharing Progress in Cancer Care (SPCC) and Indica Labs; advisory board of Takeda; and sponsored research agreements with Roche and IAG, all unrelated to the current study. V.H.K. is a participant in a patent application on the assessment of cancer immunotherapy biomarkers by digital pathology; a patent application on multimodal deep learning for the prediction of recurrence risk in cancer patients, and a patent application on predicting the efficacy of cancer treatment using deep learning. G.R. is a participant in a patent application on matching cells from different measurement modalities which is not directly related to the current work. Moreover, G.R. is a cofounder of Computomics GmbH, Germany, and one of its shareholders. C.B. is a cofounder and scientific advisor of Myllia Biotechnology and Neurolentech which is not directly related to the current work. The remaining authors declare no competing interests.
BACKGROUND:Computational simulation of biological processes can be a valuable tool for accelerating biomedical research, but usually requires extensive domain knowledge and manual adaptation. Large language models (LLMs) such as GPT-4 have proven surprisingly successful for a wide range of tasks. This study provides proof-of-concept for the use of GPT-4 as a versatile simulator of biological systems. METHODS:We introduce SimulateGPT, a proof-of-concept for knowledge-driven simulation across levels of biological organization through structured prompting of GPT-4. We benchmarked our approach against direct GPT-4 inference in blinded qualitative evaluations by domain experts in four scenarios and in two quantitative scenarios with experimental ground truth. The qualitative scenarios included mouse experiments with known outcomes and treatment decision support in sepsis. The quantitative scenarios included prediction of gene essentiality in cancer cells and progression-free survival in cancer patients. RESULTS:In qualitative experiments, biomedical scientists rated SimulateGPT's predictions favorably over direct GPT-4 inference. In quantitative experiments, SimulateGPT substantially improved classification accuracy for predicting the essentiality of individual genes and increased correlation coefficients and precision in the regression task of predicting progression-free survival. CONCLUSION:This proof-of-concept study suggests that LLMs may enable a new class of biomedical simulators. Such text-based simulations appear well suited for modeling and understanding complex living systems that are difficult to describe with physics-based first-principles simulations, but for which extensive knowledge is available as written text. Finally, we propose several directions for further development of LLM-based biomedical simulators, including augmentation through web search retrieval, integrated mathematical modeling, and fine-tuning on experimental data.
Single-cell RNA-seq characterizes biological samples at unprecedented scale and detail, but data interpretation remains challenging. Here we introduce CellWhisperer, a multimodal machine learning model and software that connects transcriptomes and text for interactive single-cell RNA-seq data analysis. CellWhisperer enables the chat-based interrogation of transcriptome data in English language. To train our model, we created an AI-curated dataset with over a million pairs of RNA-seq profiles and matched textual annotations across a broad range of human biology, and we established a multimodal embedding of matched transcriptomes and text using contrastive learning. Our model enables free-text search and annotation of transcriptome datasets by cell types, states, and other properties in a zero-shot manner and without the need for reference datasets. Moreover, CellWhisperer answers questions about cells and genes in natural-language chats, using a biologically fluent large language model that we fine-tuned to analyze bulk and single-cell transcriptome data across various biological applications. We integrated CellWhisperer with the widely used CELLxGENE browser, allowing users to interactively explore RNA-seq data through an integrated graphical and chat interface. Our method demonstrates a new way of working with transcriptome data, leveraging the power of natural language for single-cell data analysis and establishing an important building block for future AI-based bioinformatics research assistants. ### Competing Interest Statement C.B. is a cofounder and scientific advisor of Myllia Biotechnology and Neurolentech. The remaining authors declare no competing interests.
Computational simulation of biological processes can be a valuable tool in accelerating biomedical research, but usually requires extensive domain knowledge and manual adaptation. Recently, large language models (LLMs) such as GPT-4 have proven surprisingly successful for a wide range of tasks by generating human language at a very large scale. Here we explore the potential of leveraging LLMs as simulators of biological systems. We establish proof-of-concept of a text-based simulator, SimulateGPT, that uses LLM reasoning. We demonstrate good prediction performance for various biomedical applications, without requiring explicit domain knowledge or manual tuning. LLMs thus enable a new class of versatile and broadly applicable biological simulators. This text-based simulation paradigm is well-suited for modeling and understanding complex living systems that are difficult to describe with physics-based first-principles simulation, but for which extensive knowledge and context is available as written text.
Mutations in the splicing factor SF3B1 are frequently occurring in various cancers and drive tumor progression through the activation of cryptic splice sites in multiple genes. Recent studies also demonstrate a positive correlation between the expression levels of wild-type SF3B1 and tumor malignancy. Here, we demonstrate that SF3B1 is a hypoxia-inducible factor (HIF)-1 target gene that positively regulates HIF1 pathway activity. By physically interacting with HIF1α, SF3B1 facilitates binding of the HIF1 complex to hypoxia response elements (HREs) to activate target gene expression. To further validate the relevance of this mechanism for tumor progression, we show that a reduction in SF3B1 levels via monoallelic deletion of Sf3b1 impedes tumor formation and progression via impaired HIF signaling in a mouse model for pancreatic cancer. Our work uncovers an essential role of SF3B1 in HIF1 signaling, thereby providing a potential explanation for the link between high SF3B1 expression and aggressiveness of solid tumors.
Argonaute proteins (AGOs), which play an essential role in cytosolic post-transcriptional gene silencing, have been also reported to function in nuclear processes like transcriptional activation or repression, alternative splicing and, chromatin organization. As most of these studies have been conducted in human cancer cell lines, the relevance of AGOs nuclear functions in the context of mouse early embryonic development remains uninvestigated. Here, we examined a possible role of the AGO1 protein on the distribution of constitutive heterochromatin in mouse embryonic stem cells (mESCs). We observed a specific redistribution of the repressive histone mark H3K9me3 and the heterochromatin protein HP1α, away from pericentromeric regions upon Ago1 depletion. Furthermore, we demonstrated that major satellite transcripts are strongly up-regulated in Ago1_KO mESCs and that their levels are partially restored upon AGO1 rescue. We also observed a similar redistribution of H3K9me3 and HP1α in Drosha_KO mESCs, suggesting a role for microRNAs (miRNAs) in the regulation of heterochromatin distribution in mESCs. Finally, we showed that specific miRNAs with complementarity to major satellites can partially regulate the expression of these transcripts.
MicroRNA (miRNA) loaded Argonaute (AGO) complexes regulate gene expression via direct base pairing with their mRNA targets. Previous works suggest that up to 60% of mammalian transcripts might be subject to miRNA-mediated regulation, but it remains largely unknown which fraction of these interactions are functional in a specific cellular context. Here, we integrate transcriptome data from a set of miRNA-depleted mouse embryonic stem cell (mESC) lines with published miRNA interaction predictions and AGO-binding profiles. Using this integrative approach, combined with molecular validation data, we present evidence that < 10% of expressed genes are functionally and directly regulated by miRNAs in mESCs. In addition, analyses of the stem cell-specific miR-290-295 cluster target genes identify TFAP4 as an important transcription factor for early development. The extensive datasets developed in this study will support the development of improved predictive models for miRNA-mRNA functional interactions.
MicroRNA (miRNA) loaded Argonaute (AGO) complexes regulate gene expression via direct base pairing with their mRNA targets. Previous works suggest that up to 60 Integrative analysis of transcriptome data from miRNA‐depleted mESC lines, miRNA interaction predictions and AGO‐binding profiles allows for the identification of direct and functional miRNA targets. These data support the development of algorithms to predict functional miRNA‐mRNA interactions. Integrative analysis of transcriptome data from miRNA‐depleted mESC lines, miRNA interaction predictions and AGO‐binding profiles allows forthe identification of direct and functional miRNA targets. These data support the development of algorithms to predict functional miRNA‐mRNA interactions.
The Argonaute proteins (AGOs) are well known for their role in post-transcriptional gene silencing in the microRNA (miRNA) pathway. Here we show that in mouse embryonic stem cells, AGO1&2 serve additional functions that go beyond the miRNA pathway. Through the combined deletion of both Agos, we identified a specific set of genes that are uniquely regulated by AGOs but not by the other miRNA biogenesis factors. Deletion of Ago2&1 caused a global reduction of the repressive histone mark H3K27me3 due to downregulation at protein levels of Polycomb repressive complex 2 components. By integrating chromatin accessibility, prediction of transcription factor binding sites, and chromatin immunoprecipitation sequencing data, we identified the pluripotency factor KLF4 as a key modulator of AGO1&2-regulated genes. Our findings revealed a novel axis of gene regulation that is mediated by noncanonical functions of AGO proteins that affect chromatin states and gene expression using mechanisms outside the miRNA pathway.
Background: The COVID-19 pandemic has resulted in 275 million infections and 5.4 million deaths as of December 2021. While effective vaccines are being administered globally, there is still a great need for antiviral therapies as antigenically novel SARS-CoV-2 variants continue to emerge across the globe. Viruses require host factors at every step in their life cycle, representing a rich pool of candidate targets for antiviral drug design. Methods: To identify host factors that promote SARS-CoV-2 infection with potential for broad-spectrum activity across the coronavirus family, we performed genome-scale CRISPR knockout screens in two cell lines (Vero E6 and HEK293T ectopically expressing ACE2) with SARS-CoV-2 and the common cold-causing human coronavirus OC43. Gene knockdown, CRISPR knockout, and small molecule testing in Vero, HEK293, and human small airway epithelial cells were used to verify our findings. Results: While we identified multiple genes and functional pathways that have been previously reported to promote human coronavirus replication, we also identified a substantial number of novel genes and pathways. The website https://sarscrisprscreens.epi.ufl.edu/ was created to allow visualization and comparison of SARS-CoV2 CRISPR screens in a uniformly analyzed way. Of note, host factors involved in cell cycle regulation were enriched in our screens as were several key components of the programmed mRNA decay pathway. The role of EDC4 and XRN1 in coronavirus replication in human small airway epithelial cells was verified. Finally, we identified novel candidate antiviral compounds targeting a number of factors revealed by our screens. Conclusions: Overall, our studies substantiate and expand the growing body of literature focused on understanding key human coronavirus-host cell interactions and exploit that knowledge for rational antiviral drug development.
SUMMARY Mutations in the splicing factor SF3B1 are frequently occurring in various cancers and drive tumor progression through the activation of cryptic splice sites in multiple genes. Recent studies also demonstrate a positive correlation between expression levels of wildtype SF3B1 and tumor malignancy, but underlying mechanisms remain elusive. Here, we report that SF3B1 acts as an activator of HIF signaling through a splicing-independent mechanism. We demonstrate that SF3B1 forms a heterotrimer with HIF1α and HIF1β, facilitating binding of the HIF1 complex to hypoxia response elements (HREs) to activate target gene expression. We further validate the relevance of this mechanism for tumor progression. Monoallelic deletion of Sf3b1 impedes formation and progression of hypoxic pancreatic cancer via impaired HIF signaling, but is well tolerated in normoxic chromophobe renal cell carcinoma. Our work uncovers an essential role of SF3B1 in HIF1 signaling, providing a causal link between high SF3B1 expression and aggressiveness of solid tumors.
MicroRNAs (miRNAs) are well-studied small noncoding RNAs involved in post-transcriptional gene regulation in a wide range of organisms, including mammals. Their function is mediated by base pairing with their target RNAs. Although many features required for miRNA-mediated repression have been described, the identification of functional interactions is still challenging. In the last two decades, numerous Machine Learning (ML) models have been developed to predict their putative targets. In this review, we summarize the biological knowledge and the experimental data used to develop these ML models. Recently, Deep Neural Network-based models have also emerged in miRNA interaction modeling. We thus outline established and emerging models to give a perspective on the future developments needed to improve the identification of genes directly regulated by miRNAs.
SUMMARY:Single-guide RNAs (sgRNAs) targeting the same gene can significantly vary in terms of efficacy and specificity. PAVOOC (Prediction And Visualization of On- and Off-targets for CRISPR) is a web-based CRISPR sgRNA design tool that employs state of the art machine learning models to prioritize most effective candidate sgRNAs. In contrast to other tools, it maps sgRNAs to functional domains and protein structures and visualizes cut sites on corresponding protein crystal structures. Furthermore, PAVOOC supports homology-directed repair template generation for genome editing experiments and the visualization of the mutated amino acids in 3D.AVAILABILITY AND IMPLEMENTATION:PAVOOC is available under https://pavooc.me and accessible using modern browsers (Chrome/Chromium recommended). The source code is hosted at github.com/moritzschaefer/pavooc under the MIT License. The backend, including data processing steps, and the frontend are implemented in Python 3 and ReactJS, respectively. All components run in a simple Docker environment.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
This paper aims to explore the neural patterns between Chinese and Germans for electroencephalogram (EEG)-based emotion recognition. Both Chinese and German subjects, wearing electrode caps, watched video stimuli that triggered positive, neutral, and negative emotions. Two emotion classifiers are trained on Chinese EEG data and German EEG data, respectively. The experiment results indicate that: a) German neural patterns are basically in accordance with Chinese ones; b) the main difference lies in the upper temporal region in Delta band which activates more when a German is in positive mood; and c) the Chinese positive emotion achieves the best accuracy while German emotions share the approximate accuracy. Moreover, Gamma band serves as the critical band for both German and Chinese emotion recognition.