We propose LOGDIFF (Logical Guidance for the Exact Composition of Diffusion Models), a guidance framework for diffusion models that enables principled constrained generation with complex logical expressions at inference time. We study when exact score-based guidance for complex logical formulas can be obtained from guidance signals associated with atomic attributes and constraints. First, we derive an exact Boolean calculus that provides a sufficient condition for exact logical guidance. Specifically, if a formula admits a circuit representation in which conjunctions combine conditionally independent subformulas and disjunctions combine subformulas that are either conditionally independent or mutually exclusive, exact logical guidance is achievable. In this case, the guidance signal can be computed exactly from atomic scores and posterior probabilities using an efficient recursive algorithm. Moreover, we show that, for commonly encountered classes of distributions, any desired Boolean formula is compilable into such a circuit representation. Second, by combining atomic guidance scores with posterior probability estimates, we introduce a hybrid guidance approach that bridges classifierguidance and classifier-free guidance, applicable to both compositional logical guidance and standard conditional generation. We demonstrate the effectiveness of our framework on multiple image and protein structure generation tasks.
Diverse risk genes have been identified for neurodevelopmental disorders (NDDs), but how these genes converge on similar biological pathways in neurons, and thus give rise to similar phenotypes, is unclear. Here we apply a pooled CRISPR approach to successfully target 23 NDD loss-of-function genes with roles in chromatin biology and examine convergent effects on gene expression across human induced pluripotent stem cell-derived neural progenitor cells, glutamatergic neurons and GABAergic neurons. Points of convergence vary between these cell types, with the greatest number of convergent genes and strongest convergent networks in mature glutamatergic neurons, where they broadly represent synaptic, epigenetic and, unexpectedly, mitochondrial pathways. The most convergent networks were observed between NDD genes with shared biological annotations, clinical associations and co-expression patterns in human post-mortem brain. Drugs that were predicted to reverse convergent transcriptomic signatures and/or arousal and sensory processing behaviors ameliorated behavioral phenotypes in zebrafish NDD gene mutants. These results suggest that convergent effects of NDD risk genes could provide clinically useful insights.
Functional magnetic resonance imaging (fMRI) is a widely used technique for studying the brain. Recent methods that utilize graph neural networks (GNNs) for analysis of brain functional connectivity have shown great potential for the classification of brain disorders, such as Alzheimer's disease (AD). However, these methods often assume a preset number of functional modules across all subjects, which overlooks inter-subject variability. In addition, the discovered modules are rarely used to directly guide the learned connectivity patterns. Here, to address these issues, we propose a Meta Probabilistic Pooling GNN (MPP-GNN). We frame the model's task as a coupled, bilevel optimization that performs adaptive graph partitioning hierarchically to discover subject-specific modules and then uses the discovered brain modules as an explicit prior to guide edge refinement and representation learning. We validate MPP-GNN on two public datasets for AD classification, achieving the highest AUC in comparison to established baselines for both datasets. Furthermore, our analysis demonstrates that MPP-GNN shows significant alignment with the canonical functional-network organization defined by the Yeo brain atlas and reveals a network-level dedifferentiation pattern for AD.
We introduce a framework to analyse interpretability in deep learning, by drawing on a formal notion of model semantics from the philosophy of science. We argue that interpretability is only one aspect of a model’s semantics and illustrate our framework with examples from biomedicine.
A major challenge in near-term quantum computing is its application to large real-world datasets due to scarce quantum hardware resources. One approach to enabling tractable quantum models for such datasets involves finding low-dimensional representations that preserve essential information for downstream analysis. In classical machine learning, variational autoencoders (VAEs) facilitate efficient data compression, representation learning for subsequent tasks, and novel data generation. However, no quantum model has been proposed that exactly captures all of these features for direct application to quantum data on quantum computers. Some existing quantum models for data compression lack regularization of latent representations, thus preventing direct use for generation and control of generalization. Others are hybrid models with only some internal quantum components, impeding direct training on quantum data. To address this, we present a fully quantum framework,-QVAE, which encompasses all the capabilities of classical VAEs and can be directly applied to map both classical and quantum data to a lower-dimensional space, while effectively reconstructing much of the original state from it. Our model utilizes regularized mixed states to attain optimal latent representations. It accommodates various divergences for reconstruction and regularization. Furthermore, by accommodating mixed states at every stage, it can utilize the full data density matrix and allow for a training objective defined on probabilistic mixtures of input data. Doing so, in turn, makes efficient optimization possible and has potential implications for private and federated learning. In addition to exploring the theoretical properties of-QVAE, we demonstrate its performance on representative genomics and synthetic data. Our results indicate that-QVAE consistently learns representations that better utilize the capacity of the latent space and exhibits similar or better performance compared with matched classical models.
Non-small cell lung cancer (NSCLC) shows variable responses to immunotherapy, highlighting the need for biomarkers to guide patient selection. We applied a spatial multi-omics approach to 234 advanced NSCLC patients treated with programmed death 1-based immunotherapy across three cohorts to identify biomarkers associated with outcome. Spatial proteomics (n = 67) and spatial compartment-based transcriptomics (n = 131) enabled profiling of the tumor immune microenvironment (TIME). Using spatial proteomics, we identified a resistance cell-type signature including proliferating tumor cells, granulocytes, vessels (hazard ratio (HR) = 3.8, P = 0.004) and a response signature, including M1/M2 macrophages and CD4 T cells (HR = 0.4, P = 0.019). We then generated a cell-to-gene resistance signature using spatial transcriptomics, which was predictive of poor outcomes (HR = 5.3, 2.2, 1.7 across Yale, University of Queensland and University of Athens cohorts), while a cell-to-gene response signature predicted favorable outcomes (HR = 0.22, 0.38 and 0.56, respectively). This framework enables robust TIME modeling and identifies biomarkers to support precision immunotherapy in NSCLC.
Triple-negative breast cancer (TNBC) is an aggressive subtype with poor prognoses and limited biomarkers for predicting treatment outcomes [1]. Tumor-infiltrating lymphocytes and PD-L1 expression, while associated with immune checkpoint inhibitor (ICI) efficacy, fail to predict responses reliably [2, 3]. Spatial transcriptomics is a cutting-edge technology that enables the precise mapping of gene expression within tissue samples. However, the high cost of sequencing limits its clinical utility. This study aims to develop cost-effective predictive biomarkers by integrating spatial transcriptomic data and machine learning to impute gene expression from widely available Hematoxylin and Eosin (H&E) images. This study involves two cohorts of TNBC patients treated with pembrolizumab (N=75) and durvalumab (N=57) in the neoadjuvant setting. Digital Spatial Profiling (DSP) was used to generate Spatial transcriptomics data. Gene signatures predicting ICI outcomes were developed and validated using Least Absolute Shrinkage and Selection Operator (LASSO) regression models. Imputed gene expression was derived from H&E images, employing adaptive spatial graph neural networks (asGNN [4]), and evaluated against DSP data on a subset of patients. Signatures were trained using imputed transcriptomics directly and in combination with a hold-out set of H&E images. Prediction accuracies were evaluated against spatial transcriptomics and clinical outcomes using mean squared error and AUC. Imputed gene expression from H&E images closely mirrors spatial transcriptomic profiles. Imputed spatial gene expression using asGNN achieved better prediction of DSP expression than competing methods (HE2RNA [5] and Sequoia [6]) and non-adaptive GNN methods (p=0.031). Signatures trained directly using H&E imputed spatial expression achieved accurate prediction of treatment outcomes (AUC=0.75±0.24). Integrating the DSP transcriptomics further increased the accuracy of the predictions; the signatures trained using imputed transcriptomics combined with DSP data achieved higher prediction accuracy (AUC=0.85±0.08), which significantly outperformed the non-DSP trained models (p=0.021). This study demonstrates the feasibility of leveraging H&E images for gene expression imputation in TNBC. Future work will optimize models across diverse cancer types and expand validation with independent datasets. This innovation advances personalized medicine by bridging the gap between cutting-edge spatial transcriptomics and routine clinical diagnostics. Thazin Nwe Aung, Tianci Song, Lajos Pusztai, Mark Gerstein, David L. Rimm, Jonathan H. Warrell. Integrating digital spatial profiling and H&E images to develop predictive biomarkers for immunotherapy outcomes in triple-negative breast cancer from imputed spatial gene expression [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2424.
Quantum neural networks (QNNs) are gaining increasing interest due to their potential to detect complex patterns in data by leveraging uniquely quantum phenomena. This makes them particularly promising for biomedical applications. In these applications and in other contexts, increasing statistical power often requires aggregating data from multiple participants. However, sharing data, especially sensitive information like personal genomic sequences, raises significant privacy concerns. Quantum federated learning offers a way to collaboratively train QNN models without exposing private data. However, it faces major limitations, including high communication overhead and the need to retrain models when the task is modified. To overcome these challenges, we propose a privacy-preserving QNN training scheme that utilizes mixed quantum states to encode ensembles of data. This approach allows for the secure sharing of statistical information while safeguarding individual data points. QNNs can be trained directly on these mixed states, eliminating the need to access raw data. Building on this foundation, we introduce protocols supporting multi-party collaborative QNN training applicable across diverse domains. Our approach enables secure QNN training with only a single round of communication per participant, provides high training speed and offers task generality, i.e., new analyses can be conducted without reacquiring information from participants. We present the theoretical foundation of our scheme's utility and privacy protections, which prevent the recovery of individual data points and resist membership inference attacks as measured by differential privacy. We then validate its effectiveness on three different datasets with a focus on genomic studies with an indication of how it can used in other domains without adaptation.
Over three hundred and seventy-three risk genes, broadly enriched for roles in neuronal communication and gene expression regulation, underlie risk for autism spectrum disorder (ASD) and developmental delay (DD). Functional genomic studies of subsets of these genes consistently indicate a convergent role in neurogenesis, but how these diverse risk genes converge on a smaller number of biological pathways in mature neurons is unclear. To uncover shared downstream impacts between neurodevelopmental disorder (NDD) risk genes, here we apply a pooled CRISPR approach to contrast the transcriptomic impacts of targeting 29 NDD loss-of-function genes across human induced pluripotent stem cell (hiPSC)-derived neural progenitor cells, glutamatergic neurons, and GABAergic neurons. Points of convergence vary between the cell types of the brain and are greatest in mature glutamatergic neurons, where they broadly target not just synaptic and epigenetic, but unexpectedly, mitochondrial biology. The strongest convergent networks occur between NDD genes with common co-expression patterns in the post-mortem brain, biological annotations, and clinical associations, suggesting that convergence may one-day inform patient stratification and treatment. Towards this, ten out of eleven drugs tested that were predicted to reverse convergent signatures in human cells and/or arousal and sensory processing behaviors in zebrafish ameliorated at least one behavioral phenotype in vivo. Altogether, robust convergence in post-mitotic neurons represents a clinically actionable therapeutic window.
AbstractPurpose: We aim to improve the prediction of response or resistance to immunotherapies in patients with melanoma. This goal is based on the hypothesis that current gene signatures predicting immunotherapy outcomes show only modest accuracy due to the lack of spatial information about cellular functions and molecular processes within tumors and their microenvironment. Experimental Design: We collected gene expression data spatially from three cellular compartments defined by CD68+ macrophages, CD45+ leukocytes, and S100B+ tumor cells in 55 immunotherapy-treated melanoma specimens using Digital Spatial Profiling–Whole Transcriptome Atlas. We developed a computational pipeline to discover compartment-specific gene signatures and determine if adding spatial information can improve patient stratification. Results: We achieved robust performance of compartment-specific signatures in predicting the outcome of immune checkpoint inhibitors in the discovery cohort. Of the three signatures, the S100B signature showed the best performance in the validation cohort (N = 45). We also compared our compartment-specific signatures with published bulk signatures and found the S100B tumor spatial signature outperformed previous signatures. Within the eight-gene S100B signature, five genes (PSMB8, TAX1BP3, NOTCH3, LCP2, and NQO1) with positive coefficients predict the response, and three genes (KMT2C, OVCA2, and MGRN1) with negative coefficients predict the resistance to treatment. Conclusions: We conclude that the spatially defined compartment signatures utilize tumor and tumor microenvironment–specific information, leading to more accurate prediction of treatment outcome, and thus merit prospective clinical assessment.
Supplementary Figure S5. Testing of S100B spatial signature in two published datasets.
MOTIVATION:Spatial transcriptomics technologies, which generate a spatial map of gene activity, can deepen the understanding of tissue architecture and its molecular underpinnings in health and disease. However, the high cost makes these technologies difficult to use in practice. Histological images co-registered with targeted tissues are more affordable and routinely generated in many research and clinical studies. Hence, predicting spatial gene expression from the morphological clues embedded in tissue histological images provides a scalable alternative approach to decoding tissue complexity. RESULTS:Here, we present a graph neural network based framework to predict the spatial expression of highly expressed genes from tissue histological images. Extensive experiments on two separate breast cancer data cohorts demonstrate that our method improves the prediction performance compared to the state-of-the-art, and that our model can be used to better delineate spatial domains of biological interest. AVAILABILITY AND IMPLEMENTATION:https://github.com/song0309/asGNN/.
Single-cell genomics is a powerful tool for studying heterogeneous tissues such as the brain. Yet little is understood about how genetic variants influence cell-level gene expression. Addressing this, we uniformly processed single-nuclei, multiomics datasets into a resource comprising >2.8 million nuclei from the prefrontal cortex across 388 individuals. For 28 cell types, we assessed population-level variation in expression and chromatin across gene families and drug targets. We identified >550,000 cell type–specific regulatory elements and >1.4 million single-cell expression quantitative trait loci, which we used to build cell-type regulatory and cell-to-cell communication networks. These networks manifest cellular changes in aging and neuropsychiatric disorders. We further constructed an integrative model accurately imputing single-cell expression and simulating perturbations; the model prioritized ~250 disease-risk genes and drug targets with associated cell types.
Abstract The widespread use of immunotherapy in lung cancer, and its more recent approval for early stages, underscores the need for biomarkers that can identify the most responsive patients. Spatial transcriptomics, which maps gene expression in its spatial tissue context, offers a unique approach over bulk transcriptomics. By incorporating spatial information, the predictive power of the signature can be enhanced. Here, we aim to develop spatially informed gene signatures that could be translated into clinical RNA in situ assays, distinguishing patients unlikely to benefit from immunotherapy and thereby sparing them unnecessary side effects. We utilized the NanoString GeoMX Whole Transcriptome Atlas for spatially resolved transcriptomic profiling of retrospectively collected lung cancer tissue samples from patients treated with immunotherapy in an advanced-stage setting (N=60). By targeting 18,190 genes within distinct areas of interest (AOIs)—including stromal (macrophages/CD68+ and leukocytes/CD45+) and tumor (cytokeratin, CK+) cells—we developed AOI-specific gene signatures to predict objective responses. These were derived from a robust computational framework employing LASSO logistic regression on a split-sample approach, yielding predictive models for treatment outcome.We achieved high predictive accuracy on the training set, with the area under the curve (AUC) exceeding 0.86 for all AOI-specific spatial signatures, indicating strong potential for clinical application. Validation against an independent cohort (N=42) corroborated the efficacy of these signatures. Our 6-gene tumor signature was validated with an AUC of 0.73 (95% CI: 0.67-0.89, p = 0.009**), while the CD45 5-gene signature showed an AUC of 0.75 (95% CI: 0.53-0.97, p = 0.022*). Our 18-gene CD68 signature trended towards validation but lacked statistical significance. A combined CD68 and CD45 signature predicted outcomes with greater accuracy, achieving an AUC of 0.79 (95% CI: 0.62-0.98, p = 0.0088*). Following gene set enrichment analysis on our differentially expressed and signature genes, we identified genes with positive coefficients in both tumor and stroma signatures that were associated with glucocorticoid response and glycolytic processes linked to T-cell homeostasis, while genes with negative coefficients were associated with epithelial cell differentiation and cytokine production. These associations are concordant with the observation that genes with positive coefficients are predictors of treatment response, while those with negative coefficients indicate resistance.Our findings indicate that AOI-specific signatures predict the immunotherapy outcome in lung cancer with high accuracy, suggesting that spatial assessment can provide substantial predictive information. The high performance of these signatures indicates their potential for prospective clinical applications. Citation Format: Thazin N. Aung, Myrto Moutafi, Ioannis Trontzas, Arutha Kulasinghe, James Monkman, Niki Gavrielatou, Ioannis Vathiotis, Jonathan H. Warrell, David L. Rimm. Spatial-specific gene signatures to predict immunotherapy outcomes in lung cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 1141.
Supplementary Figure S3. Performance of alternative 8-gene signature tested in the S100B compartment of the validation cohort and Heatmaps illustrating the quantile normalized expression counts of the signature genes.
Cultural processes of change bear many resemblances to biological evolution. The underlying units of non-biological evolution have, however, remained elusive, especially in the domain of music. Here, we introduce a general framework to jointly identify underlying units and their associated evolutionary processes. We model musical styles and principles of organization in dimensions such as harmony and form as following an evolutionary process. Furthermore, we propose that such processes can be identified by extracting latent evolutionary signatures from musical corpora, analogously to identifying mutational signatures in genomics. These signatures provide a latent embedding for each song or musical piece. We develop a deep generative architecture for our model, which can be viewed as a type of variational autoencoder with an evolutionary prior constraining the latent space; specifically, the embeddings for each song are tied together via an energy-based prior, which encourages songs close in evolutionary space to share similar representations. As illustration, we analyse songs from the McGill Billboard dataset. We find frequent chord transitions and formal repetition schemes and identify latent evolutionary signatures related to these features. Finally, we show that the latent evolutionary representations learned by our model outperform non-evolutionary representations in such tasks as period and genre prediction.
Graph neural networks (GNNs) have emerged as powerful tools for representation learning. Their efficacy depends on their having an optimal underlying graph. In many cases, the most relevant information comes from specific subgraphs. In this work, we introduce a GNNbased framework (graph-partitioned GNN [GP-GNN]) to partition the GNN graph to focus on the most relevant subgraphs. Our approach jointly learns task-dependent graph partitions and node representations, making it particularly effective when critical features reside within initially unidentified subgraphs. Protein liquid-liquid phase separation (LLPS) is a problem especially well-suited to GP-GNNs because intrinsically disordered regions (IDRs) are known to function as protein subdomains in it, playing a key role in the phase separation process. In this study, we demonstrate how GP-GNN accurately predicts LLPS by partitioning protein graphs into task-relevant subgraphs consistent with known IDRs. Our model achieves state-of-the-art accuracy in predicting LLPS and offers biological insights valuable for downstream investigation.