DNA methylation is a central epigenetic modification that regulates gene expression, maintains genomic stability, and guides cellular differentiation. However, direct measurements of DNA methylation, such as whole genome bisulfite sequencing or DNA methylation arrays, are costly and require substantial DNA input, limiting their scalability for large cohorts and their applicability to emerging modalities such as single cell and spatially resolved transcriptomics. In this study, motivated by the fact that DNA methylation is fundamentally a metabolic process, we investigate whether sample-wise DNA methylation activity can be inferred directly from transcriptomic profiles of genes involved in methionine and one-carbon metabolism. We show that a compact metabolic model comprising seven core reaction steps and 98 genes accurately predicts total DNA methylation activity across matched transcriptomic and methylation datasets from CCLE, TCGA, GTEx, and an independent single-cell multi-omics data set. Building on this framework, we develop Total DNA Methylation Activity (TDMA), a physics-informed neural network based score that enables robust estimation of DNA methylation activity from bulk, single-cell, and spatial transcriptomics data. We demonstrate that TDMA captures methylation-dependent transcriptional regulation and identifies genes and pathways under epigenetic control. Applying TDMA to GTEx, we further reveal strong associations between the predicted total methylation activity, chronological aging, and established epigenetic clocks. We also demonstrated that TDMA can serve as a transcriptomics-derived epigenetic clock and highlights age-dependent roles of folate and glutathione metabolism in epigenetic aging. Applying TDMA to single cell and spatial transcriptomics data collected from pancreatic adenocarcinoma (PDAC), we identified that methionine metabolism and DNA methylation regulates T cell cytotoxicity in the tumor microenvironment of PDAC. Together, this work establishes a scalable, modality-agnostic framework for estimating DNA methylation activity from transcriptomics and provides new insights into the metabolic regulation of epigenetic aging. ### Competing Interest Statement The authors have declared no competing interest.
Background: Accurate forecasting of Alzheimer's Disease (AD) progression is critical for personalized patient management and clinical trial stratification. However, current predictive models often struggle to effectively integrate high-dimensional neuroimaging with longitudinal clinical data. We introduce AD-LLaVA-3D, a novel multimodal framework designed to bridge this gap by adapting large vision-language models for volumetric and temporal forecasting. Methods: We leveraged the LLaVA-NeXT-Video architecture to treat 3D MRI volumes as temporal sequences, enabling the model to process volumetric imaging alongside longitudinal Tabular Clinical Records (TCR). The model was trained on the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort (n=764) and evaluated using a rigorous patient-level split. We assessed its ability to forecast a suite of future clinical indicators (e.g., CDR-SB, MMSE) against traditional machine learning baselines (Lasso, Random Forest, Gradient Boosting) and specialized deep learning models (ResNet-3D, Med-Flamingo). Results: AD-LLaVA-3D demonstrated superior predictive accuracy on the ADNI test set, achieving a Coefficient of Determination ($R^2$) of 0.68 for the critical CDR-SB score, surpassing the best-performing baseline (R^2=0.66). Crucially, in an independent external validation on the Open Access Series of Imaging Studies (OASIS) cohort (n=76), our model exhibited exceptional generalization (R^2=0.82$, $MSE=0.54), whereas comparison models showed significant performance degradation (R^2 < 0.60). Conclusions: This study presents the first application of a video-based multimodal architecture for AD progression forecasting. By effectively integrating 3D MRI with tabular clinical records, AD-LLaVA-3D offers a robust, generalizable tool for monitoring disease trajectories, significantly advancing predictive capabilities beyond current unimodal or static methods. Highlights: First-in-Class Architecture: We introduce the first application of video-based Large Vision-Language Models (LVLMs) to interpret 3D volumetric MRI as a temporal sequence, capturing longitudinal neurodegeneration more effectively than static 3D-CNNs. Robust External Validation: The model achieved superior predictive accuracy (R^2=0.82) on an independent external cohort (OASIS), demonstrating exceptional generalization beyond the training population (ADNI). Data-Efficient Multimodal Integration: We developed a novel prompting strategy that integrates sparse Tabular Clinical Records (TCR) without artificial imputation, allowing the model to leverage incomplete real-world medical history. Clinical Trial Enrichment: By accurately forecasting future cognitive scores (CDR-SB, MMSE), AD-LLaVA-3D serves as a precise screening tool to identify "rapid progressors" for clinical trials, potentially reducing failure rates in drug development. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement NIH ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: ADNI DATA: https://adni.loni.usc.edu/data-samples/adni-data/ OASIS DATA: https://sites.wustl.edu/oasisbrains/ I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
Metastasis remains the leading cause of cancer-related mortality, yet predicting future metastasis is a major clinical challenge due to the lack of validated biomarkers and effective assessment methods. Here, we present EmitGCL, a deep-learning framework that accurately predicts future metastasis and its corresponding biomarkers. Based on a comprehensive benchmarking comparison, EmitGCL outperforms other computational tools across six cancer types from seven cohorts of patients with superior sensitivity and specificity. It captures occult metastatic cells in a patient with a lymph node-negative breast cancer, who was declared to have no evidence of disease by conventional imaging methods but was later confirmed to have metastatic disease. Notably, EmitGCL identifies HSP90AA1 and HSP90AB1 as predictable biomarkers for future breast cancer metastasis, which we validate by in-vitro pharmacological inhibition of HSP90 that reduced breast cancer cell migration and further support across five independent cohorts of patients (n = 420). Furthermore, we demonstrate YY1 transcription factor as a key driver of breast cancer metastasis, which we corroborate with in-silico, CRISPR-based migration assays, and in vivo mouse lung colonization experiments, suggesting that YY1 is a potential therapeutic target for further investigation.
Abstract Modern platforms such as Xenium and Visium HD measure the spatial coordinates and identities of millions of individual RNA molecules per tissue section, providing unprecedented resolution to study tumor heterogeneity and the spatial organization of cellular compartments. However, most current analyses rely on cell segmentation, which can be error-prone in densely packed or morphologically complex tumors. Here, we introduce a segmentation-free statistical framework for modeling RNA molecules in high-resolution spatial transcriptomics data using Log-Gaussian Cox Processes (LGCPs). Building on our recently developed efficient variational method for fitting LGCPs, we treat each RNA molecule as a point in continuous space and model gene-specific log-intensity surfaces with a latent Gaussian random field prior. This framework enables us to (i) infer smooth intensity fields for selected genes, (ii) estimate spatial cross-covariance between these fields as a direct measure of gene-gene co-localization, and (iii) perform direct association analyses quantifying how the local abundance of one gene varies as a function of genes or spatial features, all without requiring cell boundaries. These quantities define molecule-level association scores that capture local enrichment or exclusion of RNA species, facilitating the discovery of ligand-receptor hotspots, metabolic niches, and immune-tumor interaction zones at subcellular resolution. By aggregating molecule-level association patterns over marker gene sets, our approach further supports inference of cell type and subtype-level spatial organization and interactions. Simulation studies demonstrate that our segmentation-free LGCP approach accurately recovers underlying intensity surfaces and spatial associations with favorable runtimes. Overall, this work provides a scalable, model-based tool for leveraging the full richness of high-resolution spatial transcriptomics to map RNA molecule associations in cancer tissues without relying on cell segmentation. Citation Format: Xiao Wang, Nan Zhang, Chi Zhang, Sha Cao. A segmentation-free method for modeling high-resolution spatial transcriptomics at the molecule level [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 6853.
Abstract Metabolism strongly influences how immune cells become activated, differentiate, and lose function in the tumor microenvironment. Understanding the metabolic programs that define different tumor-infiltrating immune cell subtypes is important for identifying pathways that could be targeted to improve immune responses in cancer, yet the distinct metabolic features of major immune cell types—including T cells, myeloid cells, NK cells, and B cells—are still not well defined. To address this, we collected multiple tumor-derived scRNA-seq datasets (>500,000 cells across ∼30 immune cell subtypes). Using these datasets, we conducted metabolic flux inference to identify immune cell-specific metabolic states in the tumor microenvironment. scRNA-seq datasets of tumor-infiltrating immune cells were first analyzed using Seurat, where cell type and subtype identities were assigned based on canonical marker genes and confirmed with additional annotation tools. Pseudobulk and meta-cell representations were created to reduce sparsity, and metabolic flux was estimated using our in-house tool MPOCtrL, which infers reaction-level activity from gene expression using curated metabolic gene lists. Dimensionality reduction and clustering were applied to compare metabolic patterns across immune cell types and subtypes. In particular, for T cells, we have discovered distinct metabolic phenotypes for different T cell subtypes. Exhausted T cells showed high serine and glutamate metabolism, low glucose uptake, and reduced β-oxidation, while T follicular helper cells showed opposite trends with low serine and glutamate metabolism but high glucose uptake and elevated glycolysis. Glycolysis was also high in Th1-like cells but reduced in CD4 cytotoxic effector and Tn-like cells. Lactate-associated flux was enriched in Treg and Trm cells, while Tcm and Tn-like cells showed low lactate production. Ketone body metabolism was high in Th17-biased CD4 T cells and low in Trm cells. Fatty acid pathways varied across subsets, with Tn-like cells showing low fatty acid synthesis and Th1-like cells showing high β-oxidation. Pentose phosphate pathway activity remained uniformly low across T-cell subsets. This analysis reveals clear metabolic patterns across immune cell types and subtypes in the tumor microenvironment. By defining these metabolic programs, our work provides a basis for identifying metabolic pathways that could be targeted to change immune cell behavior and improve anti-tumor immunity. Citation Format: Yue Fang, Changlin Wan, Haiqi Zhu, Zheng An, Pengtao Dang, Chi Zhang, Sha Cao. Single-cell metabolic flux analysis defines distinct metabolic programs across tumor-infiltrating immune cells [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5458.
The exponential trajectory of biomedical literature has precipitated a fundamental "synthesis gap" in metabolic research, where critical mechanistic insights remain fragmented across hundreds of thousands of disjointed full-text articles, preventing the consolidation of a global mechanistic view. Here, we present MetaKnogic-Alpha, a foundational mechanistic knowledge substrate designed to bridge this gap by transforming unstructured literature into a navigable, logic-based resource. MetaKnogic-Alpha synthesizes over 100K full-text articles into a hyper-relational hypergraph structure, preserving the n-ary relational logic inherent in complex metabolic pathways. To ensure biological rigor, we implemented a hierarchical discovery protocol: an autonomous reasoning agent first enriches query nomenclature for domain-specific precision, followed by a multi-hop topological expansion within the hypergraph to surface functional neighbors, such as enzymatic co-factors and distal regulators, often lost in traditional search paradigms. Crucially, the system subjects all literature-derived insights to a deterministic biochemical grounding against a curated metabolic reaction network, significantly mitigating the risk of probabilistic hallucinations common in standalone generative models. In rigorous benchmarking, MetaKnogic-Alpha achieved a mechanistic accuracy of 0.98 in scenarios where supporting evidence was present, providing a robustly attributable audit trail back to the primary literature via PubMed Central Identifiers. We designate this primary release as "alpha" to establish the foundational architectural logic for a burgeoning million-scale resource. By compressing the synthesis of thousands of papers from a multi-month manual effort into several hours of automated discovery, MetaKnogic-Alpha serves as a high-fidelity research companion that augments the human expert's ability to resolve complex metabolic interactions and identify novel therapeutic drivers in precision oncology. ### Competing Interest Statement The authors have declared no competing interest.
Multiple kernel clustering (MKC) effectively extracts intrinsic and complementary information from data by integrating diverse kernel functions. The allocation of kernel weights is crucial for MKC performance and is closely related to the relationships among kernel matrices. However, it is very difficult to fully capture the intricate relationships among high-dimensional matrices because previous research mostly relies on predefined metrics to characterize the correlation among kernel matrices. To address this challenge, a novel MKC model called AMKC-LRR is proposed that adaptively learns the interrelations among kernel matrices using low-rank representation and unifies this learning process with the clustering task within an optimization framework. Furthermore, an effective alternate optimization algorithm is designed to solve the resulting problem. Extensive experiments and statistical tests conducted on twelve commonly used benchmark datasets show that our proposed model performs favorably in comparison to state-of-the-art MKC methods. The source code for the proposed model is available at https://github.com/bala23-w/AMKC-LRR/.
Skeletal muscle loss in pancreatic cancer is a significant cause of morbidity and mortality for patients. In order to understand myocytes changes we examined myonuclei- and myofiber-specific dynamics during pancreatic cancer cachexia progression. Single-nucleus RNA-seq was used to interrogate myonuclear gene expression, and RNAscope and immunofluorescence characterized myofiber-specific changes. Bulk RNA-seq of skeletal muscle provided a whole-muscle transcriptomic profile. Cachexia induces a progressive loss of muscle differentiation factor Maf and its target Myh4 , accompanied by increased expression of Myh1 and Myh2 . This myofiber dedifferentiation occurs without evidence for fiber type shifting, regeneration, or proliferation. Single-nuclei analysis reveals global shifts in myofiber gene expression identity including the identification of a cachexia only myonuclear subpopulation. Cachexia gene expression was not restricted solely to this PDAC-specific myonuclear subpopulation and did not overlap with Myh1 and Myh2 expressing myonuclei early in cachexia. Altogether, PDAC cachexia elicits distinct transcriptional responses across different myonuclear populations. These results reveal population-specific heterogeneity in cachexia gene activation, rather than a uniform upregulation of cachexia mediators across muscle tissue. Our data suggest that myonuclei fate occurs prior to overt muscle wasting when cachexia gene expression only modestly overlaps with differentiation factors, with a strong association after irreversible muscle wasting. These findings explain the challenge of effectively targeting skeletal muscle wasting in cancer cachexia requires addressing the changing cell population induced through non overlapping mechanisms.
Spatially resolved transcriptomics (SRT) offers insights into tumor microenvironments (TME), mapping gene and cell activities across tissue regions. While systems biology methods utilize SRT data to characterize tissue and cell neighborhood variations, high costs and limited scalability hinder large-scale applications and clinical correlations. In contrast, histological imaging (HI) provides accessible, high-resolution morphological insights with clinical context. Existing HI-based gene expression imputation methods rely on small datasets and cover a limited gene fraction, leaving much of the transcriptome unexplored. This study develops a robust framework for imputing spatially resolved gene expression profiles from HI data by integrating HI-based imputation with scGPT - a foundation models trained on single-cell RNA-seq (scRNA-seq) data. We hypothesize HI data can impute genes linked to TME, tissue structure, and cell types, while scGPT infers additional genes using gene and cell type-specific relationships. The proposed framework merges an advanced HI-based gene imputation algorithm with scGPT to predict spatially resolved full-transcriptome profiles. The HI-based model first imputes 500-1,000 tissue structure- and cell type-dependent genes. We enhanced existing methods by (1) incorporating cross-pixel multi-attentions for spatial adjustments, (2) enabling residual connections to preserve early features, and (3) introducing a composite loss function for spatial coherence, gene-gene interaction preservation, and distribution alignment. Tissue annotations provide contextual information that enhances model predictivity. scGPT refines and expands predictions into whole gene profiles, leveraging its robust architecture to capture complex molecular relationships. The framework demonstrates strong performance across unseen specimens, achieving low mean squared error (0.38) and high correlations (0.79) while mitigating batch effects. Spatial heatmaps capture tissue heterogeneity, with tumor, immune, and stromal boundaries aligning well with annotations. Gene Enrichment analysis shows the model's capability to capture immune, stromal, and tumor compartments, though challenges remain for underrepresented genes. Integrating scGPT generates accurate and meaningful whole-genome profiles. This enables implementation of systems biology models for transcriptomics - including pathway, metabolism, stemness and CNV analysis - into HI data analysis. This framework bridges histology and transcriptomics by combining advanced imputation algorithms with powerful capabilities of scGPT, enabling scalable, spatially resolved gene expression mapping. Its robust extrapolation transforms understanding of tissue biology and the TME complexity. Paveethran Swaminathan, Pengtao Dang, Haiqi Zhu, Zheng An, Xin Lu, Zhi Li, Yue Fang, Min Yang, Yuhui Wei, Tingbo Guo, Xinyu Zhou, Xiao Wang, Jia Wang, Chi Zhang, Sha Cao. Characterizing biological functional changes in cancer tumor microenvironment by integrating histology-based gene imputation with scRNA-seq foundation models [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 1353.
Metastasis remains the leading cause of cancer-related mortality, yet predicting future metastasis is a major clinical challenge due to the lack of validated biomarkers and effective assessment methods. Here, we present EmitGCL, a deep-learning framework that accurately predicts future metastasis and its corresponding biomarkers. Based on a comprehensive benchmarking comparison, EmitGCL outperformed other computational tools across six cancer types from seven cohorts of patients with superior sensitivity and specificity. It captured occult metastatic cells in a patient with a lymph node-negative breast cancer, who was declared to have no evidence of disease by conventional imaging methods but was later confirmed to have a metastatic disease. Notably, EmitGCL identified HSP90AA1 and HSP90AB1 as predictable biomarkers for future breast cancer metastasis, which was validated across five independent cohorts of patients (n=420). Furthermore, we demonstrated YY1 transcription factor as a key driver of breast cancer metastasis which was validated through in-silico and CRISPR-based migration assays, suggesting that YY1 is a potential therapeutic target for deterring metastasis.
MOTIVATION:As the SARS-CoV-2 virus rapidly evolves, predicting the trajectory of viral mutations has become a critical yet complex task. A deep understanding of future mutation patterns, in particular the mutations that will prevail in the near future, is vital in steering diagnostics, therapeutics, and vaccine strategies for disease control. RESULTS:In this study, we developed a model to forecast future SARS-CoV-2 mutation surges in real-time, using historical mutation frequency data from the USA. We transformed the temporal prediction problem into a supervised learning framework using a sliding window approach. This involved breaking the time series of mutation frequencies into very short segments. Considering the time-dependent nature of the data, we focused on modeling the first-order derivative of the mutation frequency. We predicted the final derivative in each segment based on the preceding derivatives, employing various machine learning methods, including random forest, XGBoost, support vector machine, and neural network models. Empowered by the novel transformation strategy and the high capacity of machine learning models, we observed low prediction error that is confined within 0.1% and 1% when making predictions of mutation rates for the future 30 and 80 days, respectively. In addition, the method also led to a notable increase in prediction accuracy compared to traditional time-series models, as evidenced by much lower MAE (Mean Absolute Error) and MSE (Mean Squared Error) for predictions made within different time horizons. To further assess the method's effectiveness and robustness in predicting mutation patterns for unforeseen mutations, we first designed a synthetic case where we categorized all mutations into three major patterns. The model demonstrated its robustness by accurately predicting unseen mutation patterns when training on data from two pattern categories while testing on the third pattern category, showcasing its potential in forecasting a variety of mutation trajectories. We then applied our method to prediction for a recent time frame between 1 January 2025 and 10 June 2025, for both the USA and UK, where the model training was conducted using frequency sequence data collected between 12 December 2019 and 26 January 2023 in the USA. The model demonstrated superior performance for both datasets. AVAILABILITY AND IMPLEMENTATION:To enhance accessibility and utility, we built our methodology into a GitHub package (https://github.com/ZhouXY199502/SWD). Our method has the potential applicability to study other infectious diseases or forecasting tasks, thus extending its relevance beyond the current COVID pandemic.
Antigen processing and presentation via major histocompatibility complex (MHC) molecules are central to immune surveillance. Yet, quantifying the dynamic activity of MHC class I and II antigen presentation remains a critical challenge, particularly in diseases like cancer, infection and autoimmunity where these pathways are often disrupted. Current methods fall short in providing precise, sample-specific insights into antigen presentation, limiting our understanding of immune evasion and therapeutic responses. Here, we present PSAA (PINN-empowered Systems Biology Analysis of Antigen Presentation Activity), which is designed to estimate sample-wise MHC class I and class II antigen presentation activity using bulk, single-cell, and spatially resolved transcriptomics or proteomics data. By reconstructing MHC pathways and employing pathway flux estimation, PSAA offers a detailed, stepwise quantification of MHC pathway activity, enabling predictions of gene-specific impacts and their downstream effects on immune interactions. Benchmarked across diverse omics datasets and experimental validations, PSAA demonstrates a robust prediction accuracy and utility across various disease contexts. In conclusion, PSAA and its downstream functions provide a comprehensive framework for analyzing the dynamics of MHC antigen presentation using omics data. By linking antigen presentation to immune cell activity and clinical outcomes, PSAA both elucidates key mechanisms driving disease progression and uncovers potential therapeutic targets.
BACKGROUND: Despite much research, advances in early prediction of spontaneous preterm birth (sPTB) has been slow. The evolving field of circulating microparticle (CMP) biology may identify novel blood-based, and clinically useful, biomarkers. OBJECTIVE: To test the ability of a previously identified, 7-marker set of CMP-derived proteins from the first trimester of pregnancy, in the form of an in vitro diagnostic multivariate index assay (IVDMIA), to stratify pregnant patients according to their risk for sPTB. STUDY DESIGN: We employed a previously validated set of CMP protein biomarkers, utilizing mass spectrometry assays and a nested case- control design in a subset of participants from the Nulliparous Pregnancy Outcomes Study: monitoring mothers-to-be (nuMoM2b). We evaluated these biomarkers in the form of an IVDMIA to predict risk for sPTB at different gestational ages. Plasma samples collected at 9- to 13-weeks' gestation were analyzed. The IVDMIA assigned subjects to 1 of 3 sPTB risk categories: low risk (LR), moderate risk (MR), or high risk (HR). Independent validation on a set-aside set confirmed the IVDMIA's performance in risk stratification. RESULTS: Samples from 400 participants from the nuMoM2b cohort were used for the study; of these, 160 delivered<37 weeks and 240 delivered at term. Through Monte Carlo simulation in which the validation results were adjusted based on actual weekly sPTB incidence rates in the nuMoM2b cohort, the IVDMIA stratifications demonstrated statistically significant differences among the risk groups in time-to-event (birth) analysis (P<.0001). The incidence- rate adjusted cumulative risks of sPTB at <= 32 weeks' gestation were 0.4%, 1.6%, and 7.5%, respectively for the LR, M R, and HR groups, respectively. Compared to the LR group, the corresponding risk ratios of the IVDMIA assigned MR and HR group were 4.25 (95% confidence interval [CI] 2.2-7.9) and 19.92 (95% CI 10.4-37.4), respectively. CONCLUSION: A first trimester CMP protein biomarker panel can be used to stratify risk for sPTB at different gestational ages. Such a multitiered stratification tool could be used to assess risk early in pregnancy to enable timely clinical management and interventions, and, ultimately, to enable the development of tailored care pathways for sPTB prevention.
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general statistical framework that estimates location-specific coefficients linking a response feature to high-dimensional spatial predictors. S3R unites structured sparsity with a minimum-spanning-tree-guided smoothness penalty, yielding coefficient fields that are coherent within neighborhoods yet permit sharp boundaries. In synthetic data, S3R accurately recovers spatially varying effects, selects relevant predictors, and preserves known boundaries. Applied to Visium-based ST data, S3R recapitulates layer-specific target-TF associations in human dorsolateral prefrontal cortex with concordant layer-wise correlations in matched single-cell data. In acute Haemophilus ducreyi skin infection, S3R converts spot-level gene expression mixtures into cell type-attributed expression fields, revealing per-cell type spatial gradients, and improving concordance of spatially variable gene calls when tests are applied to these demixed fields. In pancreatic ductal adenocarcinoma, S3R builds cross-cell-type, cross-gene co-variation tensors that quantify cell-cell interaction strength at gene-pair resolution and nominate interacting genes whose pathway enrichments align with established stromal-epithelial and immune crosstalk. An efficient implementation scales to large assays, and on a Xenium-based breast cancer dataset, S3R delineates the contributions at gene-gene, local neighborhood, and global context-level to target gene expression. Because responses and predictors in S3R are user-defined, it could flexibly address diverse biological questions within a single, scalable, and interpretable regression framework.
Background and Aims:Over 80% of patients with pancreatic cancer experience cachexia, characterized by severe muscle and fat loss. While all the mechanistic understanding comes from preclinical models, the translatable nature of these findings to humans remains a critical gap due to the limited knowledge of human cachexia biology. Methods:We generated matched gene and microRNA profiles from rectus abdominis muscle of 55 pancreatic ductal adenocarcinoma and 18 control subjects. Differentially expressed genes and microRNAs were identified at 1.5-fold change and p<0.05. Results:Gene expression results revealed a striking sex-specific difference at the expression and pathway levels. In both sexes, co-expression gene network analysis identified more significant modules and hub genes at 1-month of weight loss than the traditionally used six months, suggesting that gene alterations may be more dynamic in the early stages of the disease progression. When comparing hub genes from humans to experimental models of cachexia, genes such as RELA, DDX21, WDR75, PTPN1, and CRIP3 exhibited similar patterns of expression, suggesting their potential role in cachexia. microRNAs also exhibited sex-specific expression. Although several common miRNAs were identified between sexes, their gene targets differed, indicating that microRNAs may regulate gene targets in a sex-specific manner. Conclusions:The dataset can serve as a resource for validating preclinical findings and exploring previously unexplored molecules in cachexia. Future studies will functionally characterize the role of the hub genes and microRNAs in cachexia. This is the first study to identify sex-specific genes and microRNAs from a single cancer type.
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general framework that estimates location-specific coefficients linking a response feature to high-dimensional spatial predictors. S3R unites structured sparsity with a minimum-spanning-tree–guided smoothness penalty, yielding coefficient fields that are coherent within neighborhoods yet permit sharp boundaries. S3R enables large scale data analysis with an efficient implementation using a reduced MST graph, multi-GPU training, and parallel hyperparameter search. In synthetic data, S3R accurately recovers spatially varying effects, selects relevant predictors, and preserves known boundaries. Applied to human dorsolateral prefrontal cortex, S3R recapitulates layer-specific target–TF associations with concordant layer-wise correlations in matched single-cell data. In acute Haemophilus ducreyi skin infection, S3R converts spot-level mixtures into cell type–attributed expression fields and reveals spatial gradients; applying SVG tests to these fields increases concordance and recovers gradients missed by spot-level methods. In pancreatic ductal adenocarcinoma, S3R constructs cross–cell type, cross-gene co-variation tensors that prioritize interactions among cell types with interacting genes enriching pathways consistent with known biology. Because responses and predictors in S3R are user-defined, it could flexibly address diverse ST questions within a single, scalable, and interpretable regression framework. ### Competing Interest Statement The authors have declared no competing interest.
Few-shot learning (FSL) is a machine learning paradigm that aims to generalize models from a small number of labeled examples, typically fewer than 10 per class. FSL is particularly crucial in biomedical, environmental, materials, and mechanical sciences, where samples are limited and data collection is often prohibitively costly, time-consuming, or ethically constrained. In this study, we present an innovative approach to FSL by demonstrating that a Large Multi-Modal Model (LMMM), trained on a set of independent tasks spanning diverse domains, task types, and input modalities, can substantially improve the generalization of FSL models, outperforming models based on conventional meta-learning on tasks of the same type. To support this, we first constructed a Multi-Modal Model Few-shot Dataset (M3FD, over 10K+ few-shot samples), which includes 2D RGB images, 2D/3D medical scans, tabular and time-course datasets, from which we manually curated FSL tasks such as classification. We further introduced M3F (Multi-Modal Model for Few-shot learning framework), a novel Large Multi-Modal Model framework tailored for data-constrained scientific applications. M3F supports a wide range of scientific data types through a modular pipeline. By fine-tuning the model on M3FD, M3F improves model performance, making LMMM feasible for real-world FSL deployment. The source code is located at https://github.com/ptdang1001/M3F. To democratize access to complex FSL data and promote reproducibility for public usage, M3FD is paired with a flexible and user-friendly tool that enables efficient querying, task-specific sampling, and preprocessing. Together, our dataset and framework offer a unified, scalable solution that significantly lowers the barrier to applying LMMMs in data-scarce scientific domains.
Understanding the relationship between cardiovascular burden, amyloid, and cognition in Alzheimer’s disease (AD) is essential for targeted interventions, especially in ethnically diverse populations where research remains limited. This study aimed to investigate these relationships in a cohort of Korean older adults along the AD spectrum. 526 participants from the Korean Brain Aging Study for the Early Diagnosis and Prediction of Alzheimer’s Disease (KBASE) cohort were included in this study. Vascular burden was quantified using Framingham Risk Score (FRS) and participants were categorized into four groups based on combinations of FRS (FRS High or FRS Low with a median split) and amyloid status (Aβ+ or Aβ- based on a cut-off of 1.2373). Cognitive function was evaluated using standardized neuropsychological tests processed with structural equation models to produce domain scores for memory, executive functioning, language, and visuospatial. ANOVA was employed at baseline to analyze cognitive differences among these groups and within each clinical diagnosis. Longitudinal mixed effects models spanning a period of four years from the initial visit captured cognitive changes over time within these groups (Figure 1). Significant group and pairwise differences were observed among the four groups in all cognitive domains (p < 0.0001). Stratified analysis within each clinical diagnoses group revealed that CN individuals in FRS high Aβ- demonstrated significantly lower memory scores compared to those with FRS low Aβ- (p < 0.0001), this trend was absent from MCI and AD groups (Figure 2). Longitudinally, FRS high Aβ+ and FRS low Aβ+ groups consistently demonstrated lower memory scores compared to the FRS low Aβ- group. Interestingly, no significant difference in cognition was observed between FRS high Aβ- and FRS low Aβ- groups over time. However, the most pronounced divergence in longitudinal cognition of the four FRS and Amyloid groups was observed within the MCI diagnosis group (Figure 3). This study highlights the differential impact of cardiovascular risk on cognition depending on amyloid status and clinical diagnosis group. This underscores the importance of considering both cardiovascular risk factors and amyloid pathology early-on in understanding clinical manifestation and cognitive decline in the AD spectrum, particularly in ethnically diverse populations.
BACKGROUND:Limited research has explored the effect of cardiovascular risk and amyloid interplay on cognitive decline in East Asians. METHODS:Vascular burden was quantified using Framingham's General Cardiovascular Risk Score (FRS) in 526 Korean Brain Aging Study (KBASE) participants. Cognitive differences in groups stratified by FRS and amyloid positivity were assessed at baseline and longitudinally. RESULTS:Baseline analyses revealed that amyloid-negative (Aβ-) cognitively normal (CN) individuals with high FRS had lower cognition compared to Aβ- CN individuals with low FRS (p < 0.0001). Longitudinally, amyloid pathology predominantly drove cognitive decline, while FRS alone had negligible effects on cognition in CN and mild cognitive impairment (MCI) groups. CONCLUSION:Our findings indicate that managing vascular risk may be crucial in preserving cognition in Aβ- individuals early on and before the clinical manifestation of dementia. Within the CN and MCI groups, irrespective of FRS status, amyloid-positive individuals had worse cognitive performance than Aβ- individuals. HIGHLIGHTS:Vascular risk significantly affects cognition in amyloid-negative older Koreans. Amyloid-negative CN older adults with high vascular risk had lower baseline cognition. Amyloid pathology drives cognitive decline in CN and MCI, regardless of vascular risk. The study underscores the impact of vascular health on the AD disease spectrum.
This study presents a multi-faceted approach combining stereotactic biopsy with standard clinical open-craniotomy for sample collection, voxel-wise analysis of MR images, regression-based Generalized Additive Models (GAM), whole-exome sequencing. This work aims to demonstrate the potential of machine learning algorithms to predict variations in cellular molecular tumor characteristics. This retrospective study enrolled ten treatment-naive patients with radiologically confirmed glioma (5 WHO grade II, 5 WHO grade IV). Each patient underwent a multiparametric MR scan (T1W, T1W-CE, T2W, T2W-FLAIR, DWI) prior to surgery (27.9+/-34.0 days). During standard craniotomy procedure, at least 1 stereotactic biopsy was collected from each patient, with screenshots of the sample locations saved for spatial registration to pre-surgical MR data. Whole-exome sequencing was performed on flash-frozen tumor samples, prioritizing the signatures of five glioma-related genes: IDH1, TP53, EGFR, PIK3CA, NF1. Regression was implemented with a GAM using a univariate shape function for each predictor. Standard receiver operating characteristic analyses were used to evaluate detection, with AUC (area under curve) calculated for each gene target MR contrast combination. The mean AUC for the five gene targets 31 MR contrast combinations was 0.75+/-0.11; individual AUCs were as high as 0.96 for both IDH1 TP53 with T2W-FLAIR ADC 0.99 for EGFR with T2W ADC. An average AUC of 0.85 across the five mutations was achieved using the combination of T1W, T2W-FLAIR, ADC. These results suggest the possibility of predicting exome-wide mutation events from non-invasive, in vivo imaging by combining stereotactic localization of glioma samples a semi-parametric deep learning method. This approach holds potential for refining targeted therapy by better addressing the genomic heterogeneity of glioma tumors.