Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning through iterative visual interaction. Medical VLMs often rely on static visual embeddings and single-pass inference, preventing models from re-examining, verifying, or refining visual evidence during reasoning. While tool-integrated reasoning offers a promising path forward, open-source VLMs lack the training infrastructure to learn effective tool selection, invocation, and coordination in multi-modal medical reasoning. We introduce MedVistaGym, a scalable and interactive training environment that incentivizes tool-integrated visual reasoning for medical image analysis. MedVistaGym equips VLMs to determine when and which tools to invoke, localize task-relevant image regions, and integrate single or multiple sub-image evidence into interleaved multimodal reasoning within a unified, executable interface for agentic training. Using MedVistaGym, we train MedVistaGym-R1 to interleave tool use with agentic reasoning through trajectory sampling and end-to-end reinforcement learning. Across six medical VQA benchmarks, MedVistaGym-R1-8B exceeds comparably sized tool-augmented baselines by 19.10% to 24.21%, demonstrating that structured agentic training--not tool access alone--unlocks effective tool-integrated reasoning for medical image analysis.
Spatially resolved transcriptomics (SRT) is a promising new technology that enables simultaneous analysis of gene expression and spatial information for biomedical research. However, the existing statistical and deep learning algorithms used for analyzing SRT data rely solely on two-dimensional (2D) spatial coordinates, which limits their ability to accurately identify spatial domains, spatially variable genes (SVGs), cell-to-cell communications, and developmental trajectories in a three-dimensional (3D) spatial manner. To address these limitations, we introduced Spa3D, which utilized the anti-leakage Fourier transform and graph convolutional neural network model to reconstruct 3D-based spatial structures from multiple 2D SRT slices. We demonstrate that Spa3D is appliable to analyze data from various SRT technology platforms and outperforms state-of-art methods by: (i) improving spatial domain identification through 3D reconstruction, (ii) elucidating cell-cell communication landscape in the 3D cellular organization, (iii) modeling of organ-level tempo-spatial development patterns in a 3D fashion, and (iv) annotating 3D spatial trajectory that are not captured by 2D spatial coordinates.
Extracting structured information from clinical notes requires navigating a dense web of interdependent variables where the value of one attribute logically constrains others. Existing Large Language Model (LLM)-based extraction pipelines often struggle to capture these dependencies, leading to clinically inconsistent outputs. We propose deep reflective reasoning, a large language model agent framework that iteratively self-critiques and revises structured outputs by checking consistency among variables, the input text, and retrieved domain knowledge, stopping when outputs converge. We extensively evaluate the proposed method in three diverse oncology applications: (1) On colorectal cancer synoptic reporting from gross descriptions (n=217), reflective reasoning improved average F1 across eight categorical synoptic variables from 0.828 to 0.911 and increased mean correct rate across four numeric variables from 0.806 to 0.895; (2) On Ewing sarcoma CD99 immunostaining pattern identification (n=200), the accuracy improved from 0.870 to 0.927; (3) On lung cancer tumor staging (n=100), tumor stage accuracy improved from 0.680 to 0.833 (pT: 0.842 -> 0.884; pN: 0.885 -> 0.948). The results demonstrate that deep reflective reasoning can systematically improve the reliability of LLM-based structured data extraction under interdependence constraints, enabling more consistent machine-operable clinical datasets and facilitating knowledge discovery with machine learning and data science towards digital health.
Advances in spatial transcriptomics (ST) technologies enable spatially resolved gene expression profiling, yet most platforms remain limited to spot-level resolution, where each measurement aggregates signals from multiple cells and obscures cell-specific programs and microenvironmental interactions. Existing computational approaches either provide limited sub-spot refinement or rely on matched single-cell RNA sequencing references that are costly and difficult to obtain. Here, we present DeSpaST (Deconvoluting Spatial Transcriptomics Signals to Cell-Level Resolution Using Histology Images), a dynamic edge-conditioned graph convolutional network that deconvolves spot-level ST data into cell-level gene expression profiles using only paired histology images. DeSpaST extracts nucleus-level morphological features, constructs directed cellular interaction graphs, and integrates spatial and transcriptional information through message passing. Validated across four cancer datasets with orthogonal Xenium, immunofluorescence, ablation, and interslice evaluations, DeSpaST enhances spatial resolution, facilitates downstream cellular-level analyses, and provides deeper insights into tissue biology.
Dotplot of Ezh2 expression across various samples profiled by the Mouse430_2 microarray platform
Graph Neural Networks (GNNs) provide a means for modeling inherently graphical data, such as transportation, social, and molecular networks, but also for enhancing and reducing unstructured data, including text and images. A potential shortcoming is that their predictions are opaque-a black box-hindering broader adoption and refinement. In this paper, we propose a novel architecture-agnostic algorithm, GNN-EGG (Graph Neural Network Explanations via Graph Generation), for GNN classifiers. As a model-level post-hoc explanation method, GNN-EGG can learn the data generating distribution for each class of graphs. The primary contribution of this work is the use of a differentiable approximation to Graph Edit Distance (GED) in the loss function. This term enables us to ensure consistency in both the graph space and the embedding space for our representative examples. It also reduces the random baseline issue where completely random graphs can still yield similar embeddings and strong predictions as reported in previous work. We benchmark our algorithm against the current state-of-theart models using the mutagenic molecules dataset (MUTAG) and apply our method to a large-scale GNN for malignancy detection in digital pathology tasks. The official implementation of this method can be found at this GitHub repository.
Advancements in multi-omics research have demonstrated the potential of integrating human microbiome and metabolomics data to better understand physiological processes and improve prediction accuracy in studies of human health. While conventional models utilizing single-omics data provide valuable perspectives, they often fail to capture the complexity of biological systems. Recent developments in supervised contrastive learning frameworks have enhanced predictive performance for categorical responses, yet limitations persist in extending these methods to continuous outcomes. A robust model capable of addressing these gaps could significantly enhance multi-omics predictions and provide new insights into complex biological interactions. Here, we present MB-SupCon-cont, a novel supervised contrastive learning framework designed for both categorical and continuous responses in multi-omics data. MB-SupCon-cont improves prediction accuracy by incorporating a generalized contrastive loss function that defines similarity and dissimilarity for continuous responses using three distance-based weighting methods. Through simulation studies and two real-world datasets for Type 2 Diabetes (T2D) and High-Fat Diet (HFD), we demonstrate that MB-SupCon-cont consistently achieves lower prediction errors than tuned conventional models, canonical correlation analysis, and autoencoder baselines, with most reaching statistical significance. We further provide a validation-based rule for selecting the weighting method and show that the learned embeddings align more closely with the response and recover known microbe and metabolite associations. The framework also provides superior representation learning and improves data visualization in lower-dimensional spaces. These findings suggest that MB-SupCon-cont is a powerful tool for general multi-omics prediction and may have broad applicability in biomedical research.
T cells are the central players in antitumor immunity, and effective tumor killing depends on their ability to infiltrate into the tumor microenvironment (TME) while maintaining normal cytotoxicity. However, late-stage tumors develop immunosuppressive mechanisms that impede T cell movement and induce exhaustion. Investigating T cell migration in human tumors in vivo could provide insights into tumor immune escape, although it remains a challenging task. In this study, we developed ReMiTT, a computational method that leverages spatial transcriptomics data to track T cell migration patterns within tumor tissue. Applying ReMiTT to multiple tumor samples, we identified potential migration trails. On these trails, chemokines that promote T cell trafficking displayed an increasing trend. Additionally, we identified key genes and pathways enriched on these migration trails, including those involved in cytoskeleton rearrangement, leukocyte chemotaxis, cell adhesion, leukocyte migration, and extracellular matrix remodeling. Furthermore, we characterized the phenotypes of T cells along these trails, showing that the migrating T cells are highly proliferative. Our findings introduce an approach for studying T cell migration and interactions within the TME, offering valuable insights into tumor-immune dynamics.