Gene regulatory networks (GRNs) underlie maintenance of cellular phenotypes and responses to stimuli. Modern single-cell profiling methods offer high-throughput datasets to infer GRNs en masse but do not capture dynamic information. Cell-state heterogeneity further confounds correlation-based inference approaches. We addressed these challenges with TwINFER, a conceptual framework leveraging information from recently divided sister cells, or “twins,” identifiable via recently developed barcoding techniques. We show that twin information discriminates regulatory from non-regulatory correlations and resolves interaction direction and type (activation/repression). We performed a diverse set of simulations, covering common network motifs and large-scale networks, where TwINFER outperforms state-of-the-art inference capabilities. Crucially, TwINFER resolved the commonly observed false positives in fan-out and feed-forward loop motifs where most methods perform poorly. Lastly, we applied TwINFER to a lineage-barcoded hematopoiesis dataset which refined the network inference, flagged multi-state genes, and determined causal relations. Our work exploits cellular twins as untapped information, readily complementing existing inference approaches.
Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross-modal supervision yields larger gains for volumetric models. Channel cross-attention and frequency-domain regularization are critical for leveraging 3D spatial context. On a protein–protein interaction task, MAE-3D achieves a ROC–AUC of 0.865, outperforming prior methods by up to +0.025. For protein localization, our best 3D model attains state-of-the-art AUC_micro (0.952) and F1_micro (0.742), improving over previous approaches by +0.003 and +0.010 absolute, respectively. Overall, these results demonstrate the advantages of native 3D modeling and multimodal alignment for representation learning in single-cell microscopy.
Mitochondrial dysfunction is implicated in a wide range of disorders, including cancer, neurodegeneration, and cardiovascular diseases. Conventional assays typically assess mitochondrial function by measuring bulk respiration rates across thousands of cells. While informative, these approaches cannot resolve the behavior of individual mitochondria, may overlook rare mitochondrial events, and often lack sensitivity to subtle changes. Fluorescence microscopy provides single-organelle resolution but is usually low throughput; for example, imaging 1000 cells can require from 20 minutes to an entire day depending on imaging mode. Additionally, analysis of fluorescence images frequently relies on manually selected thresholds, introducing potential bias and variability due to differences in signal-to-background ratios across cells. Here, we present a High-throughput Mitochondrial Imaging Platform (H-MIP) combined with a deep learning segmentation model, enabling rapid imaging and automated characterization of mitochondrial morphology in millions of mitochondria from tens of thousands of cells. We demonstrate the utility of H-MIP by assessing mitochondrial morphology following treatment with Mdivi-1, an inhibitor of mitochondrial division. Furthermore, we employ an in vitro disease model (vascular calcification) to show that mitochondria become elongated in vascular smooth muscle cells undergoing pathological calcification. Overall, H-MIP provides a scalable and robust tool for investigating mitochondrial structure, with the potential to accelerate therapeutic discovery targeting mitochondrial health. ### Competing Interest Statement The authors have declared no competing interest. Biotechnology and Biological Sciences Research Council, BB/J004316/1, BBS/E/D/20221657, BBS/E/RL/230001C, BB/Y513982/1 British Heart Foundation, RE/18/5/34216 Wellcome Trust, grant 210752 European Research Council, 866411 & 101113551 & 101213822 German Research Foundation, TRR-359
Out-of-Distribution (OOD) classification is a domain generalization task in computer vision. Deep learning models are typically developed and tested under the implicit assumption that training and test data are drawn independently and identically distributed (IID) from the same distribution. Overlooking OOD images can lead to poor performance under unseen or adverse viewing conditions, which are common in real-world scenarios. In this work, the proposed solution can be described as a data-driven approach to solve the OOD classification task in computer vision. The proposed approach consists of three stages, a training stage for exploiting labeled source data with different data augmentation strategies using powerful pretrained vision transformer models, an intermediate stage for weighted model ensemble and post-processing strategies, and finally an inference stage for exploiting unlabeled target data by using test-time learning. The proposed data-driven approach enhances the OOD generalization ability of deep models that withstand shifts in nuisances such as shape, pose, context, texture, occlusion, and weather in OOD or rare scenarios. Extensive data-augmentation strategies are used to improve the OOD generalization of deep models across various nuisances. The effectiveness of the proposed approach is evaluated using two standard computer vision benchmarks: ROBIN and a test set provided by the OOD-CV Challenge 2023. The experimental results show that the proposed approach demonstrates a performance improvement of 2.73% the ROBIN test set and achieves accuracy of 94.04% for the Challenge test set in terms of OOD robustness evaluation with classification accuracy. Furthermore, the proposed solution has secured a position within the top three OOD-based rankings on the OOD-CV Challenge Image Classification Leaderboard, 2023.
The dynamic environment of laboratories and clinics, with streams of data arriving on a daily basis, requires regular updates of trained machine learning models for consistent performance. Continual learning is supposed to help train models without catastrophic forgetting. However, state-of-the-art methods are ineffective for multiple instance learning (MIL), which is often used in single-cell-based hematologic disease diagnosis (e.g., leukemia detection). Here, we propose the first continual learning method tailored specifically to MIL. Our method is rehearsal-based over a selection of single instances from various bags. We use a combination of the instance attention score and distance from the bag mean and class mean vectors to carefully select which samples and instances to store in exemplary sets from previous tasks, preserving the diversity of the data. Using the real-world input of one month of data from a leukemia laboratory, we study the effectiveness of our approach in a class incremental scenario, comparing it to well-known continual learning methods. We show that our method considerably outperforms state-of-the-art methods, providing the first continual learning approach for MIL. This enables the adaptation of models to shifting data distributions over time, such as those caused by changes in disease occurrence or underlying genetic alterations.
Neural Cellular Automata (NCA) offer a robust and interpretable approach to image classification, making them a promising choice for microscopy image analysis. However, a performance gap remains between NCA and larger, more complex architectures. We address this challenge by integrating attention pooling with NCA to enhance feature extraction and improve classification accuracy. The attention pooling mechanism refines the focus on the most informative regions, leading to more accurate predictions. We evaluate our method on eight diverse microscopy image datasets and demonstrate that our approach significantly outperforms existing NCA methods while remaining parameter-efficient and explainable. Furthermore, we compare our method with traditional lightweight convolutional neural network and vision transformer architectures, showing improved performance while maintaining a significantly lower parameter count. Our results highlight the potential of NCA-based models an alternative for explainable image classification.
The detection and segmentation of white blood cells in blood smear images is a key step in medical diagnostics, supporting various downstream tasks such as automated blood cell counting, morphological analysis, cell classification, and disease diagnosis and monitoring. Training robust and accurate models requires large amounts of labeled data, which is both time-consuming and expensive to acquire. In this work, we propose a novel approach for weakly supervised segmentation using neural cellular automata (NCA-WSS). By leveraging the feature maps generated by NCA during classification, we can extract segmentation masks without the need for retraining with segmentation labels. We evaluate our method on three white blood cell microscopy datasets and demonstrate that NCA-WSS significantly outperforms existing weakly supervised approaches. Our work illustrates the potential of NCA for both classification and segmentation in a weakly supervised framework, providing a scalable and efficient solution for medical image analysis.
Distinguishing infiltrative basal cell carcinoma (BCC) from poorly differentiated cutaneous squamous cell carcinoma (cSCC) remains a significant histopathological challenge. Automated deep learning approaches hold promise for improving diagnostic reliability, yet robust external validation is essential. In this study, we developed a weakly supervised deep learning model to classify these diagnostically challenging subtypes and evaluated its generalizability across internal and external cohorts, as well as in comparison to a dermatopathology foundation model (HistoGPT). The model employed a multiple-instance learning framework (CLAM) using the histopathology-specific transformer Phikon for feature extraction from whole-slide images. Slide-level ground-truth diagnoses from the collected images (n = 335, University Hospital Erlangen) were derived from routine clinical practice and re-evaluated by two board-certified dermatopathologists. Performance was assessed on an internal test set of 84 whole-slide images (27 cSCC and 57 BCC) and two external datasets: Queensland cohort (n = 10, curated in-distribution cases) and the COBRA cohort (n = 200, broad, partly out-of-distribution cases). Model discrimination was quantified using ROC curves, while accuracy, sensitivity, and specificity were reported alongside 95% Wilson confidence intervals (CIs). On the internal test set, the model achieved perfect classification [area under the receiver operating characteristic (AUC) = 1.0; 100% accuracy, sensitivity, and specificity]. Similarly, strong performance was observed in the Queensland cohort (AUC = 1.0), although limited by sample size. In the more heterogeneous COBRA cohort, discrimination remained high (AUC = 0.923, 95% CI 0.885-0.961), requiring threshold adjustment to correct for marked calibration shift (balanced accuracy 86.5% at Youden's J). Attention heatmaps highlighted histologically meaningful regions. In zero-shot evaluation on the internal test set, HistoGPT achieved an overall accuracy of 77%, with high class-wise sensitivity for BCC (98%, 95% CI 91-100) but markedly reduced sensitivity for cSCC (33%, 95% CI 19-52). Fine-tuning a task-specific classifier on the HistoGPT backbone substantially improved performance, achieving near-perfect discrimination and 98% balanced accuracy. These findings demonstrate that weakly supervised deep learning enables highly accurate classification of diagnostically challenging BCC and cutaneous squamous cell carcinoma subtypes. However, reliable deployment across institutions necessitates careful calibration and domain adaptation, and even powerful foundation models such as HistoGPT benefit from targeted fine-tuning to ensure robust performance in dermatopathology.
Peripheral blood smears remain a cornerstone in the diagnosis of hematological neoplasms, offering rapid and valuable insights that inform subsequent diagnostic steps. However, since neoplastic transformations typically arise in the bone marrow, they may not manifest as detectable aberrations in peripheral blood, presenting a diagnostic challenge. In this paper, we introduce cAItomorph, an explainable transformer-based AI model, trained to classify hematological malignancies based on peripheral blood cytomorphology. Our data comprises peripheral blood single-cell images from 6115 patients with diagnoses confirmed by cytomorphology, cytogenetics, molecular genetics, and immunophenotyping from bone marrow samples, and 495 healthy controls, eight coarse classes. cAItomorph leverages the DinoBloom hematology foundation model and aggregates image encodings via a transformer-based architecture into a single vector. It achieves an overall accuracy of 0.72 in eight disease classification, with F1 scores of 0.76 for acute leukemia, 0.80 for myeloproliferative neoplasms and 0.94 for healthy cases. The overall accuracy increases to 0.87 in top-2 predictions. cAItomorph achieves high sensitivity for acute leukemia cases in external test sets. By analyzing attention heads, we demonstrate clinically relevant cell-level attentions in both internal and external test sets. Moreover, our model's calibrated prediction probabilities reduce the false discovery rate from 13.5% to 8.7% without missing any acute leukemia cases, thereby decreasing the number of unnecessary bone marrow aspirations based on peripheral blood smears. This study highlights the potential of AI-assisted diagnostics in hematological malignancies, illustrating how models trained on real-world data could enhance diagnostic accuracy and reduce invasive procedures.
The integration of multi-stain histopathology images through deep learning poses a significant challenge. Current approaches struggle with data heterogeneity and missing data, as concatenating multi-stain features may not effectively model stain-specific and cross-stain interactions. We introduce UNICORN (UNiversal stain Integration network for CORonary classificatioN), a two-stage, end-to-end trainable model comprising transformer self-attention blocks to process multi-stain histopathology for atherosclerosis severity prediction. The initial stage employs domain-specific expert models to extract features from each staining. An aggregation expert model then integrates features by learning their interactions. On a multi-class, multi-stain whole slide images (WSIs) dataset of atherosclerotic lesions from Munich Cardiovascular Studies Biobank (MISSION), UNICORN achieved a classification accuracy of 0.68, significantly outperforming state-of-the-art models. UNICORN identifies relevant tissue phenotypes across stainings and implicitly models disease progression. Its explainability and effectiveness in predicting atherosclerosis progression highlight the potential for broader applications in medical research and decision support.
Neural cellular automata (NCA) provide a lightweight alternative to encoder-decoder segmentation networks. However, it can be difficult to decide when a prediction should be trusted. Here, we study uncertainty estimation for NCA-based medical image segmentation without modifying the underlying architecture or retraining the model. Our approach is motivated by viewing the NCA as a dynamical system where convergent attractors correspond to confident predictions. Concretely, we propose resilience, a simple measure that leverages the intrinsic iterative structure of NCAs by probing the stability of the final prediction under small perturbations of the automaton state. Predictions that return to the same solution are deemed confident, while those that change substantially are flagged as uncertain. We evaluate uncertainty by its ability to predict segmentation quality using selective prediction metrics (ΔDice@90 and AURC) and ranking metrics (AUROC and AUPRC). Across multiple medical segmentation benchmarks, resilience identifies failure cases more reliably than baselines, improving trust and safety in NCA-based models.
Multiple instance learning (MIL) is a framework for weakly supervised classification, where labels are assigned to sets of instances, i.e., bags, rather than to individual data points. This paradigm has proven effective in tasks where fine-grained annotations are unavailable or costly to obtain. However, the effectiveness of MIL drops sharply when training data are scarce, such as for rare disease classification. To address this challenge, we propose incorporating topological inductive biases into the data representation space within the MIL framework. This bias introduces a topology-preserving constraint that encourages the instance encoder to maintain the topological structure of the instance distribution within each bag when mapping them to MIL latent space. As a result, our Topology Guided MIL (TG-MIL) method enhances the performance and generalizability of MIL classifiers across different aggregation functions, especially under scarce-data regimes. Our evaluations show average performance improvements of 15.3
Abstract Microscopic images of cells and tissues are central to disease diagnosis. In computational pathology, multiple instance learning (MIL) has emerged as a key paradigm for analyzing numerous images within a single patient sample. While the representative distribution of cells in a sample is important for diagnosis, existing MIL frameworks largely overlook it. We introduce TopoMIL, a framework that extracts the representative topological structure of the sample and integrates it into the MIL classifier. Three topological representations are assessed, each with distinct advantages and computational costs. We evaluate TopoMIL on four histopathology and cytomorphology datasets, each presenting unique challenges. Integrating the sample’s topological information into MIL enhances classification across average, max, attention-based, and transformer pooling, yielding AUCROC gains of 3.3%, 4.2%, 5.9%, and 0.5%, respectively, with moderate computational cost. Our work underscores the potential of TopoMIL as a scalable extension to existing morphology-based models in computational pathology.
Data scarcity is a major bottleneck in medical Multiple Instance Learning (MIL), especially for rare diseases or expensive modalities. We introduce a statistically grounded patient augmentation approach that generates realistic patients directly in embedding space. Using Gaussian Mixture Models as a probabilistic clustering approach on pooled instance embeddings from all patients, our method learns disease-specific "recipes"-statistical distributions of instances across unsupervised clusters. New patients are then generated by sampling embeddings from clusters based on learned recipes. Unlike existing methods that require examples from all categories, our method can generate patients offline by re-mixing pooled embeddings. Generated patients are further selected based on uncertainty quantification to improve MIL performance. We evaluate our method across three clinically relevant scarcity scenarios: (i) cross-dataset transfer, where an entirely missing "healthy" class is generated using statistics from an external cohort; (ii) low-data regimes, where class sizes are extremely limited; and (iii) small-cohort non-image tasks, including single-cell RNA-seq and flow cytometry. Across all experiments, our method improves performance over baseline, often outperforming other bag-mixing strategies. Notably, in the missing-class scenario, a performance comparable to full-dataset training is achieved, demonstrating its potential for rare disease diagnostic and privacy-preserving patient augmentation. The code is available at https://github.com/marrlab/RECIPE
Oncogenes such as KRAS display marked tissue specificity in their oncogenic potential, genetic interactions and phenotypic effects, but the underlying determinants remain largely unresolved1-5. Here, to address these questions, we developed the Mouse Cancer Cell line Atlas, a broad-utility resource of 590 comprehensively characterized models across a wide range of entities ( www.mcca.tum.de ). Comparative and functional studies using this platform, human cohorts and mice identified core principles underlying tissue-specific evolution of KRAS-initiated cancers. First, we show that mutant KRAS dosage gain through allelic imbalance exerts cell-type-specific effects, defining its timing across entities, as exemplified by dosage-sensitive developmental reprogramming during pancreatic cancer initiation. Second, we highlight how tissue- and stage-specific evolutionary requirements, such as block of differentiation in the intestine, select for KRAS-collaborating alterations. Third, we identified context-dependent epistatic KRAS-tumour suppressor interactions and show that reciprocal dosage sensitivities dictate the entity-specific patterns of cancer gene alterations, explaining their frequency, zygosity and acquisition chronology. These findings highlight how intrinsic and acquired determinants instruct cancer evolution in different tissues, with predictable molecular patterns, temporal dynamics and phenotypic outcomes. Our study provides major advances towards a mechanistic understanding of cancer genomes.
Multimodal alignment of histopathology encoders with transcriptomic and genomic data has been shown to significantly improve performance in downstream diagnostic tasks. Hematological cytology is unique in that visual single-cell evaluation is often paired with cytogenetics and molecular genetics for blood cancer diagnosis. In this study, we present a framework to align single white blood cell images with chromosomal aberrations (karyotype) and somatic mutations from targeted gene panels. Our training strategy follows a two-stage approach: (i) self-supervised, vision-only pretraining of a transformer aggregator using an iBOT head on a cohort of over 1500 patients, and (ii) genetic alignment via supervised contrastive loss on acute myeloid leukemia patients. Our genetically aligned patient encoder improves hematological diagnostic tasks, outperforming slide-level histopathology foundation models. Additionally, the model provides off-the-shelf retrieval capabilities for diseases and genetic alterations. Incorporating genetic data into patient encoders increases the quality of patient representations, providing a framework that aligns with clinical diagnostic workflows and paves the way for future multimodal hematology-specific AI. The code and model weights are available at https://github.com/marrlab/GenBloom.
Abstract Quantitative fluorescence microscopy is frequently confounded by spatially varying illumination and temporal intensity drift. Although BaSiC is a widely adopted retrospective correction method, it can fail when foreground content is strongly correlated across images—a common regime in time-lapse, tiled and volumetric acquisitions—and its application often requires manual parameter tuning that limits reproducibility and scalability. We introduce BaSiCPy, a foreground-aware implementation of BaSiC that improves illumination profile estimation under correlated foreground structures, provides automatic hyperparameter selection and accelerates large-scale processing through GPU support. BaSiCPy is distributed as an open-source Python package with graphical and programmatic interfaces, facilitating integration into contemporary bioimage analysis workflows.
Hierarchical structure is common in image data, where fine-grained clusters often merge into larger, coarser semantic groups. In biological cell images, current self-supervised learning models often suppress this hierarchy, as coarse factors such as imaging modality can obscure finer morphological attributes in the latent space. We propose a hierarchy-aware self-supervised training framework to address this problem. Our method combines two components: a distillation framework with a segmentation teacher to improve morphological awareness in the latent space, and a hierarchy-aware contrastive loss based on HDBSCAN to improve decision boundaries between closely related subtypes at different hierarchical levels. Together, these components reduce the tendency of self-supervised learning to overemphasize coarse factors and instead align embeddings with semantic and morphological cues. This yields biologically meaningful sub-clusters driven by fine morphological detail. We train and evaluate our method on a curated corpus of 2.3 million single cells aggregated from 20 microscopy datasets, both labeled and unlabeled, covering 208 cell classes. Our method improves over baseline and counterpart methods, increasing average top-K accuracy by 2.8
The microscopic observation of blood cells is a crucial step in diagnosing pathologies such as leukemia. DINOv2 models have been employed to extract features from blood cell images, but they do not include biological knowledge, nor do they allow multi-granular labels. To enhance the representation of these cells, we propose leveraging a biologically informed hierarchy of white blood cell types. We train a DINOv2-based foundation model with a semi-supervised framework that uses hierarchical supervision. It enables using datasets with varying levels of label precision within a structure that represents the process of cell differentiation. To support multi-level label precision, we modify the original hierarchical loss function, allowing any hierarchy level to serve as a ground truth class. We evaluate our model on three external datasets, including an out-of-domain set of cervical cells. Our approach improves generalization of the model to new datasets, improving by 1 percentage point the balanced accuracy on the two blood cell external datasets, and by 2.5 percentage point the balanced accuracy on the out-of-domain dataset. In addition the proposed strategy better aligns the model’s latent space with biological properties, leading to more acceptable misclassifications
Pathology foundation models (FMs) have driven significant progress in computational pathology. However, these high-performing models can easily exceed a billion parameters and produce high-dimensional embeddings, thus limiting their applicability for research or clinical use when computing resources are tight. Here, we introduce Pathryoshka, a multi-teacher distillation framework inspired by RADIO distillation and Matryoshka Representation Learning to reduce pathology FM sizes while allowing for adaptable embedding dimensions. We evaluate our framework with a distilled model on ten public pathology benchmarks with varying downstream tasks. Compared to its much larger teachers, Pathryoshka reduces the model size by 86-92