Electrocardiograms (ECGs) are among the most widely available clinical signals and play a central role in cardiovascular diagnosis. While recent foundation models (FMs) have shown promise for learning transferable ECG representations, most existing pretraining approaches treat leads as independent channels and fail to explicitly leverage their strong structural redundancy. We introduce the latent attention masked autoencoder (LAMAE) FM that directly exploits this structure by learning cross-lead connection mechanisms during self-supervised pretraining. Our approach models higher-order interactions across leads through latent attention, enabling permutation-invariant aggregation and adaptive weighting of lead-specific representations. We provide empirical evidence on the Mimic-IV-ECG database that leveraging the cross-lead connection constitutes an effective form of structural supervision, improving representation quality and transferability. Our method shows strong performance in predicting ICD-10 codes, outperforming independent-lead masked modeling and alignment-based baselines.
Echocardiography is a widely used modality for cardiac assessment due to its non-invasive and cost-effective nature, but the sparse and heterogeneous spatiotemporal views of the heart pose distinct challenges. Existing masked autoencoder (MAE) approaches typically process images or short clips independently, failing to capture the inherent multi-view structure required for coherent cardiac representation. We introduce Latent Attention Masked Autoencoder (LAMAE), a foundation model architecture tailored to the multi-view nature of medical imaging. LAMAE augments the standard MAE with a latent attention module that enables information exchange across frames and views directly in latent space. This allows the model to aggregate variable-length sequences and distinct views, reconstructing a holistic representation of cardiac function from partial observations. We pretrain LAMAE on MIMIC-IV-ECHO, a large-scale, uncurated dataset reflecting real-world clinical variability. To the best of our knowledge, we present the first results for predicting ICD-10 codes from MIMIC-IV-ECHO videos. Furthermore, we empirically demonstrate that representations learned from adult data transfer effectively to pediatric cohorts despite substantial anatomical differences. These results provide evidence that incorporating structural priors, such as multi-view attention, yields significantly more robust and transferable representations.
We introduce Concept Bottleneck Reward Models (CB-RM), a reward modeling framework that enables interpretable preference learning through selective concept annotation. Unlike standard RLHF methods that rely on opaque reward functions, CB-RM decomposes reward prediction into human-interpretable concepts. To make this framework efficient in low-supervision settings, we formalize an active learning strategy that dynamically acquires the most informative concept labels. We propose an acquisition function based on Expected Information Gain and show that it significantly accelerates concept learning without compromising preference accuracy. Evaluated on the UltraFeedback dataset, our method outperforms baselines in interpretability and sample efficiency, marking a step towards more transparent, auditable, and human-aligned reward models.
Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages the natural multi-view organization of radiology studies to learn view-invariant and disease-relevant representations. MVMAE combines masked image reconstruction with cross-view alignment, transforming clinical redundancy across projections into a powerful self-supervisory signal. We further extend this approach with MVMAE-V2T, which incorporates radiology reports as an auxiliary text-based learning signal to enhance semantic grounding while preserving fully vision-based inference. Evaluated on a downstream disease classification task on three large-scale public datasets, MIMIC-CXR, CheXpert, and PadChest, MVMAE consistently outperforms supervised and vision-language baselines. Furthermore, MVMAE-V2T provides additional gains, particularly in low-label regimes where structured textual supervision is most beneficial. Together, these results establish the importance of structural and textual supervision as complementary paths toward scalable, clinically grounded medical foundation models.
Synthetic data have emerged as an attractive option for developing machine-learning methods in human neuroimaging, particularly in magnetic resonance imaging (MRI)-a modality where image contrast depends enormously on acquisition hardware and parameters. This retrospective paper reviews a family of recently proposed methods, based on synthetic data, for generalizable machine learning in brain MRI analysis. Central to this framework is the concept of domain randomization, which involves training neural networks on a vastly diverse array of synthetically generated images with random contrast properties. This technique has enabled robust, adaptable models that are capable of handling diverse MRI contrasts, resolutions, and pathologies, while working out-of-the-box, without retraining. We have successfully applied this method to tasks such as whole-brain segmentation (SynthSeg), skull-stripping (SynthStrip), registration (SynthMorph, EasyReg), super-resolution, and MR contrast transfer (SynthSR). Beyond these applications, the paper discusses other possible use cases and future work in our methodology. Neural networks trained with synthetic data enable the analysis of clinical MRI, including large retrospective datasets, while greatly alleviating (and sometimes eliminating) the need for substantial labeled datasets, and offer enormous potential as robust tools to address various research goals.
We introduce ExpLIMEable for enhancing the understanding of Local Interpretable Model-Agnostic Explanations (LIME), with a focus on medical image analysis. LIME is a popular and widely used method in explainable artificial intelligence (XAI) that provides locally faithful and interpretable post-hoc explanations for black box models. However, LIME explanations are not always robust due to variations in perturbation techniques and the selection of interpretable functions. The proposed visual analytics application aims to address these concerns by enabling the users to freely explore and compare the explanations generated by different LIME parameter instances. The application utilizes a convolutional neural network (CNN) for brain MRI tumor classification and allows users to customize post-hoc LIME parameters to gain insights into the model's decision-making process. The developed application assists machine learning developers in understanding the limitations of LIME and its sensitivity to different parameters, as well as the doctors in providing an explanation to machine learning models, enabling more informed decision-making, with the ultimate goal of improving its robustness and explanation quality.
Background: Portable, low-field-strength (0.064-T) MRI has the potential to transform neuroimaging but is limited by low spatial resolution and low signal-to-noise ratio. Purpose: To implement a machine learning super-resolution algorithm that synthesizes higher spatial resolution images (1-mm isotropic) from lower resolution T1-weighted and T2-weighted portable brain MRI scans, making them amenable to automated quantitative morphometry.Materials and Methods: An external high-field-strength MRI data set (1-mm isotropic scans from the Open Access Series of Imaging Studies data set) and segmentations for 39 regions of interest (ROIs) in the brain were used to train a super-resolution convolu-tional neural network (CNN). Secondary analysis of an internal test set of 24 paired low-and high-field-strength clinical MRI scans in participants with neurologic symptoms was performed. These were part of a prospective observational study (August 2020 to December 2021) at Massachusetts General Hospital (exclusion criteria: inability to lay flat, body habitus preventing low-field -strength MRI, presence of MRI contraindications). Three well-established automated segmentation tools were applied to three sets of scans: high-field-strength (1.5-3 T, reference standard), low-field-strength (0.064 T), and synthetic high-field-strength images generated from the low-field-strength data with the CNN. Statistical significance of correlations was assessed with Student t tests. Correlation coefficients were compared with Steiger Z tests.Results: Eleven participants (mean age, 50 years +/- 14; seven men) had full cerebrum coverage in the images without motion arti-facts or large stroke lesion with distortion from mass effect. Direct segmentation of low-field-strength MRI yielded nonsignificant correlations with volumetric measurements from high field strength for most ROIs (P > .05). Correlations largely improved when segmenting the synthetic images: P values were less than .05 for all ROIs (eg, for the hippocampus [r = 0.85; P < .001], thalamus [r = 0.84; P = .001], and whole cerebrum [r = 0.92; P < .001]). Deviations from the model (z score maps) visually correlated with pathologic abnormalities.Conclusion: This work demonstrated proof-of-principle augmentation of portable MRI with a machine learning super-resolution algorithm, which yielded highly correlated brain morphometric measurements to real higher resolution images.(c) RSNA, 2022
The recent introduction of portable, low-field MRI (LF-MRI) into the clinical setting has the potential to transform neuroimaging. However, LF-MRI is limited by lower resolution and signal-to-noise ratio, leading to incomplete characterization of brain regions. To address this challenge, recent advances in machine learning facilitate the synthesis of higher resolution images derived from one or multiple lower resolution scans. Here, we report the extension of a machine learning super-resolution (SR) algorithm to synthesize 1 mm isotropic MPRAGE-like scans from LF-MRI T1-weighted and T2-weighted sequences. Our initial results on a paired dataset of LF and high-field (HF, 1.5T-3T) clinical scans show that: (i) application of available automated segmentation tools directly to LF-MRI images falters; but (ii) segmentation tools succeed when applied to SR images with high correlation to gold standard measurements from HF-MRI (e.g., r = 0.85 for hippocampal volume, r = 0.84 for the thalamus, r = 0.92 for the whole cerebrum). This work demonstrates proof-of-principle post-processing image enhancement from lower resolution LF-MRI sequences. These results lay the foundation for future work to enhance the detection of normal and abnormal image findings at LF and ultimately improve the diagnostic performance of LF-MRI. Our tools are publicly available on FreeSurfer (surfer.nmr.mgh.harvard.edu/).
Introduction: Flow cytometry is a laser-based technology used to quantify, sort, and analyze particles in a solution [1].This work presents a custom-made micro-flow cytometry platform to quantify Mycobacterium tuberculosis (Mtb), the agent that causes tuberculosis, using light-sheet fluorescence microscopy (LSFM).This project aims to quantify the number of living bacteria remaining in a solution after exposure to new potential compounds, as part of an effort to reduce the time burden of developing new antibiotics against Mtb. Methods:The light-sheet flow cytometer is composed of a microflow cytometry platform and a single-plane illumination microscopy (SPIM) system.This SPIM is characterized by two Galvo motors, which allow for 3D scanning of the sample and shadow reduction without moving the sample stage, avoiding unnecessary forces acting on the cytometry platform.Two electro-tunable lenses allow fine focusing from the software, improving the acquisition focus and 3D scanning control system by coordinating the laser movement and focal point.The configuration is shown in Figure 1.