Disease heterogeneity and commonality pose critical challenges to precision medicine, as traditional approaches frequently focus on single disease entities and overlook shared mechanisms across conditions. Here, inspired by pan-cancer and multi-organ research, we introduce the concept of ‘pan-disease’ to investigate the heterogeneity and shared etiology in brain, eye and heart diseases. Leveraging individual-level data from 129,340 participants and summary-level data, curated from the MULTI consortium, we applied a weakly supervised deep learning model (Surreal-GAN) to multi-organ imaging, genetic and proteomic data, identifying 11 artificial intelligence (AI)-derived biomarkers, called multi-organ AI endophenotypes, for the brain (Brain 1–6), eye (Eye 1–3) and heart (Heart 1–2). We found Brain 3 to be a risk factor for Alzheimer’s disease progression and mortality, whereas Brain 5 was protective against Alzheimer’s disease progression. In data from an anti-amyloid Alzheimer’s disease drug (solanezumab), heterogeneity in cognitive decline trajectories was observed across treatment groups. At week 240, patients with lower Brain 1–3 expression had slower cognitive decline, whereas patients with higher expression had faster cognitive decline. A multilayer causal pathway pinpointed Brain 1 as a mediational endophenotype linking the FLRT2 protein to migraine, exemplifying new therapeutic targets and pathways. In addition, genes associated with Eye 1 and Eye 3 were enriched in cancer drug-related gene sets with causal links to specific cancer types and proteins. Finally, Heart 1 and Heart 2 had the highest mortality risk and unique medication history profiles, with Heart 1 showing favorable responses to antihypertensive medications and Heart 2 to digoxin treatment. The 11 multi-organ AI endophenotypes provide new AI dimensional representations for precision medicine and highlight the potential of AI-driven patient stratification for disease risk monitoring, clinical trials and drug discovery. Disease heterogeneity complicates precision medicine, which focuses on single conditions and ignores shared mechanisms. Here the authors introduce ‘pan-disease’ analysis using a deep learning model on multi-organ data, identifying 11 AI-derived biomarkers that reveal new therapeutic targets and pathways, enhancing patient stratification for disease risk monitoring and drug discovery.
Deep neural networks can classify ECGs with high accuracy when training data is abundant. Rare conditions like Brugada syndrome, an inherited arrhythmia syndrome predisposing to sudden death, pose challenges due to data scarcity hindering model training. We evaluated multiple machine learning (ML) approaches to optimise a Brugada ECG classification model using limited training data. The baseline model was trained on a dataset comprising 176 Brugada, 176 right bundle branch block (RBBB) and 352 normal ECGs from Zhongshan Hospital (Zhongshan-baseline dataset), framed as a binary classification task to distinguish Brugada from non-Brugada ECGs. A 25%-75% train-test split was used to exacerbate data scarcity. To enhance training, we incorporated three additional datasets: (i) a different, labelled ECG dataset from Zhongshan Hospital including normal and RBBB ECGs (Zhongshan-pretrain), (ii) an unlabelled ECG dataset from Hammersmith Hospital including Brugada and non-Brugada ECGs (Imperial), (iii) an open-access labelled ECG dataset (PTB-XL). Three strategies were tested: (1) supervised pretraining, (2) self-supervised pretraining with data augmentation, and (3) oversampling using SMOTE (synthetic minority oversampling technique). Each model was evaluated on the unseen internal test set and an external Brugada mimic dataset. The models were re-trained using an 80%-20% train-test split as a secondary analysis. The baseline model achieved 92.2% accuracy, F1-score 0.837, and area under the Receiver Operating Characteristic curve (AUC) 0.962. Supervised pretraining significantly improved performance when training data was scarce, with the best model pretrained on the Zhongshan-pretrain dataset boosting accuracy (+3.2%), F1-score (+0.071) and AUC + 0.019), with consistent cross-validation performance. Self-supervised pretraining produced smaller and more variable gains, although select models better mitigated against false positives on the Brugada mimic dataset. SMOTE oversampling showed inconsistent effects on performance. Incorporating pretraining and oversampling may facilitate the development of more accurate AI-ECG models for rare diseases when training data is limited but provides diminishing returns when adequate labelled data is available.
Surface registration plays an important role for anatomical shape analysis in medical imaging. Existing surface registration methods often face a trade-off between efficiency and robustness. Local point matching methods are computationally efficient, but vulnerable to noise and initialisation. Methods designed for global point set alignment tend to incur a high computational cost. To address the challenge, here we present a fast surface registration method, which formulates surface meshes as probability measures and surface registration as a distributional optimisation problem. The discrepancy between two meshes is measured using an efficient sliced Wasserstein distance with log-linear computational complexity. We propose a novel optimisation method, AdamFlow, which generalises the well-known Adam optimisation method from the Euclidean space to the probability space for minimising the sliced Wasserstein distance. We theoretically analyse the asymptotic convergence of AdamFlow and empirically demonstrate its superior performance in both affine and non-rigid surface registration across various anatomical structures.
Cardiac imaging enables quantitative assessment of cardiac structure and function but remains constrained by cost, infrastructure and specialist expertise. In contrast, electrocardiogram (ECG) is widely accessible yet underexploited, despite encoding latent information about cardiac physiology. Here we introduce visionECG, a conditional flow matching framework that learns a probabilistic mapping between two biological distributions - the space of cardiac electrical signals and the space of cardiac geometries. Using 71,132 paired ECG and cardiac mesh sequence datasets from the UK Biobank, with external assessment in 5,000 patients with ECG-echocardiogram pairs, the model reconstructs quantitatively accurate spatiotemporal representations of the left ventricle using ECG inputs and basic demographic information alone. These reconstructions enable discrimination of structural abnormalities and disease labels, provide visualisations of functional abnormalities, and support flexible quantification of both global and regional parameters. By reframing the ECG as a generative source of patient-specific left ventricular geometry and motion, this work establishes a scalable framework for translating low-dimensional signals into high-dimensional, physiologically grounded structured representations.
Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and interpreted in natural language. However, the field of medical MLLMs is constrained by non-trivial challenges, notably the scarcity of high-quality training data and the frequent occurrence of missing data in the real-world clinical setting. Here, we propose a novel unified multimodal model, UniBrain, for brain magnetic resonance image (MRI) analysis. To address potential missing brain MRI modalities, we employ a unified training strategy to perform joint imaging modality imputation and brain image understanding. During training, an interleaved and description-enriched data flow is constructed to train the model in an autoregressive manner, enabling medical reasoning with generated multimodal data. A self-alignment strategy is introduced to leverage dense image embeddings to learn fine-grained anatomical features without requiring detailed image captions. Furthermore, we propose a dynamic hidden state mechanism to alleviate the exposure bias during long-context multimodal inference. Extensive experiments on multi-disease brain MRI dataset demonstrate that UniBrain achieves high performance for brain image imputation, understanding, and disease diagnosis under various extents of modality incompleteness.
Background Long-term electrocardiogram (ECG) monitoring with wearable devices enables large-scale characterisation of cardiac rhythms, but population-based evidence remains limited. The UK Biobank Cardiac Monitoring Study integrates 14-day patch-based ECG monitoring with accelerometry and detailed phenotypic and lifestyle data. Here, we report the acquisition protocol, data processing, and initial findings from 27,658 participants. Methods Participants in the UK Biobank imaging study were invited to undergo 14-day cardiac monitoring using a Zio XT (pilot phase; 2015-18) or BodyGuardian MINI (main phase; 2019-ongoing) monitor. ECGs were analysed by certified technicians and automated algorithms to identify atrial, ventricular, and conduction arrhythmias. In parallel, beat-to-beat RR intervals were derived using in-house algorithms, and physical activity from calibrated triaxial accelerometer data. Analyses assessed wear time, arrhythmia prevalence, circadian patterns, and repeatability. Findings In total, 27,658 participants (mean age 71 years; 49.9% women) were analysed, including 7,795 from the pilot phase and 21,141 from the main phase; 1,353 (4.9%) had repeat recordings. In the main phase, median wear time was 13.2 days (IQR 11.9-13.9), and undiagnosed atrial fibrillation (AF) was detected more frequently in men than women (3.2% vs 1.7%; p<0.001); 68% was paroxysmal, with 27.4% detected during week two. Ventricular tachycardia occurred in 12.1% (8.4% in women), with sustained episodes rare (0.4%) but observed. Arrhythmia timing varied markedly with activity, with AF peaking during nocturnal inactivity and ventricular ectopy increasing during activity, peaking at midday. Repeat assessments showed strong reproducibility of diurnal heart rate and activity profiles, with more modest arrhythmia consistency. Interpretation Extended ECG monitoring enables detection of subclinical arrhythmias and long-term physiological rhythms in older adults. Linkage to imaging, multi-omics, and clinical outcomes in UK Biobank will enable unprecedented evaluation of the natural history of asymptomatic rhythm disturbances and their impact on brain health. Funding British Heart Foundation and Wellcome Trust. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement British Heart Foundation [RG/18/6/33576] British Heart Foundation Center of Research Excellence [RE/18/3/34214] Wellcome Trust [223100/Z/21/Z] ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study was approved by the North-West Research Ethics Committee (06/MRE08/65). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study will be made available in the future upon request to UK Biobank.
Clinical abnormality grounding for rare diseases is often hindered by data scarcity, rendering supervised fine-tuning infeasible and single-pass inference highly unstable. Thus, we propose Dynamic Decision Learning (DDL), a framework that enables frozen LVLMs to refine their decisions across language and visual spaces by optimizing instructions and consolidating predictions under visual perturbations, thereby improving localization quality and producing a consensus‑based reliability score that quantifies the model’s confidence. Results on brain‑imaging benchmarks, including a rare‑disease dataset with 281 pathology types across 3B-72B models, show that DDL improves mAP@75 by up to 105\% on rare‑disease cases and surpasses adaptation baselines and supervised fine‑tuning. Moreover, we show that DDL yields stronger calibration between consensus‑based reliability scores and localization accuracy under severe distribution shifts and increasing task difficulty. The code will be open-sourced.
BACKGROUND:Segmentation of lower limb bones is essential for accurate preoperative planning in total knee arthroplasty (TKA), yet fully annotated CT datasets are rarely available in clinical practice. This study evaluated whether a partially supervised deep learning framework can leverage incompletely annotated CT data to generate anatomically accurate femur and tibia segmentations suitable for TKA planning. METHOD:A 3D nnU-Net model was trained using 205 healthy full-leg CT scans with mixed annotation completeness, including 17 fully annotated cases and partially labelled femur or tibia in the remaining scans. Performance was evaluated on an internal healthy dataset (n = 40), a cadaveric dataset (n = 15), and an osteoarthritis (OA) dataset acquired for robotic TKA planning (n = 10). Accuracy was assessed using Dice similarity coefficient (DSC), Hausdorff distance (HD), HD95, and root-mean-square surface distance (RMSE). Clinical relevance was evaluated using landmark localisation errors and joint alignment measurements (mLDFA and mPTA). RESULTS:On the cadaveric dataset, mean DSC values were 96.53% (femur) and 97.41% (tibia), with RMSE < 1 mm. On the OA dataset, mean DSC remained approximately 96.5% with HD95 < 1.7 mm across acquisition windows. Alignment measurements derived from automatic and manual segmentations showed small differences (mLDFA: 86.81° vs 87.02°, p = 0.11; mPTA: 86.74° vs 87.29°, p = 0.08), with no statistically significant differences observed. CONCLUSION:Partially supervised training enables accurate lower limb segmentation from incompletely annotated CT datasets and preserves clinically relevant alignment measurements required for TKA planning.
Spatio-temporal (3D+t) generative modelling of cardiac shape and motion is crucial for understanding heart structure and function at population scale. Existing generative models for cardiac shape synthesis either adopt volumetric shape representations that lack anatomical correspondence across different time points and subjects, or rely on VAE-based frameworks that suffer from a trade-off between reconstruction fidelity and generative diversity. In this work, we propose Cardiac Mesh Flow, a novel generative flow model for 3D+t cardiac four-chamber mesh generation with anatomical correspondence, temporal coherence, and periodic consistency. Leveraging the flow matching technique, Cardiac Mesh Flow performs efficient one-step generation of multi-scale free-form deformation fields, which warp a template mesh to generate cardiac four-chamber meshes across a cardiac cycle. Furthermore, Cardiac Mesh Flow enables controllable generation conditioned on cardiac chamber volumes, allowing precise control of the synthetic heart. Experimental results demonstrate that Cardiac Mesh Flow achieves high fidelity and diversity on both unconditional and conditional generation, compared to state-of-the-art 3D+t cardiac mesh generation methods.
The human heart is a sophisticated system composed of four cardiac chambers with distinct shapes, which function in a coordinated manner. Existing shape models of the heart mainly focus on the ventricular chambers and they are derived from relatively small datasets. Here, we present a spatio-temporal (3D+t) statistical shape model of all four cardiac chambers, learnt from a large population of nearly 100,000 participants from the UK Biobank. A deep learning-based pipeline is developed to reconstruct 3D+t four-chamber meshes from the cardiac magnetic resonance images of the UK Biobank imaging population. Based on the reconstructed meshes, a 3D+t statistical shape model is learnt to characterise the shape variations and motion patterns of the four cardiac chambers. We reveal the associations of the four-chamber shape model with demographics, anthropometrics, cardiovascular risk factors, and cardiac diseases. Compared to conventional image-derived phenotypes, we validate that the four-chamber shape-derived phenotypes significantly enhance the performance in downstream tasks, including cardiovascular disease classification and heart age prediction. Furthermore, we demonstrate the effectiveness of shape-derived phenotypes in novel applications such as heart shape retrieval and heart re-identification from longitudinal data. To facilitate future research, we will release the learning-based mesh reconstruction pipeline, the four-chamber cardiac shape model, and return all derived four-chamber meshes to the UK Biobank.
Optimal sleep has a vital role in promoting healthy ageing and enhancing longevity. Here we propose Sleep Chart to assess the relationship between self-reported sleep duration and 23 biological ageing clocks derived from in vivo imaging1, plasma proteomics2 and metabolomics3. First, a systemic, U-shaped pattern emerges between sleep duration and biological age gaps across nine brain and body systems and three omics technologies. The sample-specific lowest biological age gaps are achieved between 6.4 and 7.8 h of sleep duration, varying by organ and sex in the UK Biobank (aged 37-84 years). Furthermore, short (<6 h) and long (>8 h) sleep duration, compared with a normal sleep duration (6-8 h), are associated with increased risk of systemic diseases beyond the brain and all-cause mortality, with evidence from genetic correlations and time-to-incident survival predictions, such as depression and diabetes. Finally, the pathways by which long and short sleep duration are associated with late-life depression differ: ageing clocks may partially mediate the pathway for long sleep duration, while short sleep duration shows a more direct link. Although Mendelian randomization does not provide strong evidence that disease causally affects sleep, it cannot completely exclude such reverse causality. Our findings suggest a cross-organ, multi-omics U-shaped relationship between sleep duration and biological ageing clocks, highlighting the potential of sleep optimization to promote healthy ageing, lower disease risk and extend longevity.
BACKGROUND:Cardiac magnetic resonance imaging is central to cardiovascular diagnosis and management, yet extracting key clinical measurements remains time-consuming, subjective, and of limited reproducibility. Current deep learning methods often require a separate model trained from scratch for each task, and generating sufficient labelled training data demands substantial clinical expertise. METHODS:We developed CineMA, a multi-view conv-transformer masked autoencoder foundation model, pre-trained on 15 million cine cardiac magnetic resonance images from 74,916 studies. The model was fine-tuned and evaluated on eight independent datasets for segmentation, landmark localisation, disease diagnosis, and prognostication, representing the largest such benchmark to date. Performance was compared against convolutional neural network baselines, including nnUNet. RESULTS:Here we show, without dataset-specific hyperparameter tuning, CineMA approaches nnUNet performance in ventricle segmentation and ejection fraction estimation while achieving higher consistency across repeated scans. CineMA surpasses convolutional baselines in cardiovascular disease detection with notably improved specificity, and matches their performance in long-axis function measurement. Beyond cardiac diseases, CineMA shows potential for predicting systemic conditions and survival outcomes, with comparable performance across demographic subgroups. CONCLUSIONS:CineMA demonstrates accuracy, learning efficiency, adaptability, and fairness across diverse cardiac image analysis tasks, offering a strong alternative to task-specific model training for automated cardiac image analysis.
Abstract Characterisation of the motion dynamics of the left ventricle is key to understanding pathophysiological mechanisms and transitions from health to disease. Conventional volumetric assessments of the heart using imaging represent mainly aggregate global features of function that are poorly discriminating. Here we present a novel approach to quantify and visualise how the left ventricle is affected by cardiovascular risk factors through efficient representations of motion trajectories. We use computer vision to survey four-dimensional cardiac motion traits using densely sampled point clouds of the left ventricle in over 20,000 participants of UK Biobank. We developed a computational framework for dimensionality reduction of spatiotemporal information to derive a human-interpretable signature summarising variation in complex patterns of motion. We found six phenogroups representing a novel classification of heterogeneous motion phenotypes with differential enrichment of cardiovascular outcomes and genetic risk. Low dimensional representations of motion are visualised as a simple spatial signature capturing deviation from an average state. Discovering compact cardiac motion signatures of health and disease from dynamic point clouds enables efficient classification of patient risk and predisposing polygenic factors.
Magnetic resonance imaging (MRI) is a powerful and versatile imaging technique, offering a wide spectrum of information about the anatomy by employing different acquisition modalities. However, in the clinical workflow, it is impractical to collect all relevant modalities due to the scan time and cost constraints. Virtual full-stack scanning aims to impute missing MRI modalities from available but incomplete acquisitions, offering a cost-efficient solution to enhance data completeness and clinical usability. Existing imputation methods often depend on global conditioning or modality-specific designs, which limit their generalisability across patient cohorts and imaging protocols. To address these limitations, we propose CodeBrain, a unified framework that reformulates various ``any-to-any'' imputation tasks as a region-level full-stack code prediction problem. CodeBrain adopts a two-stage pipeline: (1) it learns the compact representation of a complete MRI modality set by encoding it into scalar-quantised codes at the region level, enabling high-fidelity image reconstruction after decoding these codes along with modality-agnostic common features; (2) it trains a projection encoder to predict the full-stack code map from incomplete modalities via a grading-based design for diverse imputation scenarios. Extensive experiments on two public brain MRI datasets, i.e., IXI and BraTS 2023, demonstrate that CodeBrain consistently outperforms state-of-the-art methods, establishing a new benchmark for unified brain MRI imputation and enabling virtual full-stack scanning.Code will be released.
Cardiovascular health is vital to human well-being, and cardiac magnetic resonance (CMR) imaging is considered the clinical reference standard for diagnosing cardiovascular disease. However, its adoption is hindered by long scan times, complex contrasts, and inconsistent quality. While deep learning methods perform well on specific CMR imaging sequences, they often fail to generalize across modalities and sampling schemes. The lack of benchmarks for high-quality, fast CMR image reconstruction further limits technology comparison and adoption. The CMRxRecon2024 challenge, attracting over 200 teams from 18 countries, addressed these issues with two tasks: generalization to unseen modalities and robustness to diverse undersampling patterns. We introduced the largest public multi-modality CMR raw dataset, an open benchmarking platform, and shared code. Analysis of the best-performing solutions revealed that prompt-based adaptation and enhanced physics-driven consistency enabled strong cross-scenario performance. These findings establish principles for generalizable reconstruction models and advance clinically translatable AI in cardiovascular imaging.
Identifying robust associations between cardiac imaging phenotypes and clinical diseases is fundamental to population-scale cardiovascular research and reliable risk stratification. However, current phenome-wide association studies rely on pre-defined, single-variable phenotypes or expert-crafted features, which limits their ability to capture clinically meaningful non-linear effects and cross-phenotype interactions. To address this, we propose CPAgents, an iterative phenotype-Composition framework for cardiovascular Phenome-wide association study (PheWAS) that automatically constructs and validates interpretable composite phenotypes (e.g., polynomial, ratio, and interaction forms) from base imaging features. Specifically, our system coordinates three agents: (i) an Analyst that identifies statistical pathologies and nominates candidate transformations; (ii) a Proposer that generates constrained, medically and statistically motivated expressions under numerical safety rules; and (iii) a Verifier that evaluates candidates using multi-stage criteria and produces transparent evidence trails for accepted phenotypes. Evaluated on a population-scale cardiac imaging cohort, the discovered composite phenotypes markedly improve disease discrimination: across 72 classifier-disease-metric combinations, our variants achieve the top rank in 56 cases versus 18 for baselines, with gains observed across all nine clinical disease categories. Our framework yields compact, clinically interpretable phenotype formulas with transparent evidence trails, enabling scalable discovery of stronger phenotype-disease associations beyond expert-driven feature selection.
In this work, we address the problem of grounding abnormalities in medical images, where the goal is to localize clinical findings based on textual descriptions. While generalist Vision-Language Models (VLMs) excel in natural grounding tasks, they often struggle in the medical domain due to rare, compositional, and domain-specific terms that are poorly aligned with visual patterns. Specialized medical VLMs address this challenge via large-scale domain pretraining, but at the cost of substantial annotation and computational resources. To overcome these limitations, we propose \textbf{Knowledge to Sight (K2Sight)}, a framework that introduces structured semantic supervision by decomposing clinical concepts into interpretable visual attributes, such as shape, density, and anatomical location. These attributes are distilled from domain ontologies and encoded into concise instruction-style prompts, which guide region-text alignment during training. Unlike conventional report-level supervision, our approach explicitly bridges domain knowledge and spatial structure, enabling data-efficient training of compact models. We train compact models with 0.23B and 2B parameters using only 1.5\% of the data required by state-of-the-art medical VLMs. Despite their small size and limited training data, these models achieve performance on par with or better than 7B+ medical VLMs, with up to 9.82\% improvement in $mAP_{50}$. Code and models: \href{https://lijunrio.github.io/K2Sight/}{\textcolor{SOTAPink}{https://lijunrio.github.io/K2Sight/}}.