Motion artifacts in magnetic resonance imaging (MRI) degrade diagnostic reliability. Existing deep learning methods are typically contrast-specific and fail to generalize across diverse modalities and artifact severities. We propose a unified framework combining parameter-informed contrast disentanglement with severity-aware adaptive correction. ScanCLIP, pretrained on over 30,000 MRI text-image pairs, derives contrast embeddings from acquisition parameters to disentangle contrast style from anatomical content, yielding contrast-free features. A Vision Transformer then estimates motion severity and routes features through a Mixture-of-Experts network, enabling targeted artifact correction. A dual-pathway decoder reconstructs both the clean image and residual artifact map, enforcing image-space consistency. On IXI and HCP benchmarks, our method improves PSNR by 0.75 dB and SSIM by up to 0.0279 over state-of-the-art approaches, with larger gains at higher artifact severities. It further demonstrates robust zero-shot generalization on real-world clinical data acquired with unseen scanning parameters, where existing methods either fail to remove artifacts or introduce additional distortions.
Magnetic resonance imaging (MRI) is an indispensable tool for clinical knee examination, which often scans 2D stacked slices from multiple views. Radiologists typically locate lesion regions in one view, and then refer to other views to formulate a comprehensive diagnosis. However, existing computer-aided diagnosis methods fall short of identifying and fusing local regions in multi-view scans, leading to a decline in diagnostic performance and a heavy reliance on extensively annotated data. This paper introduces a novel framework that represents multi-view MRI scans as a knee graph, and conducts diagnosis using the proposed Knee Graph Network (KGNet). Moreover, KGNet is greatly enhanced by multi-task pre-training, which requires KGNet to reconstruct masked knee local patches and segment unmasked ones working alongside corresponding decoders. Experimental evaluations on public and in-house clinical datasets confirm that our framework outperforms existing approaches in diagnosing cartilage defects, anterior cruciate ligament tears, and knee abnormalities. In conclusion, our framework demonstrates the potential of enhancing knee disease diagnosis by representing multi-view MRI scans as a graph and employing multi-task pre-training in the graph network. The code is publicly available at https://github.com/zixuzhuang/KGNet.
Cervical abnormality screening is pivotal for prevention and treatment. However, the substantial size of whole slide images (WSIs) makes examination labor-intensive and time-consuming. Current deep learning-based approaches struggle with the morphological diversity of cervical cytology and require specialized models for distinct diagnostic tasks, leading to fragmented workflows. Here, we present UniCAS, a cytology foundation model pre-trained on 48,532 cervical WSIs encompassing diverse patient demographics and pathological conditions. UniCAS enables various clinical analysis tasks, achieving state-of-the-art performance in slide-level diagnosis, region-level analysis, and pixel-level image enhancement. In particular, by integrating a multi-task aggregator for slide-level diagnosis, UniCAS achieves area under the curve (AUC) values of 92.60%, 92.58%, and 98.39% for cancer screening, candidiasis testing, and clue cell diagnosis, respectively, while reducing diagnostic time by 70% compared with conventional approaches. This work establishes a paradigm for efficient multi-scale analysis in automated cervical cytology, bridging the gap between computational pathology and clinical diagnostic workflows.
Cervical cancer screening involves the classification of whole slide images (WSI) and the detection of cervical lesion cells. Deep learning for cervical lesion analysis typically demands extensive cell-level annotations, which are costly and time-consuming to obtain. Conversely, weak slide-level annotations, assigning a single label to a gigapixel WSI, are more accessible but lack detailed lesion cell information. In this paper, we propose the Unified Feature Fusion Learning (UFFL) framework to optimize both WSI classification and cell detection in cervical cancer screening. By bridging slide- and cell-level annotations, UFFL employs weakly supervised learning via the CerviBank Refinement Unit (CBRU), which leverages weak slide-level annotations (WSA) to integrate limited cell-level annotations (LCA). It leverages context-enhanced Multiple Instance Learning (MIL) to address the annotation scale gap between tasks. The CBRU captures distinct positive and negative cell patterns to enhance feature representation within MIL. Additionally, Noise-Suppressed Positive Guidance (NSPG) improves latent space robustness by prioritizing positive instances, reducing noise, and enhancing separability from ambiguous regions. Experimental results demonstrate that UFFL significantly improves detection and classification performance across diverse annotation levels.
Numerous applications regard motion segmentation as a fundamental and vital process. A plethora of motion segmentation techniques have been introduced, with the subspace clustering-based method standing out, particularly because of its unsupervised nature. However, these methods often face a challenge in effectively handling nonlinear data with hybrid noise. In the present study, we propose a novel robust subspace clustering methodology, specifically designed to address the complexities inherent in motion segmentation tasks. We’ve termed it as Robust Subspace Clustering with Noise Suppression (RSCNS),which integrates hybrid noise reconstruction with a representation of data relationships. Specifically, we propose a hybrid noise modeling method by joining Correntropy and Cauchy function to suppress noise and outlier pollution. To restore the corrupted data, we treat the motion trajectory feature data matrix as an approximate low-rank matrix and design a truncated weighting nuclear norm regularization constraint. Meanwhile, the block diagonal regularizer (BDR) is incorporated into our model to ensure that motion trajectory features from the same moving object are clustered together. Experimental evaluations are conducted on various video datasets, demonstrating that RSCNS can effectively handle motion segmentation tasks not only in visible light video, but also in invisible light (infrared) video.
Magnetic Resonance Imaging (MRI) is highly susceptible to motion artifacts due to prolonged acquisition times, posing a persistent challenge to image fidelity and diagnostic reliability. Existing deep learning-based correction methods typically rely on a fixed strategy, applying uniform models regardless of motion severity. This lack of adaptability often leads to over-correction in mild cases and insufficient restoration under severe motion. To overcome these limitations, we introduce FLEX-MoCo, a flexible motion correction pipeline that dynamically adapts its strategy based on both motion location and severity. FLEX-MoCo comprises two key components: a Motion Representation Learner that identifies motion artifacts from anatomical structures and identifies motion patterns, and a Correction Path Router that selects an optimal correction path tailored to the identified motion characteristics, ranging from lightweight models to multi-stage frameworks. Extensive experiments across diverse datasets containing both simulated and real-world motion artifacts demonstrate that FLEX-MoCo consistently surpasses state-of-the-art methods in motion correction and generalizes to unseen domains, underscoring the effectiveness of its motion-aware adaptive routing in handling varying motion patterns and severities.
The postnatal white matter connectome undergoes profound reorganization, yet the topological principles governing its spatiotemporal maturation remain largely unknown. Using connectome mapping, machine learning, and neurobiological annotation, we show hierarchical network development from birth to childhood and its association with neurobiological signatures. We identify two cardinal topological transformations that change rapidly during infancy and continue to refine into childhood, as characterized by nonlinear global increases in network efficiency and robustness to nodal attack, and regional reorganization with accelerated hub consolidation and prolonged modular reconfiguration, predominantly involving the prefrontal and insular cortices. Early developmental trajectories of these association cortices predict late childhood network architecture through local microstructural maturation of connected white matter tracts. These patterns align with well-established multiscale cortical hierarchies, including anatomical, evolutionary, and energy metabolism axes. Our findings reveal critical neurotopological milestones after postnatal development and establish a unified multiscale framework linking macroscale network dynamics to biologically constrained rules. This study reveals a hierarchical development of the brain’s structural connectome from infancy to childhood, characterized by distinct sensorimotor-association trajectories and alignment with multiple neurobiological hierarchies.
Ultrasound-based computer-aided diagnosis (CAD) has indicated effectiveness for developmental dysplasia of the hip (DDH). Recently, foundation models have shown promising potential to further improve the performance of a CAD model. However, their deployment in clinical applications remains challenging due to the high computational cost and overfitting risk of full fine-tuning. To address this issue, a new memory-efficient adaptation framework, named External Spatial Adapter Tuning (ESAT), is proposed for the CAD of DDH. The ESAT develops spatial adapters based on lowrank depthwise separable convolution at each layer of foundation model, so as to extract and refine hierarchical spatial features. These adapter outputs are further fused externally for the classification task, so that the gradient flow is fully decoupled from the foundation model. In real-world DDH diagnosis experiments, the ESAT requires only 30.16 % and 0.304 % of the memory and parameter usage on ViT-Base (as foundation model) compared with the full fine-tuning, yet achieves higher diagnostic accuracy than the full fine-tuning and other parameter-efficient tuning methods.
Background There are few studies on the brain structure of preterm infants based on low-grade intraventricular hemorrhage. The purpose of this study was to report changes in total and regional brain structure abnormalities in preterm neonates (PN) with low-grade intraventricular hemorrhage (grades I and II) and no other associated MRI abnormalities, and to correlate these changes with gestational age (GA). Methods We examined 76 preterm neonates (26 with low-grade IVH and 50 without IVH) who showed no focal abnormalities on magnetic resonance imaging (MRI) at term-equivalent age. The structural assessment of the brain involves measuring the surface area, thickness, average curvature, and volume of different regions. Retrospective analysis of brain magnetic resonance images of 25 healthy full-term infants and comparison with premature newborns, whose age after postmenstrual was similar. Results Compared with the control preterm infants, the infants with low-grade IVH had decreases in the following:1) Total brain surface area;2) the surface area in Orbitofrontal-Med Right; 3) brain volume in Right Orbitofrontal-Med and Right Hippocampus and Right Thalamus. Conclusions Our study reveals the potential harmful effects of low-grade IVH on surface area and volume development in preterm infants compared to those without IVH at term-equivalent age, underscoring its clinical significance for neurodevelopment in infants with low-grade IVH.
Accurate automatic segmentation of medical images typically requires large datasets with high-quality annotations, making it less applicable in clinical settings due to limited training data. One-shot segmentation based on learned transformations (OSSLT) has shown promise when labeled data is extremely limited, typically including unsupervised deformable registration, data augmentation with learned registration, and segmentation learned from augmented data. However, current one-shot segmentation methods are challenged by limited data diversity during augmentation, and potential label errors caused by imperfect registration. To address these issues, we propose a novel one-shot medical image segmentation method with adversarial training and label error rectification (AdLER), with the aim of improving the diversity of generated data and correcting label errors to enhance segmentation performance. Specifically, we implement a novel dual consistency constraint to ensure anatomy-aligned registration that reduces registration errors. Furthermore, we develop an adversarial training strategy to augment the atlas image, which ensures both generation diversity and segmentation robustness. We also propose to rectify potential label errors in the augmented atlas images by estimating segmentation uncertainty, which can compensate for the imperfect nature of deformable registration and improve segmentation authenticity. Experiments on the CANDI and ABIDE datasets demonstrate that the proposed AdLER outperforms previous state-of-the-art methods by 0.5% (CANDI), 3.9% (ABIDE "seen"), and 5.0% (ABIDE "unseen") in segmentation based on Dice scores, respectively.
Cervical cancer remains a significant global health challenge, particularly in low- and middle-income countries. While the Pap smear is a highly effective screening tool, its manual evaluation is time-consuming and prone to human error. Recent deep learning advancements automate cytology tasks; however, early detectors often struggle with fine-grained classification and severe class imbalances, particularly for underrepresented lesions. In this paper, we present Team jht010312's benchmarking methodology and results for the ISBI 2026 RIVA Cervical Cytology Challenge. We evaluate multiple state-of-the-art object detection models on the RIVA dataset, including Co-DETR, D-FINE, EVA-02, Mask R-CNN, and RetinaNet. Our fine-tuned Co-DETR model achieves the best overall performance, yielding a Preliminary and Final Phase mAP of 0.2273 and 0.1937 in Track A, and 0.6036 and 0.5906 in Track B, securing 1st and 3rd place, respectively. Code is available at https://github.com/peter-fei/ISBI-RIVA.
RATIONALE AND OBJECTIVES:Dextro-Transposition of the great arteries (d-TGA), the second most common cyanotic congenital heart defect, requires early risk stratification to identify high-risk arterial switch operation (ASO) patients for long-term prognosis. The aim of this study was to 1) determine biventricular strain and rapid long-axis strain (RLAS) could detect impaired biventricular mechanical in D-TGA patients and 2) establish and validate strain-based prognostic models for improved risk stratification and prediction of composite adverse events. MATERIALS AND METHODS:All patients with D-TGA were divided into training test set and external test set who underwent cardiovascular magnetic resonance (CMR) examination were included in the study. The composite adverse events were defined as complications that necessitated reintervention by echocardiography after CMR. Cardiac function and strain variables were initially considered for inclusion in three models. L1 regularization feature selection was constrained to a maximum of five variables. Its predictive performance was evaluated using the integrated area under the receiver operating characteristic curve (AUC). RESULTS:This study recruited 52 patients in the training set and 20 in the external test set. The primary adverse events were observed in 22 patients (42.3%) and five patients (25%). Compared with controls, the training set patients had significantly greater right ventricle (RV) global circumferential strain (GCS) (-16.32±3.40 vs. -9.41±2.89, p <0.001) and RV GRS (26.93±7.49 vs. 15.03±5.02, p < 0.001) and had significantly reduced left ventricle (LV) atrioventricular junction longitudinal strain (AJLS) (-17.32±3.62 vs. -19.87±2.93, p = 0.033). LV AJLS >18% had significant risk stratification of experiencing composite adverse events (p = 0.004) in the training set group. The model 3 yielded higher diagnostic performance than the other models (AUC=0.87,95% CI:0.83,0.93) in the training set. The model 3 (AUC=0.91, 95% CI:0.87,0.92) integrating biventricular strain and RLAS demonstrated excellent prognostic performance compared with model 1 (AUC=0.71, 95% CI:0.66,0.74) and model 2(AUC=0.67, CI:0.62,0.74) in external set group. CONCLUSION:LV AJLS and RV GCS were found to be independent predictors of composite adverse events. LV AJLS had significant risk stratification of experiencing composite adverse events. The model integrating biventricular strain and RLAS demonstrated excellent prognostic performance compared with traditional models in asymptomatic D-TGA patients.
Brain network is commonly divided into modules for analyzing their functionally segregated roles for group-level analysis in neuroimaging studies. Here, we introduce stochastic modules within brain networks for a robust probabilistic measurement of structural-functional module consistency (SFMC) in a group of subjects. Specifically, a stochastic module can be regarded as the chance of a brain region across subjects potentially being assigned to a group-level sub-network, characterized as an assignment probability for this brain region. This novel method has two advantages for evaluating inhomogeneous modules in brain networks. The first is that it can robustly evaluate the consistency between brain structural and functional modules whose population sizes are not necessary the same, and the second is that it is able to take into account the inter-individual variability of the modules for the groups. Moreover, compared with the conventional structural-functional coupling approach, our stochastic module-based method reveals a more pronounced decline in the coupling between structure and function, indicating stronger developmental reorganization. Our results using the dataset from Baby Connectome Project (BCP) show that the SFMC decreases from 0 to 5 years old, and is greater in primary brain regions, such as visual areas, while lower in more advanced cognitive regions, including those related to attention, control, and default mode network.
Reconstructing missing modalities of magnetic resonance images (MRIs) is a significant challenge in the medical imaging field. Current generative approaches such as generative adversarial networks (GANs) and diffusion models (DFs) have shown promise in synthesizing high-fidelity 2-D slices. However, they fall short in producing high-quality 3-D results due to their inability to leverage the contextual information across adjacent slices, resulting in low-quality images with poor interplane consistency. This problem is exacerbated when trained on datasets obtained with 2-D scanning protocols, where different MRI modalities have varying resolution between slices, leading to blurry and low-resolution results. To overcome these issues, we propose a novel fine-tuning strategy that can enhance the 2-D multimodal synthesis models to improve both consistency and interplane resolution in 3-D images. We begin by developing a novel attention-based module that can effectively empower generative models to produce 3-D results with high consistency between adjacent slices. Furthermore, we incorporate self-supervised super-resolution (SR) to deal with the varying resolution problem and eventually improve the interplane resolution of generated images. Instead of directly applying SR models that may bring about potential artifacts, we design a novel uncertainty modulation strategy, which optimizes the utilization of super-resolved images for high-quality reconstruction. We extensively evaluate our method through supervised and unsupervised multimodal synthesis using different generative models, including GAN- and DM-based models, all of which demonstrate the superiority and flexibility of the proposed method.
Quantitative PET underpins diagnosis and treatment monitoring in neurodegenerative disease, yet systematic biases between PET-MRI and PET-CT preclude threshold transfer and cross-site comparability. We developed and validated the first unified, anatomically guided deep-learning framework to harmonize PET-MRI quantification to PET-CT standards across multiple tracers and scanner manufacturers. The model learns CT-anchored attenuation representations using a vision transformer autoencoder, aligns MRI features to the CT space via contrastive objectives, and performs attention-guided residual correction. In paired same-day scans (N = 70; 18F-FDG, 18F-florbetaben, and 18F-florzolotau), cross-platform bias fell by >80% while preserving inter-regional biological topology. The framework generalized zero-shot to held-out tracers (18F-florbetapir and 18F-FP-CIT) without retraining. Multicenter validation (N = 420; three sites, four vendors) reduced amyloid Centiloid discrepancies from 23.6 to 4.1 (close to, though slightly above, PET-CT test-retest variability) and aligned tau SUVR thresholds. These results support more consistent cross-platform diagnostic cut-offs and reliable longitudinal monitoring when patients transition between modalities, establishing a practical route to scalable, radiation-sparing quantitative PET in therapeutic workflows.
Multi-label text classification (MLTC) is a vital task in natural language processing (NLP), often requiring high-quality text representations generated by pre-trained language models (PLMs). However, the inherent input length constraints of PLMs limit their capacity to handle long texts effectively. To address this challenge, we propose an innovative framework for multi-label long text classification. Our approach incorporates a dynamic text segmentation algorithm that optimally partitions long texts, thereby mitigating the input length limitations of PLMs. Additionally, we enhance both text and label representations by integrating external knowledge, modeling label co-occurrence relationships, and employing attention mechanisms. Extensive experiments conducted on diverse MLTC datasets demonstrate the superior performance of our method and uncover intricate relationships between texts and their associated labels. The code is available at https://github.com/Coder-Jeffrey/SKFRL
To evaluate the feasibility of an AI system for identifying active tuberculosis (ATB) in TB-specialized hospitals in high-prevalence settings. An AI system designed to identify ATB was retrospectively validated using a multi-center dataset of 1741 CT images from three TB-specialized hospitals. The dataset included ATB, pneumonia, pulmonary nodules and normal cases. The system’s utility and generalizability were assessed across four application scenarios, and pairwise comparisons of the system’s performance were conducted among the three hospitals. The system demonstrated good generalizability across three settings. It achieved an AUC over 0.9 for distinguishing between abnormal and normal, over 0.95 for distinguishing between ATB and normal, over 0.8 for distinguishing between ATB and non-ATB, and an AUC ranging from 0.762 to 0.906 for distinguishing between ATB and other abnormalities (pneumonia and pulmonary nodules). For all evaluation matrices, at least one pairwise comparison showed no significant difference in performance among the three hospitals across different scenarios. Using an AI system to identify ATB in CT images is feasible in TB-specialized hospitals. This evaluation provides valuable insights for those looking to implement AI to support clinical decision-making and optimize resource utilization in hospitals overwhelmed by TB cases.
Multiple cameras can provide comprehensive multi-view video coverage of a person. Fusing this multi-view data is crucial for tasks like behavioral analysis, although it traditionally requires camera calibration—a process that is often complex. Moreover, previous studies have overlooked the challenges posed by self-occlusion under multiple views and the continuity of human body shape estimation. In this study, we introduce a method to reconstruct the 3D human body from multiple uncalibrated camera views. Initially, we utilize a pre-trained human body encoder to process each camera view individually, enabling the reconstruction of human body models and parameters for each view along with predicted camera positions. Rather than merely averaging the models across views, we develop a neural network trained to assign weights to individual views for all human body joints, based on the estimated distribution of joint distances from each camera. Additionally, we focus on the mesh surface of the human body for dynamic fusion, allowing for the seamless integration of facial expressions and body shape into a unified human body model. Our method has shown excellent performance in reconstructing the human body on two public datasets, advancing beyond previous work from the SMPL model to the SMPL-X model. This extension incorporates more complex hand poses and facial expressions, enhancing the detail and accuracy of the reconstructions. Crucially, it supports the flexible ad-hoc deployment of any number of cameras, offering significant potential for various applications.
Pap smear cell classification represents a complex diagnostic challenge fundamentally distinct from natural image analysis. This clinical task demands dual capabilities: precise morphological characterization at cellular resolution and robust suppression of non-target tissue interference. Although recent pathological foundation models have created new paradigms for medical image analysis, performance gaps remain significant when applying histopathology-pretrained models to cervical cytopathology tasks. In this paper, we detail the cjdbehumble team's winning solution for the Pap Smear Cell Classification Challenge (PS3C). Our approach addresses these challenges through two key innovations: First, we employ parameter-efficient adaptation of multiple foundation models using low-rank adaptation (LoRA). Then, we implement an ensemble learning strategy that integrates differential features from multiple foundational models through a unified decision-making mechanism. Our method achieved groundbreaking results in the PS3C, securing first place with outstanding performance metrics. The code is available at https://github.com/peter-fei/ISBI-PS3C.
Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video and audio while balancing appropriateness, realism, and diversity. While prior works have mostly focused on uni-modal inputs or simplified reaction mappings, recent approaches such as PerFRDiff have explored multi-modal inputs and the one-to-many nature of appropriate reaction mappings. In this work, we propose the Facial Reaction Diffusion (ReactDiff) framework that uniquely integrates a Multi-Modality Transformer with conditional diffusion in the latent space for enhanced reaction generation. Unlike existing methods, ReactDiff leverages intra- and inter-class attention for fine-grained multi-modal interaction, while the latent diffusion process between the encoder and decoder enables diverse yet contextually appropriate outputs. Experimental results demonstrate that ReactDiff significantly outperforms existing approaches, achieving a facial reaction correlation of 0.26 and diversity score of 0.094 while maintaining competitive realism. The code is open-sourced at github.