The rapid development of deep learning-based computational pathology and genomics has demonstrated the significant promise of effectively integrating whole slide images (WSIs) and genomic data for cancer survival prediction. However, the substantial heterogeneity between pathological and genomic features makes exploring complex cross-modal relationships and constructing comprehensive patient representations challenging. To address this, we propose the Information Compression-based Multimodal Confidence-guided Fusion Network (iMCN). The framework is built around two key modules. First, the Adaptive Pathology Information Compression (APIC) module employs learnable information centers to dynamically cluster image regions, removing redundant information while maintaining discriminative survival-related patterns. Second, the Confidence-guided Multimodal Fusion (CMF) module utilizes a learned sub-network to estimate the confidence of each modality's representation, allowing for dynamic weighted fusion that prioritizes the most reliable features in each case. Evaluated on the TCGA-LUAD and TCGA-BRCA cohorts, iMCN achieved average concordance index (C-index) values of 0.691 and 0.740, respectively, outperforming existing state-of-the-art methods by an absolute improvement of 1.65%. Qualitatively, the model generates interpretable heatmaps that localize high-association regions between specific morphological structures (e.g., tumor cell nests) and functional genomic pathways (e.g., oncogenesis), offering biological insights into genomic-pathologic linkages. In conclusion, iMCN significantly advances multimodal survival analysis by introducing a principled framework for information compression and confidence-based fusion. Besides, correlation analysis reveal that tissue heterogeneity influences optimal retention rates differently across cancer types, with higher-heterogeneity tumors (e.g., LUAD) benefiting more from aggressive information compression. Beyond its predictive performance, the model's ability to elucidate the interplay between tissue morphology and molecular biology enhances its value as a tool for translational cancer research.
Granular flows are ubiquitous in nature and industrial applications, yet a complete continuum theory remains a long-standing challenge. The leading empirical approach, μ(I) rheology, lacks microscopic foundations and becomes multivalued in dense, slowly sheared flows where nonlocal corrections are required. Exploiting state-of-the-art high-speed X-ray tomography to investigate microscopic dynamics of dense granular flows in a Couette geometry, we establish a new, universal constitutive law spanning quasi-static to inertial regimes based on structural relaxation, resolving the fundamental difficulty in the original μ(I) framework. By further establishing a non-equilibrium statistical framework for granular flows, we demonstrate an intrinsic analogy between driven granular matter and hard-sphere liquids owing to their identical Carnahan-Starling equation of state, naturally explaining our rheological approach and the emergence of glassy behaviors. Our framework unifies granular rheology with the broader physics of disordered systems and provides a complete, microscopically-based theoretical framework for dense granular flow.
Obtaining multiple CT scans from the same patient is required in many clinical scenarios, such as lung nodule screening and image-guided radiation therapy. Repeated scans would expose patients to higher radiation dose and increase the risk of cancer. In this study, we aim to achieve ultra-low-dose imaging for subsequent scans by collecting extremely undersampled sinogram via regional few-view scanning, and preserve image quality utilizing the preceding fullsampled scan as prior. To fully exploit prior information, we propose a two-stage framework consisting of diffusion model-based sinogram restoration and deep learning-based unrolled iterative reconstruction. Specifically, the undersampled sinogram is first restored by a conditional diffusion model with sinogram-domain prior guidance. Then, we formulate the undersampled data reconstruction problem as an optimization problem combining fidelity terms for both undersampled and restored data, along with a regularization term based on image-domain prior. Next, we propose Prior-aided Alternate Iterative NeTwork (PAINT) to solve the optimization problem. PAINT alternately updates the undersampled or restored data fidelity term, and unrolls the iterations to integrate neural network-based prior regularization. In the case of 112 mm field of view in simulated data experiments, our proposed framework achieved superior performance in terms of CT value accuracy and image details preservation. Clinical data experiments also demonstrated that our proposed framework outperformed the comparison methods in artifact reduction and structure recovery.
Objective.The scanning process of mono-window intraoral scanner (Mono-IOS) is often cumbersome, and the impression accuracy is insufficient for edentulous arches implant rehabilitation. To address these limitations, this paper proposes the design and evaluation of a novel tri-window IOS (Tri-IOS), which is capable of simultaneously capturing occlusal, buccal, and lingual data.Approach.Tri-IOS features a U-shaped configuration of three plane mirrors in its scanning head, creating three virtual cameras. The relative poses of the three virtual cameras are calibrated using a sphere-based fitting method. The system integrates three-directional point cloud inputs per frame by relative poses, followed by iterative closest point algorithm for precise tracking and reconstruction. The system's effectiveness is validated on constructed dataset, and a prototype is built to conduct calibration and imaging experiments.Main Results.The design achieves a scanner head size under 2.5 times the thickest molar width, with synchronized three-directional data acquisition. Results on the constructed datasets showed that the tri-window design reduced impression time to one-third of the mono-window setup. And compared to mono-window, tri-window design enhanced both local detail and overall shape fidelity of the reconstruction results and significantly improved accuracy in edentulous arch implant impressions.Significance.The Tri-IOS demonstrates significant advantages in both impression acquisition speed and accuracy. Furthermore, its mirror-based design exhibits strong potential for miniaturization and productization. It is of significant importance for improving the performance of IOS and advancing product realization.
Cone-beam Computed Tomography (CBCT) is essential for target localization and treatment planning in image-guided radiotherapy for lung cancer. Dynamic reconstruction is an ill-posed inverse problem, as each motion state is captured by only one single projection. Several supervised learning methods have been proposed, but they rely on training with paired simulated or real datasets, and reconstruct images at discrete respiratory phases. Therefore, we propose a self-supervised learning method for lung imaging in the radiotherapy, adaptive Dynamic implicit neural representation (aDiner), to reconstruct dynamic CBCT images for each timepoint. Specifically, an adaptive composite representation (ACR) is introduced. In ACR, the inter- and intra-cycle encoding (iICE) captures the inter- and intra-cycle similarities of lung motion and ensures temporal coherence. Moreover, a probability-based feature fusion (PFF) is proposed to fuse decoupled static and dynamic features, composing the spatiotemporal representation. To improve the motion estimation, we further design a motion-guided coarse-to-fine sampling (MCS) strategy to focus on dynamic regions, with adaptive projection-ray sampling (APS) and adaptive deformed-point sampling (ADS). Compared with the previous methods, aDiner separately estimates static and dynamic regions, and takes advantage of the inter- and intra-cycle similarities to achieve better reconstruction quality. aDiner was evaluated on the simulated dataset, the physical phantom, and the clinical dataset. The results showed that aDiner could accurately and robustly reconstruct dynamic CBCTs for each timepoint. It also captured the lung motion in both regular and irregular respirations. The dynamic demos are available at: https://github.com/Henryzyf/aDiner.
Deep image prior (DIP) has been proven to be an effective method to improve the image quality of positron emission tomography (PET) images when large datasets are not available. However most DIP-type methods are based on 3D U-net, which may be not suitable for unsupervised PET image denoising. In this work, we developed a neural architecture search method to find a better neural architecture for unsupervised PET image denoising. The cell level architecture and the network level architecture are searched alternatively. In each searching step, we designed a super network as a full set of all possible networks. Then the high-count PET images were used to update architecture and the low-count PET images were used to update weights. The performance of the searched architecture was evaluated through phantom study. The results showed that the searched neural architecture can better preserve image details while maintaining a lower noise level.
Diabetic retinopathy (DR) grading is crucial for early diagnosis and treatment, but severe class imbalance in DR datasets impedes grading accuracy and compromises classification performance. While deep learning-based medical image synthesis has been widely adopted to address it, existing methods still struggle to capture the correlation of critical lesion features and preserve fine-grained pathological details, yielding images inconsistent with authentic pathology. To overcome these limitations, this paper proposes an Adaptive Cross-layer Fusion and Dense Generative Adversarial Network (ACFD-GAN) to synthesize realistic DR images and employ them for DR grading. Specifically, we develop a Lightweight Residual Dense Block (LRDB) to effectively capture and integrate local multi-level features with reduced computational overhead. To enhance information exchange between the encoder and decoder, we design an Adaptive Cross-layer Feature Fusion (ACFF) module that dynamically weights and integrates shallow detailed and deep semantic features based on spatial and channel attention, reinforcing spatial and semantic correlations with key lesion regions while preventing the loss of fine-grained details. Furthermore, we employ spatially continuous latent vectors sampled from the latent space combined with random noise to adaptively modulate fused features through an adaptive modulation module (AMM), yielding generated images with more diversity and globally consistent structure. High-quality images are generated through weighted moving average model selection to augment the APTOS 2019 and Messidor datasets for DR grading. Experimental results demonstrate that the proposed model can synthesize high-quality DR images from limited training data, and that leveraging ACFD-GAN generated images can significantly improve classifier performance.
Cardiac CT provides comprehensive structural and functional heart imaging, but motion artifacts remain a fundamental challenge. Although modern scanners have improved temporal resolution through hardware advancements and electrocardiogram-gated scan protocols, cardiac CT is still limited to specific cardiac phases and may fail in clinical practice. Here, to overcome this challenge, we propose the MARVEL (Motion-Aware Reconstruction Via Embedded Learning of motion prior) framework for time-resolved cardiac CT imaging. The MARVEL synergistically integrates a powerful data-driven model (MPNet) with learned Motion Prior into a robust model-based reconstruction process. This hybrid design preserves the interpretability and flexibility of model-based reconstruction methods, while benefiting from the high performance and computational efficiency of data-driven approaches. Specifically, the proposed MP-Net with specially designed mixed spatiotemporal convolutions extracts motion information from 4D pre-reconstructed images. Through an embedded learning strategy, the model learns to predict a reconstruction-oriented cardiac motion vector field (MVF). Then, a motion-aware reconstruction is implemented according to the MVF to compensate for cardiac motion. As a result, the MARVEL achieves effective reduction of motion artifacts throughout the entire heart and breaks through the conventional phase limitations of cardiac CT imaging. Furthermore, MARVEL integrates seamlessly into standard CT workflows, facilitating its potential for clinical application. Extensive evaluations, including qualitative and quantitative assessments on both simulated and clinical datasets, as well as blinded reader studies, demonstrate MARVEL’s superiority over the comparison methods. Code and dynamic reconstruction demos will be available at https: //github.com/ZihengD/MARVEL_CardiacCT.
Accurate segmentation and classification of brain tumors from magnetic resonance imaging (MRI) are critical for effective diagnosis and treatment planning. This article proposes a novel framework for brain tumor segmentation and classification using deep learning. The segmentation model is based on a modified U-Net architecture, called residual feature pyramid attention U-net model (RFAU-Net), which incorporates residual blocks to enhance training depth, attention mechanisms to focus on relevant features, and a feature pyramid module to improve segmentation of small and complex tumor regions. To address class imbalance and pixel degradation during training, we introduce a combined loss (CL) function that integrates weighted focal loss (WFL), assigning higher weights to minority classes and reducing the influence of majority classes. The model is evaluated on two publicly available datasets, achieving state-of-the-art performance with a segmentation accuracy of 97%, a dice similarity coefficient (DSC) of 92.5%, and an intersection over union (IoU) of 92%. For tumor classification, we employ a multiheaded convolutional neural network (MHCNN), achieving 99.8% accuracy in classifying the O6-methylguanine-DNA methyltransferase (MGMT) methylation status. These results demonstrate the superiority of the RFAU-Net model over traditional U-Net and RESU-Net architectures, particularly in handling small tumor regions and class imbalance. Additionally, a user-friendly web application programming interface (API) is developed to classify brain tumors into MGMT methylated and unmethylated categories, enabling efficient integration of this model into clinical practice for improved diagnosis and treatment of gliomas.
Segmenting retinal blood vessels is critical for the early detection of retinal abnormalities. While significant progress has been achieved in vessel segmentation through deep learning techniques, existing methodologies still struggle with the effectiveness of extracting and integrating local-global features. To overcome these challenges, this paper introduces SAFH-Net, a hybrid end-to-end network architecture that synergistically integrates Swin Transformer and CNN with innovative shuffle attention mechanisms and adaptive feature fusion. Specifically, a parallel encoder architecture employs a Convolutional (Conv) block with a residual structure for local feature extraction alongside a hierarchical Swin Transformer with Shifted-Window Multi-head Self-Attention (SW-MSA) for global context modeling, thereby achieving comprehensive feature capture with minimal additional parameter overhead. Then, an improved Spatial Attention Feature Fusion (SAFF) module is used to enable pixel-level adaptive weighting for optimal local-global feature integration. Additionally, cross-channel and cross-spatial shuffle operations enhance the interaction between local details and global information efficiency while suppressing redundant information. We achieved accuracies of 97.28%, 97.55%, and 97.60% on the DRIVE, STARE, and CHASE_DB1 datasets, respectively. A series of experimental results demonstrates that the proposed model significantly outperforms other advanced methods in segmentation performance.
Stereo-electroencephalography (SEEG)-guided radiofrequency thermocoagulation (RF-TC) emerges as a feasible alternative for drug-resistant focal epilepsy with epileptogenic zones (EZs) located in functional cerebral areas, which is ineligible for resection surgery. However, the efficacy of RF-TC was considered limited. This study described a three-dimensional conformal SEEG-guided RF-TC method to enhance thermocoagulation efficacy and assessed its surgical outcomes in patients with EZs located in functional cerebral areas. Ten patients were retrospectively reviewed, and the pathogeneses included focal cortical dysplasia (FCD, n = 5) and gray matter heterotopia (GMH, n = 5). Three-dimensionally symmetrical electrode placement was designed based on reconstructed cerebral model integrating the etiology and surrounding cerebral cortices, vessels and fibers. The three-dimensional conformal SEEG-guided RF-TC procedure in the functional area was performed by these electrodes penetrating the etiology conformally. Two electrophysiological indicators extracted from SEEG recordings, specifically line length (LL) and approximate entropy (ApEn), were calculated and a support vector machine (SVM) classifier was used to predict outcomes. Eight patients (80
Four-dimensionalcone-beam computed tomography (4-D CBCT) provides respiration-resolved images and facilitates image-guided radiation therapy. However, the ability to reveal respiratory motion comes at the cost of image artifacts. As raw projection data are sorted into multiple respiratory phases, the reconstructed 4-D CBCT images are covered by severe streak artifacts. Although several deep learning-based methods have been proposed to address this issue, most algorithms formulate it as a 2-D image enhancement task, neglecting the dynamic nature of 4-D CBCT. In this article, we first identify the origin and appearance of streak artifacts in 4-D CBCT images. We find that streak artifacts exhibit a unique "rotational motion" along with the patient's respiration, distinguishable from diaphragm-driven respiratory motion in 4-D space. Therefore, we introduce RSTAR4D-Net, a 4-D model that performs rotational streak artifact reduction by exploring the dynamic prior of 4-D CBCT images. Specifically, we overcome the computational and training difficulties of a 4-D neural network. The specially designed model decomposes the 4-D convolutions into multiple lower-dimensional operations and thus efficiently processes a whole 4-D image. Additionally, a Tetris training strategy is proposed to effectively train the model using limited 4-D data. Extensive experiments substantiate the superior performance of RSTAR4D-Net compared to existing methods.
Radiation therapy is regarded as the mainstay treatment for cancer in clinic. Kilovoltage cone-beam CT (CBCT) images have been acquired for most treatment sites as the clinical routine for image-guided radiation therapy (IGRT). However, repeated CBCT scanning brings extra irradiation dose to the patients and decreases clinical efficiency. Sparse CBCT scanning is a possible solution to the problems mentioned above but at the cost of inferior image quality. To decrease the extra dose while maintaining the CBCT quality, deep learning (DL) methods are widely adopted. In this study, planning CT was used as prior information, and the corresponding strictly structure-preserved CBCT was simulated based on the attenuation information from the planning CT. We developed a hyper-resolution ultra-sparse-view CBCT reconstruction model, known as the planning CT-based strictly-structure-preserved neural network (PSSP-NET), using a generative adversarial network (GAN). This model utilized clinical CBCT projections with extremely low sampling rates for the rapid reconstruction of high-quality CBCT images, and its clinical performance was evaluated in head-and-neck cancer patients. Our experiments demonstrated enhanced performance and improved reconstruction speed.
Photon-counting computed tomography (PCCT) may dramatically benefit clinical practice due to its versatility such as dose reduction and material characterization. However, the limited number of photons detected in each individual energy bin can induce severe noise contamination in the reconstructed image. Fortunately, the notable low-rank prior inherent in the PCCT image can guide the reconstruction to a denoised outcome. To fully excavate and leverage the intrinsic low-rankness, we propose a novel reconstruction algorithm based on quaternion representation (QR), called low-rank quaternion reconstruction (LOQUAT). First, we organize a group of nonlocal similar patches into a quaternion matrix. Then, an adjusted weighted Schatten- p norm (AWSN) is introduced and imposed on the matrix to enforce its low-rank nature. Subsequently, we formulate an AWSN-regularized model and devise an alternating direction method of multipliers (ADMM) framework to solve it. Experiments on simulated and real-world data substantiate the superiority of the LOQUAT technique over several state-of-the-art competitors in terms of both visual inspection and quantitative metrics. Moreover, our QR-based method exhibits lower computational complexity than some popular tensor representation (TR) based counterparts. Besides, the global convergence of LOQUAT is theoretically established under a mild condition. These properties bolster the robustness and practicality of LOQUAT, facilitating its application in PCCT clinical scenarios. The source code will be available at https://github.com/linzf23/LOQUAT.
To develop a deep learning (DL) model for predicting disease-free survival (DFS) in clinical stage I lung cancer patients who underwent surgical resection using pre-treatment CT images, and further validate it in patients receiving stereotactic body radiation therapy (SBRT). A retrospective cohort of 2489 clinical stage I non-small cell lung cancer (NSCLC) patients treated with operation (2015–2017) was enrolled to develop a DL-based DFS prediction model. Tumor features were extracted from CT images using a three-dimensional convolutional neural network. External validation was performed on 248 clinical stage I patients receiving SBRT from two hospitals. A clinical model was constructed by multivariable Cox regression for comparison. Model performance was evaluated with Harrell’s concordance index (C-index), which measures the model’s ability to correctly rank survival times by comparing all possible pairs of subjects. In the surgical cohort, the DL model effectively predicted DFS with a C-index of 0.85 (95
Objective:This study aimed to assess the correlation and consistency between quantitative CT (QCT) and MRI asymmetric echo least squares estimation iterative water-lipid separation sequence (IDEAL-IQ) in determining pancreatic fat content in patients with type 2 diabetes. Methods:A total of 67 patients with type 2 diabetes mellitus who met the inclusion criteria were included in the study. QCT and MRIIDEAL-IQ technologies were utilized to evaluate the patients quantitatively. The pancreatic head, body, and tail regions were examined to measure the fat content and obtain the CT pancreatic fat fraction (CT-PFF) and MRI pancreatic fat fraction (MR-PFF). Pearson correlation analysis examined the relationship between diabetes-related factors and CT-PFF/MR-PFF. Additionally, Bland-Altman analysis assessed the consistency between CT-PFF and MR-PFF. Results:Among the 67 patients, 33 were males and 34 were females. The average age was (66.55±6.23) years, with an average abdominal circumference of (83.34 ± 10.10) cm. The mean values for glycated hemoglobin, fasting blood glucose, BMI, and liver fat content were (6.97±1.07) mmol • L-1, (6.83±1.82) mmol • L-1, (24.02 ± 2.96) kg/m², and (5.28±2.76)%, respectively. Pearson correlation analysis indicated a significant correlation between abdominal circumference, liver fat content, and MR-PFF (r=0.261, 0.267, P < .05). However, no significant correlation was observed between age, glycated hemoglobin, fasting blood glucose, BMI, and MR-PFF (all, P > .05). The minimum and maximum values for CT-PFF among the 67 patients were 7.3% and 60.3%, respectively, with an average value of (19.90±10.61)%. For MR-PFF, the minimum and maximum values were 2% and 48%, respectively, with an average value of (12.21±10.71)%. Pearson correlation analysis demonstrated a significant correlation between CT-PFF and MR-PFF (r = .842, P < .05). Bland-Altman analysis revealed an average bias value of 7.7% and a standard deviation of 5.6% for CT-PFF and MR-PFF. The mean 95% confidence interval ranged from 4.15% to 19.75% (P < .05), with 64 cases falling within this interval and 3 cases falling outside. Conclusion:A correlation exists between pancreatic fat content, abdominal circumference, and liver fat content. Both QCT and MRI can accurately quantify pancreatic fat content, and their correlation and consistency are relatively ideal. QCT technology is particularly suitable for patients with contraindications for magnetic resonance examination.
Objective. CT-guided interventional procedures hold a significant position in clinical practice. However, due to the high number of scans and prolonged procedure times, patients are exposed to considerable radiation doses. This study aims to utilize intraoperative x-ray imaging as an alternative to intraoperative CT guidance, integrating preoperative prior scan images to reconstruct interventional CT. Approach. We present Feature back-projection Transformer (FBFormer), a reconstruction framework using Transformer and a novel feature back-projection module to reconstruct tomographic data from ultra-sparse(less than 10 views) intraoperative x-ray images. A mixed UNet and PVT encoder is employed to extract the local and global feature of the x-ray images and preoperative image prior, supporting more various input, while a feature back-projection module is utilized to achieve 2D-3D conversion in a more natural and geometrically consistent manner. Main Results. Extensive experiments on our simulated 4DCT and clinical validation datasets demonstrate the effectiveness of the proposed method and our method achieves superior results in reconstruction performance metrics and outperforms other compelling methods. Significance. The proposed method has the potential to significantly reduce radiation exposure during interventional procedures while maintaining high-quality reconstruction. The proposed FBFormer framework demonstrates the potential to enhance the clinical reliability of interventional CT reconstruction from ultra-sparse x-ray projections by the integration of image priors and the geometric preserving feature back-projection module.
Computer-assisted methods for pathological image analysis can improve doctor's efficiency of image reading and diagnostic accuracy, effectively addressing the shortage of pathology diagnostic manpower. With the rapid development of artificial intelligence and digital pathology, deep learning technology has spurred a wealth of research in the field of histopathology. This article reviews the various applications of deep learning in digital pathological image analysis, such as pathological image segmentation, cancer auxiliary diagnosis, and cancer prognosis prediction, and discusses the challenges and solutions in its application. Furthermore, it predicts future trends in deep learning for pathological image analysis and proposes potential research directions.
Cone-beam computed tomography (CBCT) is a widely used imaging technique. In practical applications, reducing projection views can decrease radiation exposure and accelerate scanning speed, with potential benefits for stationary CT systems. However, ultra-sparse-view acquisition (e.g., ≤ 30 views) introduces severe streak artifacts that degrade image quality. This poses a critical challenge because it is difficult to differentiate artifacts from real structures. We propose a sliding volume-based streak artifact reduction network (S-STAR Net) to remove artifacts while preserving structural details. Our method introduces three key technical innovations: (1) A sliding sampling sub-volume approach to process 3D sub-volumes, fully leveraging spatial context. (2) A difference enhancement (DE) loss to help separate artifacts from real structures. (3) Novel network includes volume-attention aided residual (VAR) blocks and Fourier transform convolution (FTC) blocks for multi-domain feature learning. Evaluated under 30 projection views, the method was tested on two datasets: a walnut dataset and the CQ500 head CT dataset. Quantitative metrics (PSNR/SSIM) and qualitative assessments (multi-planar visualization, residual error maps) demonstrate that S-STAR Net achieves superior performance in both artifacts suppression and detail preservation compared to existing approaches. The proposed method effectively addresses streak artifacts and recovers subtle structures in ultra-sparse-view CBCT reconstruction. Its robustness suggests broad applicability for medical image denoising, artifact reduction, and 3D image enhancement tasks.
Lung cancer screening with computed tomography (CT) scans can effectively improve the survival rate through the early detection of lung cancer, which typically identified in the form of pulmonary nodules. Multiple sequential CT images are helpful to determine nodule malignancy and play a significant role to detect lung cancers. It is crucial to develop effective lung cancer classification algorithms to achieve accurate results from multiple images without nodule location annotations, which can free radiologists from the burden of labeling nodule locations before predicting malignancy. In this study, we proposed the sequential multi-instance learning (SMILE) framework to predict high-risk lung cancer patients with multiple CT scans. SMILE included two steps. The first step was nodule instance generation. We employed the nodule detection algorithm with image category transformation to identify nodule instance locations within the entire lung images. The second step was nodule malignancy prediction. Models were supervised by patient-level annotations, without the exact locations of nodules.We embedded multi-instance learning with temporal feature extraction into a fusion framework, which effectively promoted the classification performance. SMILE was evaluated by five-fold cross-validation on a 925-patient dataset (182 malignant, 743 benign). Every patient had three CT scans, of which the interval period was about one year. Experimental results showed the potential of SMILE to free radiologists from labeling nodule locations. The source code will be available at https://github.com/wyzhao27/SMILE.