Fusing complementary information from PET and CT is beneficial for tumor segmentation. However, few studies have considered the asymmetric roles of the two modalities: PET directly reflects tumor uptake and enables explicit tumor localization, whereas CT primarily provides anatomical context. Ignoring this modality asymmetry not only increases parameter requirements for PET/CT feature fusion, but also risks diluting discriminative PET tumor signals with abundant CT background, which is particularly detrimental for small tumors. To address these issues, we propose a novel CT-assisted PET network (CAP-Net) for efficient 3D tumor segmentation. CAP-Net integrates local features extracted by a shallow CNN branch with asymmetric feature fusion (SCAF) and global features obtained by the CAP-RWKV branch. Both branches treat PET as the dominant modality. CAP-RWKV incorporates a Local Variance Similarity-based Token Area Allocation (LVS-TAA) strategy, assigning key tokens to high-uptake PET regions while merging background into larger tokens, maximizing prior information retention. In addition, we introduce a Key Token Cross-Entropy (KT-CE) loss to mitigate class imbalance. Experiments on three public datasets demonstrate that CAP-Net achieves state-of-the-art performance, significantly improving segmentation accuracy while reducing parameters by 57.4% compared to the latest efficient model(Mobile U-ViT). Our code is available at https://github.com/truezfy/CAP-Net.
Optical coherence tomography (OCT) A-scan backscattering signals provide depth-resolved textural information about internal structures. However, conventional OCT imaging is limited by refraction-induced distortion and speckle noise, hindering fine detail resolution. While multi-angle imaging systems alleviate these issues through incoherent compounding of backscattering signals, in vivo applications face challenges: limited angular coverage during surface scanning degrades backscatter intensity compounding quality, and the absence of angular information introduces artifacts in multi-view position-intensity alignment. Furthermore, excessive smoothing during speckle suppression obscures fine textures. Consequently, reconstructing ultra-fine structures from limited-angle, sparse-view measurements remains a critical challenge. To address this, we present Backscattering-Corrected Implicit Representation Tomography (BCIRT), a framework for reconstructing multi-angle low-coherence signals. We also develop a dedicated limited-angle imaging system for intraoperative BCIRT deployment. BCIRT formulates cross-view backscattering signals as a continuous function of spatial position, utilizing implicit neural representation (INR) for fitting. A physics-informed iterative mechanism inversely models ray propagation to determine corrected ray paths, enhancing the neural representation’s robustness against distortions. Leveraging these corrected paths, we introduce a dual dynamic line mixer and a contrastive-guided discriminative deblurring module to achieve high-resolution microstructure reconstruction with reduced speckle noise. Extensive experiments on biological samples and surgical resected samples demonstrate that our method achieves state-of-the-art performance, highlighting its potential for clinical applications and biomedical research.
Accurate and reliable 3D scene reconstruction is a key component of intelligent surgery, enabling enhanced spatial understanding and data-driven analysis in minimally invasive surgery (MIS). However, existing clinical systems are often bulky and workflow-incompatible, while vision-based Structure-from-Motion methods struggle with sparse textures and specularities, leading to unstable pose estimation and high computational cost. To address these limitations, we present SurGSplat++, a progressive, pose-free Gaussian splatting framework for monocular surgical scene reconstruction that requires no auxiliary hardware or pre-computed camera poses. Experiments show that SurGSplat++ achieves improved geometric stability, reduced pose drift, and superior novel-view synthesis compared with existing approaches. By producing accurate and consistent 3D reconstructions, the proposed method provides a practical solution for post-operative analysis, pre-operative planning, and data-driven surgical modeling in clinical environments.
Magnetic resonance imaging (MRI) is a widely adopted non-invasive imaging tool for both clinical diagnosis and neuroscientific research. Nonetheless, the quality of MRI is often hampered by noise. Supervised deep learning-based denoising has proven to outperform conventional methods but requires high-signal-to-noise ratio (SNR) reference data for supervising the training, which considerably reduces its practical feasibility. To address this challenge, we propose a new iterative residual learning strategy entitled "Noise2Average" for denoising MRI data with multiple repetitions, which can be combined with transfer learning for subject-specific self-supervised training. Noise2Average learns to map each noisy repetition to the average of all noisy repetitions by fine-tuning parameters of a pre-trained convolutional neural network (CNN) and recovers higher SNR by averaging all denoised results at the first iteration, and performs this supervised residual learning-based denoising process repeatedly with the denoising results from the previous iteration as the training target for several iterations. The efficacy of Noise2Average is systematically and comprehensively demonstrated on four types of commonly acquired MRI data, including two or more consecutively acquired highly accelerated T-1-weighted (T1w) image volumes, two T1w image volumes acquired with different echo times, two diffusion-weighted image (DWI) volumes acquired with opposite phase encoding directions, and two DWI volumes synthesized using different sets of DWI volumes from a diffusion tensor imaging (DTI) scan. Quantitative evaluations show that Noise2Average preserves more image sharpness and textural details and produces more accurate quantitative microstructural metrics from DTI signal modeling than the classic Noise2Noise method and conventional benchmark denoising methods BM4D and AONLM, with denoising performance slightly inferior to that of supervised learning-based denoising method. By reducing the requirement for training data and scan time, Noise2Average substantially increases the feasibility and accessibility of deep learning-based denoising methods for MRI and potentially benefits a wider range of clinical and neuroscientific studies.
The Robotic Ultrasound System (RUSS) has the potential to transform medical imaging by addressing limitations such as operator dependency, diagnostic variability, and reproducibility in traditional ultrasound (US) examination. Despite rapid technological advancements, a substantial gap remains between RUSS research progress and clinical adoption. This review examined the clinical roles and engineering advances of RUSS, identifying key barriers to translation. Clinically, it evaluated the current applications of RUSS in supporting US procedures, while from an engineering standpoint, it summarized recent innovations and remaining technical challenges. This review examined the current state-of-the-art RUSS technologies, categorizing them based on diverse organ-specific applications while also analyzing their core functional capabilities. This review revealed a focus disparity: while abdominal US is the most commonly used in clinical practice, vascular-targeted RUSS dominates current research. It also highlighted a misalignment between research priorities and actual clinical tasks. Current studies predominantly focused on autonomous scanning and imaging, with limited attention to downstream tasks such as disease diagnosis and analysis. Building on these observations, it identified critical challenges and future trends in RUSS development. This work provides a foundation for future research, fostering collaboration between clinicians and engineers to accelerate the translation of next-generation RUSS from bench to bedside.
The increasing incidence of lower gastrointestinal neuroendocrine tumors (NETs) necessitates improved methods for early and accurate detection. Automatic segmentation of NETs in endoscopic ultrasound (EUS) images is particularly challenging due to low image contrast and indistinct tumor boundaries. This study proposes and validates GismEUS, a geometry-aware deep learning model for automated NET segmentation in EUS images. We propose GismEUS, a deep learning architecture that integrates dual features. The model employs Endoscopic Planar Geometric Feature Projection to capture fine-grained local features and Endoscopic Stereoscopic Structural Feature Modeling to extract comprehensive global features. These feature sets are fused by the Dual Structure-Guided Feature Enhancement module, which applies both semantic-aware and distance-aware attention and subsequently combines their outputs via a local–global gated fusion. The model was trained and evaluated on a private annotated dataset for EUS images of NETs and the public GIST514-DB dataset. GismEUS demonstrated superior segmentation performance across multiple evaluation metrics, achieving a Dice score of 0.6735, and significantly outperformed established benchmarks on the EUS datasets. These comprehensive results validate the effectiveness of the proposed feature integration strategy in addressing the unique challenges of EUS imaging. This study demonstrates that the novel combination of planar and stereoscopic features significantly improves segmentation accuracy and introduces a high-quality annotated dataset for NETs segmentation in EUS images. GismEUS model offers a reliable and accurate automated tool for early NETs diagnosis, supporting more consistent clinical decision-making and potentially improving patient outcomes.
Despite rapid commercialization of surgical robots, their autonomy and real-time decision-making remain limited in practice. To address this gap, we propose ArthroCut, an autonomous policy learning framework that upgrades knee arthroplasty robots from assistive execution to context-aware action generation. ArthroCut fine-tunes a Qwen--VL backbone on a self-built, time-synchronized multimodal dataset from 21 complete cases (23,205 RGB--D pairs), integrating preoperative CT/MR, intraoperative NDI tracking of bones and end effector, RGB--D surgical video, robot state, and textual intent. The method operates on two complementary token families -- Preoperative Imaging Tokens (PIT) to encode patient-specific anatomy and planned resection planes, and Time-Aligned Surgical Tokens (TAST) to fuse real-time visual, geometric, and kinematic evidence -- and emits an interpretable action grammar under grammar/safety-constrained decoding. In bench-top experiments on a knee prosthesis across seven trials, ArthroCut achieves an average success rate of 86% over the six standard resections, significantly outperforming strong baselines trained under the same protocol. Ablations show that TAST is the principal driver of reliability while PIT provides essential anatomical grounding, and their combination yields the most stable multi-plane execution. These results indicate that aligning preoperative geometry with time-aligned intraoperative perception and translating that alignment into tokenized, constrained actions is an effective path toward robust, interpretable autonomy in orthopedic robotic surgery.
Esophageal cancer, a serious malignancy, is primarily treated with radiotherapy (RT). Accurate delineation of gross tumor volume (GTV) and clinical target volume (CTV) is essential, yet manual contouring is time-consuming and automatic methods remain challenging. Adaptive radiotherapy enables personalized target adjustments, but tumor changes following initial RT complicate delineation, and automatic second-phase replanning is underexplored. We propose the Unspecified Pretraining and Modular Adaptation (UPMA) framework for context-aware second-phase esophageal target delineation. UPMA pretrains its network on diverse datasets to comprehend tumor and anatomical features, then modularly adapts it for second-phase tasks with minimal parameters. The method improves Dice scores by over 23% for GTV and over 5% for CTV, while reducing learnable parameters by over 73% during adaptation. In clinical evaluations, 73.7% and 52.6% of the predicted GTVs and CTVs required only minor edits. UPMA generalizes across diverse patient characteristics and imaging factors, supporting automated adaptive radiotherapy for precise, personalized treatment.
Tumors severely impact human health and quality of life. Intraoperative imaging, diagnosis, and treatment are often challenging and separated. Intelligent theranostic systems offer promising solutions for precise tumor real-time imaging and treatment. We present a Robotic-assisted Intelligent Optical Theranostics (RaIOT) platform, a real-time theranostic system that integrates wide-field optical coherence tomography and laser ablation for automated, precise tumor treatment. The system utilizes a 6-degree-of-freedom robotic arm to perform automated wide-field Optical Coherence Tomography (OCT) scanning, and automatically guides laser ablation. A deep learning model segments tumor tissues from the volumetric OCT data in real-time. Coordinate transformations based on robot-mediated calibration ensure high spatial registration accuracy between the imaging and treatment modules. Following segmentation, the system autonomously plans and executes an optimized laser ablation path to fully cover the target region. Experimental results on biological samples validate the system's spatially controlled resection and reliability. The results confirm that the RaIOT system achieves precise image-guided ablation, showcasing significant potential for intelligent robotic systems in clinical applications.
Brain tumors are serious neurological diseases, and MRI analysis plays a crucial role in clinical diagnosis and treatment. In clinical practice, thick-slice MRI scans are often acquired, but existing 3D segmentation networks struggle to achieve accurate delineation due to large slice thickness and limited inter-slice information. To address these challenges, we propose a Cross-layer Guided and Multi-branch Coupling Network (CGMC-Net), which tackles the problem from two complementary perspectives. A Sliced Cross-layer Mutual Attention (SCMA) module is introduced to model inter-slice dependencies by conditionally enhancing current-slice features using adjacent slices, while intra-slice representation is strengthened via Multi-dimensional Feature Expansion (MMFE). In addition, a clinically guided Region-Enhanced Branch (REB), composed of ESAM and RDFF, highlights tumor-relevant regions and improves feature discrimination. Results demonstrate that our proposed method significantly outperforms state-of-the-art (SOTA) methods, specifically, the proposed model performed well in the Dice coefficient, and Sensitivity (lesion and edema area: 91.74% and 74.18% (Highest); 92.34% and 89.86%(Highest)). Efficiency analysis further demonstrates a favorable accuracy-efficiency trade-off. Moreover, 3D tumor reconstruction based on the proposed segmentation enables reliable downstream HGG/LGG classification, supporting the practical clinical value of the framework. These results indicate that CGMC-Net provides accurate and robust segmentation for thick-slice MRI with strong potential for real-world deployment.
Confocal laser endomicroscopy (CLE) is a critical modality for the early, minimally invasive diagnosis of intraluminal diseases. However, its clinical translation is constrained by the inherent optical-mechanical trade-offs between numerical aperture (NA), probe diameter, and clinical maneuverability. Achieving high-performance imaging within the strict dimensional limits of standard biopsy channels remains a significant technical bottleneck. To address these challenges, we propose a clinically constrained optical design strategy. This integrated design approach incorporates anatomical boundary constraints directly into the optical optimization process. Based on this method, the developed miniature immersion objective achieved a 0.78 μm lateral resolution across a 300 μm field of view, while the integrated pCLE probe maintained a 1.1 μm lateral resolution. This system realizes stable cellular-level imaging under constrained geometry. It is bending-compatible and clinically deployable. Animal experiments and representative histology-correlated clinical gastric images demonstrated the feasibility of resolving tissue microstructures and clinically relevant mucosal abnormalities, supporting the translational potential of the integrated pCLE probe.
Understanding the human brain requires access to its microscopic tissue architecture. Diffusion magnetic resonance imaging (MRI) provides the only noninvasive window into whole-brain microstructure in vivo, yet reliable quantitative mapping remains confined to specialized research settings requiring dense sampling and optimized acquisition protocols. To address this gap, we present a physics-informed generative microstructure network (PIGMENT) that learns a universal generative prior of human brain microstructure and adapts it zero-shot to each participant's measured data to recover subject-specific maps. Trained on 11375 scans spanning multiple sites, vendors, and field strengths, PIGMENT enabled reliable quantitative mapping for tensor, kurtosis, and NODDI models across external datasets from five independent centers. It remains effective where conventional fitting becomes unreliable, recovering meaningful maps from extremely sparse acquisitions while supporting downstream tractography and structural connectivity mapping. PIGMENT estimates demonstrated strong biological validity, preserving submillimeter cortical microarchitectural patterns and early-childhood white matter developmental trajectories from 10-fold accelerated scans. Furthermore, PIGMENT enables reliable quantitative tensor mapping on cost-efficient low-field systems and the extraction of tumor-related biomarkers using ultra-fast clinical protocols. Together, these results establish PIGMENT as a physics-informed foundation model that extends quantitative diffusion MRI into regimes traditionally too sparse, heterogeneous, or clinically constrained for reliable analysis.
Volumetric liver ultrasound (US) plays an important role in clinical diagnosis but remains highly dependent on operator expertise, particularly during target view localization. This study aims to develop an automatic probe guidance framework that reduces reliance on manual demonstrations and additional sensing hardware, while supporting robust volumetric liver US acquisition. We propose an image-based imitation learning framework that learns probe guidance policies from a virtual expert in a simulated US scanning environment. A simulation pipeline is constructed using cross-modal medical images and a hybrid US simulator that combines physics-based ray casting with generation-based image synthesis to produce anatomically consistent and acoustically realistic US images. Optimal scanning trajectories are generated based solely on target views typically available in clinical practice. To improve robustness, pose-level and image-level data augmentations are introduced, and US observations are encoded into an anatomy-aware state representation for intercostal liver scanning. Experimental results in simulation and real clinical data demonstrate that the proposed framework achieves accurate and stable target view localization for volumetric liver US acquisition. Compared with baseline and ablated models, the method shows improved localization accuracy, increased liver coverage, and reduced rib interference, while maintaining robustness across different anatomical conditions. This work presents a data-efficient and clinically practical solution for automatic probe guidance in volumetric liver US. By leveraging realistic simulation, virtual expert demonstrations, and anatomy-aware image representations, the proposed framework enables effective learning of probe movements without requiring manual trajectory annotations or additional sensors. The results suggest strong potential for integration into computer-assisted and robotic US systems.
Bronchoscopy robots are pivotal for improving the accuracy and efficiency of early lung cancer diagnosis. However, the stiffness of existing systems is fixed, which limits the robot from accurately navigating and performing stable operations on distant peripheral nodules. This paper introduces a novel bronchoscopy robot equipped with a magnetic multi-segment variable stiffness catheter, whose stiffness can be precisely adjusted in real time. The soft state of the catheter can enhance the flexibility of movement, while the rigid state can provide stable support. The catheter's slender profile and wirelessly magnetic steerability facilitate access to the distal terminal bronchi for examination. The catheter achieved a maximum bending angle of 49.07 degrees in soft state and 1.11 degrees in rigid state. By independently controlling the stiffness of each segment, the catheter could conform to various complex shapes. In vitro experiments showed that the bronchoscopy robot could reach deep branches of the lung and operate in conjunction with tools such as lasers. The results of the live pig experiment confirmed that the bronchoscopy robot could navigate to the bronchial target location for detection. By accessing distal bronchi unreachable by commercial systems, our robot offers a superior solution for early-stage lung cancer bronchoscopy.
Surgical phase recognition is critical in computer-assisted surgery. Clinically, surgeons discriminate surgical phases through visuospatial analysis of instrument-tissue interactions. However, existing methods fail to adequately account for the critical role of the visual-neural mechanisms of the surgeons, consequently exhibiting performance limitations in dynamically complex surgical scenarios. Inspired by the information processing mechanism of dual-path visual cognitive neural mechanism, this letter proposes a novel SpatioTemporal Frequency state space modeling method (STFMamba). The method incorporates a frequency-domain balancer designed to capture and balance the low-frequency steady-state features representing tissue backgrounds and the high-frequency transient features representing instrument actions. Concurrently, it employs a dual-path state space model, where information is processed separately through a temporal-frequency path and a temporal-spatial path to decouple local high-frequency dynamic information from global steady-state information. Furthermore, a temporal-frequency guided cross-attention mechanism is introduced to effectively fuse these heterogeneous feature streams. To further enhance model robustness and accuracy, a boundary constraint mechanism is adopted to improve the classification capability for phase boundary frames. This neurocognitively inspired design, which emulates the dual-channel processing strategy of the human visual pathway, substantially enhances the robustness and accuracy of surgical phase recognition. Extensive experiments conducted on the Cholec80, AutoLaparo and CATARACTS datasets validate the superior performance of the proposed method.
Monocular endoscopy is widely used in minimally invasive surgery because of its simple hardware configuration and low system complexity. However, conventional two-dimensional displays inherently lack sufficient depth cues, hindering precise intraoperative perception. Light-field (LF) displays offer a promising glasses-free 3D alternative but require high-fidelity multi-view content. Synthesizing such content from sparse monocular observations remains a pivotal challenge, particularly under frequent tool occlusions and non-rigid tissue deformations. In this work, we propose a display-oriented framework for monocular endoscopic LF visualization based on spatiotemporally consistent 3D Gaussian Splatting (3DGS). Rather than treating reconstruction and display as independent stages, our framework explicitly optimizes scene completeness and geometric stability tailored for downstream multi-view rendering. Specifically, a canonical space is introduced to aggregate geometric information across frames for occlusion-aware geometric recovery, while anisotropic geometric regularization and cross-frame temporal anchoring are employed to enhance depth stability and disparity continuity. Leveraging this optimized representation, a virtual camera array matched to the display geometry generates high-quality elemental image arrays (EIAs). Experiments on EndoNeRF and SCARED datasets demonstrate superior performance, achieving a PSNR of 38.78 dB and an Abs Rel of 0.203. The pipeline enables efficient rendering at 188 FPS on an NVIDIA RTX 3090 GPU. Prototype validation further confirms smooth motion parallax and minimized visual artifacts during viewpoint transitions, highlighting the potential of the proposed framework for immersive endoscopic visualization.
Harnessing the full potential of capsule robots requires highly precise and reliable pose estimation. This article presents a dynamic bounding box optimization strategy based on Gaussian similarity for accurately estimating the pose of capsule robots in ultrasound (US) images. Specifically, a robot-adaptive label assignment (RALA) strategy is developed to effectively extract features through learning the geometric properties of capsule robots. Subsequently, a novel bounding box representation (BBR) is presented to accommodate variations in position and orientation. In addition, a predicted boxes similarity metric, based on the elliptical Gaussian distributions, is proposed to achieve optimal bounding box similarity beyond the conventional intersection over union (IoU) metric. This metric not only mitigates the limitations of IoU in nonoverlapping cases but also provides a more refined basis for network optimization. Finally, a dynamic and adaptive optimization strategy is introduced to enhance estimation accuracy. A localization error of 0.11 mm and an angle error of 1.00 degrees were achieved by the proposed model, demonstrating superior estimation accuracy. To evaluate the model's generalizability, a private dataset reflecting real-world clinical scenarios more closely was constructed. Without any fine-tuning, the model trained solely on the public dataset yielded a localization error of 0.39 mm and an angle estimation error of 2.60 degrees on the private dataset, confirming the robustness and generalization capability of the proposed model.
Restoring haptic feedback remains a major challenge in robot-assisted minimally invasive surgery (RMIS), especially for localizing subsurface tumors and estimating their depth. This paper presents a compact, high-density tactile sensor array based on fiber Bragg grating (FBG) sensing for robotic palpation. The array uses a honeycomb topology to increase spatial sampling within an 8.5 mm footprint while preserving high sensitivity. Static and dynamic tests show a linear force-wavelength response across seven channels, an average force resolution of 8.43 mN, and \((<)\)1% full-scale dynamic error. To estimate tumor depth from palpation signals, we propose FBG-PatchFormer. Since depth annotations are scarce and costly to obtain, FBG-PatchFormer leverages contrastive self-supervised pretraining on unlabeled palpation windows to reduce the reliance on dense labels, and is then finetuned for depth classification. On phantom palpation with embedded inclusions, FBG-PatchFormer achieves 99.60% record-level accuracy on a ten-class depth task (blank and 2--10 mm). In vivo tests on porcine liver further demonstrate robust tumor localization under physiological motion and fluid interference, supporting the clinical potential of the proposed sensing system.
ABSTRACT Minimally invasive surgery (MIS) and robot‐assisted MIS reduce access trauma but attenuate direct haptic cues at the tool–tissue interface. Consequently, interaction forces are often inferred indirectly from vision, which can increase uncertainty during force‐sensitive tasks such as grasping, suturing, palpation, and cannulation. Fiber Bragg grating (FBG) sensors are attractive for MIS instruments because they are compact, multiplexable, and intrinsically immune to electromagnetic interference, enabling distal force measurement in confined operative spaces. This review surveys FBG‐based force sensors reported for MIS and robot‐assisted MIS, covering sensing principles, mechanical transducer architectures, and interrogation strategies. We summarize representative implementations across ophthalmic microsurgery, natural‐orifice procedures, vascular intervention, and laparoscopy and compare key performance metrics reported in the literature. We further review force/temperature decoupling approaches from calibration‐matrix methods to learning‐based models and discuss practical constraints affecting translation, including packaging and sterilization, strain‐transfer drift, temperature compensation, dynamic bandwidth, and integration with surgical workflows. Finally, we outline research directions to improve robustness and validation toward clinically deployable optical force‐feedback systems.