Optical coherence tomography (OCT) A-scan backscattering signals provide depth-resolved textural information about internal structures. However, conventional OCT imaging is limited by refraction-induced distortion and speckle noise, hindering fine detail resolution. While multi-angle imaging systems alleviate these issues through incoherent compounding of backscattering signals, in vivo applications face challenges: limited angular coverage during surface scanning degrades backscatter intensity compounding quality, and the absence of angular information introduces artifacts in multi-view position-intensity alignment. Furthermore, excessive smoothing during speckle suppression obscures fine textures. Consequently, reconstructing ultra-fine structures from limited-angle, sparse-view measurements remains a critical challenge. To address this, we present Backscattering-Corrected Implicit Representation Tomography (BCIRT), a framework for reconstructing multi-angle low-coherence signals. We also develop a dedicated limited-angle imaging system for intraoperative BCIRT deployment. BCIRT formulates cross-view backscattering signals as a continuous function of spatial position, utilizing implicit neural representation (INR) for fitting. A physics-informed iterative mechanism inversely models ray propagation to determine corrected ray paths, enhancing the neural representation’s robustness against distortions. Leveraging these corrected paths, we introduce a dual dynamic line mixer and a contrastive-guided discriminative deblurring module to achieve high-resolution microstructure reconstruction with reduced speckle noise. Extensive experiments on biological samples and surgical resected samples demonstrate that our method achieves state-of-the-art performance, highlighting its potential for clinical applications and biomedical research.
Accurate and reliable 3D scene reconstruction is a key component of intelligent surgery, enabling enhanced spatial understanding and data-driven analysis in minimally invasive surgery (MIS). However, existing clinical systems are often bulky and workflow-incompatible, while vision-based Structure-from-Motion methods struggle with sparse textures and specularities, leading to unstable pose estimation and high computational cost. To address these limitations, we present SurGSplat++, a progressive, pose-free Gaussian splatting framework for monocular surgical scene reconstruction that requires no auxiliary hardware or pre-computed camera poses. Experiments show that SurGSplat++ achieves improved geometric stability, reduced pose drift, and superior novel-view synthesis compared with existing approaches. By producing accurate and consistent 3D reconstructions, the proposed method provides a practical solution for post-operative analysis, pre-operative planning, and data-driven surgical modeling in clinical environments.
The Robotic Ultrasound System (RUSS) has the potential to transform medical imaging by addressing limitations such as operator dependency, diagnostic variability, and reproducibility in traditional ultrasound (US) examination. Despite rapid technological advancements, a substantial gap remains between RUSS research progress and clinical adoption. This review examined the clinical roles and engineering advances of RUSS, identifying key barriers to translation. Clinically, it evaluated the current applications of RUSS in supporting US procedures, while from an engineering standpoint, it summarized recent innovations and remaining technical challenges. This review examined the current state-of-the-art RUSS technologies, categorizing them based on diverse organ-specific applications while also analyzing their core functional capabilities. This review revealed a focus disparity: while abdominal US is the most commonly used in clinical practice, vascular-targeted RUSS dominates current research. It also highlighted a misalignment between research priorities and actual clinical tasks. Current studies predominantly focused on autonomous scanning and imaging, with limited attention to downstream tasks such as disease diagnosis and analysis. Building on these observations, it identified critical challenges and future trends in RUSS development. This work provides a foundation for future research, fostering collaboration between clinicians and engineers to accelerate the translation of next-generation RUSS from bench to bedside.
Esophageal cancer, a serious malignancy, is primarily treated with radiotherapy (RT). Accurate delineation of gross tumor volume (GTV) and clinical target volume (CTV) is essential, yet manual contouring is time-consuming and automatic methods remain challenging. Adaptive radiotherapy enables personalized target adjustments, but tumor changes following initial RT complicate delineation, and automatic second-phase replanning is underexplored. We propose the Unspecified Pretraining and Modular Adaptation (UPMA) framework for context-aware second-phase esophageal target delineation. UPMA pretrains its network on diverse datasets to comprehend tumor and anatomical features, then modularly adapts it for second-phase tasks with minimal parameters. The method improves Dice scores by over 23% for GTV and over 5% for CTV, while reducing learnable parameters by over 73% during adaptation. In clinical evaluations, 73.7% and 52.6% of the predicted GTVs and CTVs required only minor edits. UPMA generalizes across diverse patient characteristics and imaging factors, supporting automated adaptive radiotherapy for precise, personalized treatment.
Volumetric liver ultrasound (US) plays an important role in clinical diagnosis but remains highly dependent on operator expertise, particularly during target view localization. This study aims to develop an automatic probe guidance framework that reduces reliance on manual demonstrations and additional sensing hardware, while supporting robust volumetric liver US acquisition. We propose an image-based imitation learning framework that learns probe guidance policies from a virtual expert in a simulated US scanning environment. A simulation pipeline is constructed using cross-modal medical images and a hybrid US simulator that combines physics-based ray casting with generation-based image synthesis to produce anatomically consistent and acoustically realistic US images. Optimal scanning trajectories are generated based solely on target views typically available in clinical practice. To improve robustness, pose-level and image-level data augmentations are introduced, and US observations are encoded into an anatomy-aware state representation for intercostal liver scanning. Experimental results in simulation and real clinical data demonstrate that the proposed framework achieves accurate and stable target view localization for volumetric liver US acquisition. Compared with baseline and ablated models, the method shows improved localization accuracy, increased liver coverage, and reduced rib interference, while maintaining robustness across different anatomical conditions. This work presents a data-efficient and clinically practical solution for automatic probe guidance in volumetric liver US. By leveraging realistic simulation, virtual expert demonstrations, and anatomy-aware image representations, the proposed framework enables effective learning of probe movements without requiring manual trajectory annotations or additional sensors. The results suggest strong potential for integration into computer-assisted and robotic US systems.
In robot-assisted spinal endoscopy, intraoperative imaging is frequently degraded by bleeding, irrigation fluids, bubbles, smoke, and uneven illumination, which can severely compromise surgical precision, safety, and decisionmaking. Accurate identification of anatomical structures is particularly critical in spinal procedures, yet acquiring paired clean and degraded images in real clinical settings is infeasible. To address this challenge, we propose DCP-Net, an unpaired endoscopic image restoration framework tailored for robotic spinal surgery. DCP-Net integrates Diffusion-Prior Contrastive Learning (DPCL) to leverage generative priors and contrastive objectives for robust latent representations, and Physics-Informed Constraints (PIC) to ensure anatomically consistent restoration. Furthermore, we introduce Diffusion-Prior Uncertainty Estimation (DPUE), providing pixel-wise confidence maps that quantify restoration reliability and guide risk-aware robotic perception. We further constructed a dataset comprising 21,845 paired/unpaired samples of intraoperative visual degradations in spinal endoscopy, primarily involving bleeding, bubbles, and other artifacts. Extensive experiments show that DCP-Net outperforms existing methods in both quantitative metrics and perceptual quality, significantly improving visual clarity and supporting various robotic navigation tasks. Among these tasks, accurate bleeding point detection plays a particularly critical role in ensuring safe and precise navigation in clinical practice.
Tracking moving anatomical targets with robotic ultrasound is particularly challenging when the target motion is both fast and large in scale, as the end-to-end latency of existing systems prevents the perceptioncontrol loop from closing fast enough. In this paper, we argue that overcoming this limitation calls for the joint design of perception and control, rather than optimizing each in isolation. We present a tightly-coupled framework with two main components: (1) a Decoupled DualStream Perception Network that estimates 3D translational state from 2D ultrasound images at high frequency, and (2) a Single-Step Flow Policy that outputs an entire action sequence in one forward pass, removing the need for iterative rollouts used in conventional policies. Together, the two modules enable closed-loop control at over 60 Hz. In phantom experiments with complex 3D trajectories, the system achieves a mean tracking error below 6.5 mm and re-acquires the target after resultant displacements exceeding 170 mm. It tracks targets moving at speeds up to 102 mm/s with a terminal error under 1.7 mm. In-vivo trials on a human volunteer further confirm that the approach transfers to realistic clinical conditions. To our knowledge, this is the first RUSS framework to unify high-bandwidth dynamic tracking with large-scale repositioning within a single architecture, offering a concrete step toward autonomous ultrasound operation in the presence of patient motion.
Accuracy access to peripheral lung lesions is critical for diagnosis and treatment, yet clinical bronchoscopy remains limited by operator variability and loss of spatial accuracy during respiratory motion. Conventional navigation systems depend on preoperative maps and external tracking, which fail in the deformable and continuously moving bronchial tree. We developed a fully autonomous, adaptive bronchoscopy platform that performs real‑time navigation directly from intraluminal visual-spatial cues, without predefined trajectories or external localization. A hierarchical dual‑agent framework enables a Navigator to infer dynamic waypoints and a Driver to execute fine‑scale steering with closed‑loop precision. In dynamically ventilated ex vivo lungs and live canine models, the system achieved consistent, lesion‑targeted navigation under rapid respiration without human input. Beyond bronchoscopy, this approach defines a general computational paradigm for autonomy in deformable living anatomy, offering a clinically translatable path toward operator‑independent intervention across organ systems.
Liver disease is a major global health burden. While ultrasound is the first-line diagnostic tool, liver sonography requires locating multiple non-continuous planes from positions where target structures are often not visible, for biometric assessment and lesion detection, requiring significant expertise. However, expert sonographers are severely scarce in resource-limited regions. Here, we develop an autonomous lightweight ultrasound robot comprising an AI agent that integrates multi-modal perception with memory attention for localization of unseen target structures, and a 588-gram 6-degrees-of-freedom cable-driven robot. By mounting on the abdomen, the system enhances robustness against motion. Our robot can autonomously acquire expert-level standard liver ultrasound planes and detect pathology in patients, including two from Xining, a 2261-meter-altitude city with limited medical resources. Our system performs effectively on rapid-motion individuals and in wilderness environments. This work represents the first demonstration of autonomous sonography across multiple challenging scenarios, potentially transforming access to expert-level diagnostics in underserved regions.
Objective: During minimally invasive surgery (MIS), three-dimensional (3D) endoscopes provide valuable 3D perception of the patient's internal structures. However, due to the requirement of two cameras and a relatively large baseline distance, the imaging front-end of the conventional binocular 3D (CB3D) endoscope usually lacks compactness. We aim to develop a novel compact monocular dual-view 3D (MDV3D) endoscope imaging system. Methods: We develop a novel optical design for the MDV3D endoscope that exploits the dichroic prism's reflection capability to its internal light to realize MDV3D imaging, ensuring the 3D endoscope's imaging front-end with high compactness. Additionally, we propose a 3D reconstruction optimization method (MB-BEDE) to address the challenge of insufficient accuracy of 3D surface information posed by the typically micro baseline distance between the two virtual cameras in the MDV3D endoscope. Through seamless integration of our MDV3D endoscope and MB-BEDE method, we can obtain reliable real-time 3D information. Results: Evaluation experiments demonstrate our system's capability to provide accurate 3D surface information. Notably, compared to the CB3D endoscope imaging system occupying two channels in the robotic single-port laparo-endoscopic surgery (SPLS) platform, our system only requires one channel with a 5.60 mm diameter, presenting the advantage of creating more operating space for surgical instruments during robotic SPLS procedures. Conclusion and Significance: The proposed system and method present a novel solution for developing compact and cost-effective 3D endoscope imaging systems in MIS, particularly in robotic SPLS.
Conventional single-dataset training often fails with new data distributions, especially in ultrasound (US) image analysis due to limited data, acoustic shadows, and speckle noise. Therefore, constructing a universal framework for multi-heterogeneous US datasets is imperative. However, a key challenge arises: how to effectively mitigate inter-dataset interference while preserving dataset-specific discriminative features for robust downstream task? Previous approaches utilize either a single source-specific decoder or a domain adaptation strategy, but these methods experienced a decline in performance when applied to other domains. Considering this, we propose a Universal Collaborative Mixture of Heterogeneous Source-Specific Experts (COME). Specifically, COME establishes dual structure-semantic shared experts that create a universal representation space and then collaborate with source-specific experts to extract discriminative features through providing complementary features. This design enables robust generalization by leveraging cross-datasets experience distributions and providing universal US priors for small-batch or unseen data scenarios. Extensive experiments under three evaluation modes (single-dataset, intra-organ, and inter-organ integration datasets) demonstrate COME's superiority, achieving significant mean AP improvements over state-of-the-art methods. Our project is available at: https://universalcome.github.io/UniversalCOME/.
Deformable tissue retraction is a common but time-consuming task in robotic surgery. An autonomous robotic deformable tissue retraction system has the potential to help surgeons reduce cognitive burdens and focus more on critical aspects of the surgery. However, the uncertain deformation and complex constraints of deformable tissues pose significant challenges. We propose an autonomous deformable tissue retraction framework that incorporates visual representation and learning models, along with a 7-degree-of-freedom robotic system. For extracting deformation representations and learning to manipulate deformable tissues based on 2D images, we introduce a Sequential-information-based Contrastive State Representation Learning (SC-SRL) algorithm and a reinforcement learning model with asymmetric inputs and auxiliary losses. Experimental results show that the proposed framework achieved a 93.0% success rate of tissue retraction task in a simulated environment. Furthermore, our method demonstrates a safe retraction trajectory proportion of 92.5% based on a novel evaluation method using the histogram of feature angles of the tissue particles. The proposed framework can also be deployed on a real robotic system through a sim-to-real transfer pipeline, acquire policies for nearby tasks and perform resistance to visual dynamic disturbance. This study paves a new path for the application of vision-based intelligent systems in surgical robotics.
Intraoperative navigation relies heavily on precise 3D reconstruction to ensure accuracy and safety during surgical procedures. However, endoscopic scenarios present unique challenges, including sparse features and inconsistent lighting, which render many existing Structure-from-Motion (SfM)-based methods inadequate and prone to reconstruction failure. To mitigate these constraints, we propose SurGSplat, a novel paradigm designed to progressively refine 3D Gaussian Splatting (3DGS) through the integration of geometric constraints. By enabling the detailed reconstruction of vascular structures and other critical features, SurGSplat provides surgeons with enhanced visual clarity, facilitating precise intraoperative decision-making. Experimental evaluations demonstrate that SurGSplat achieves superior performance in both novel view synthesis (NVS) and pose estimation accuracy, establishing it as a high-fidelity and efficient solution for surgical scene reconstruction. More information and results can be found on the page https://surgsplat.github.io/.
Current retinal foundation models remain constrained by curated research datasets that lack authentic clinical context, and require extensive task-specific optimization for each application, limiting their deployment efficiency in low-resource settings. Here, we show that these barriers can be overcome by building clinical native intelligence directly from real-world medical practice. Our key insight is that large-scale telemedicine programs, where expert centers provide remote consultations across distributed facilities, represent a natural reservoir for learning clinical image interpretation. We present ReVision, a retinal foundation model that learns from the natural alignment between 485,980 color fundus photographs and their corresponding diagnostic reports, accumulated through a decade-long telemedicine program spanning 162 medical institutions across China. Through extensive evaluation across 27 ophthalmic benchmarks, we demonstrate that ReVison enables deployment efficiency with minimal local resources. Without any task-specific training, ReVision achieves zero-shot disease detection with an average AUROC of 0.946 across 12 public benchmarks and 0.952 on 3 independent clinical cohorts. When minimal adaptation is feasible, ReVision matches extensively fine-tuned alternatives while requiring orders of magnitude fewer trainable parameters and labeled examples. The learned representations also transfer effectively to new clinical sites, imaging domains, imaging modalities, and systemic health prediction tasks. In a prospective reader study with 33 ophthalmologists, ReVision's zero-shot assistance improved diagnostic accuracy by 14.8
Steerable catheters offer significant advantages over conventional catheters, including enhanced control, stability, and accessibility, which reduce operational complexity, fluoroscopy time, and radiation exposure, positioning them as a promising advancement for vascular interventional procedures. Herein, a novel steerable catheter is presented, featuring a hydraulically actuated, soft, steerable tip that allows for real‐time visualization in X‐ray imaging. To optimize performance, several silicone materials were evaluated for their mechanical properties, resulting in a soft tip design with a diameter of 2.6 mm. The tip incorporates an internal tool channel and supports a large bending angle of 180°. The tip demonstrates an average response time of 1.141 s (±0.750 s), a maximum output force of 0.145 N (±0.001 N), and a maximum radial expansion of 1.121 (±0.006). A steering kinematic model of the catheter tip is developed to simulate its movement. The catheter tip's real‐time shape and position information are obtained through intelligent segmentation and neighborhood‐based endpoint detection methods, assisting the surgeon during superselective procedures. The catheter's visibility and flexibility are validated in a live porcine model, demonstrating its potential for future use in interventional procedures.
Multi-angle optical coherence tomography (OCT) has attracted increasing attention in recent years due to its higher near-isotropic resolution and enhanced penetration depth compared to conventional OCT, making it promising for biomedical applications. However, accurately reconstructing quantitative optical parameters from sparse multi-angle signals while maintaining high resolution remains a challenging task. In this paper, we introduce a novel framework for multi-angle OCT image reconstruction, termed Implicit REpresentation for OCT (IREO), enabling the recovery of structural images and optical parameter distributions of tissue from backscattered signals impacted by refraction distortion and noise. IREO regards optical parameters as a continuous function of spatial position, which is fitted by a simple neural network. Through the imaging model of OCT, the theoretical imaging intensities are calculated from optical parameters output by the network, and then compared with intensity values acquired by OCT. The neural network parameters are optimized to form an implicit representation of the imaged sample. Through extensive evaluation, we confirmed the outstanding reconstruction quality of the proposed method for two- and three-dimensional multi-angle OCT data even with sparsely-acquired data, and verified the accuracy and feasibility of quantitative optical parameter estimation with custom-built multi-angle imaging systems. The multi-angle imaging reconstruction technique introduced here achieves high-resolution, deep-penetration, distortion-corrected, and multi-contrast imaging, enhancing label-free mesoscopic visualization of biological tissues and showing promise for advancing OCT applications in biomedical research and clinical theranostics.
Contact dynamics critically influence the sim-to-real performance of deformable tissue retraction—a representative contact-rich manipulation task in robotic surgery. The uncertainty in instrument-tissue interactions further complicates its automation. To address these challenges, we propose an architecture for contact dynamics prediction and online inference. We emphasize that the pre-contact phase and the fine adjustment of the contact location are critical for achieving stable retraction performance in surgical scenarios. By leveraging characteristics of the pre-contact phase, we efficiently extract the contact dynamics with timescale-sensitive state space model based on 2D images. Then, we perform smooth switching of control strategies based on posterior beliefs for fine adjustment of the contact location and resistance to unexpected disturbances. The proposed architecture only requires a few hours of real-world data collected with minimal human intervention to learn, predict and transfer. We demonstrate the interaction and generalization ability of the proposed architecture on a real robotic system and especially evaluate the resistance and recovery ability under unexpected disturbances. We also present autonomous tissue retraction task integrated with electrosurgical cutting and tissue dissection on an ex vivo porcine model to showcase an optimized surgical workflow.
Robotic ultrasound imaging systems primarily focus on enhancing automation in grayscale image acquisition but lack essential functional information, which restricts their clinical effectiveness and efficiency. In this regard, we propose a novel robotic ultrasound system capable of automatically screening and anomaly localization using quasi-static elastography (QE). For continuous screening, a compliant force control strategy is devised to manage the complex probe operations required for elasticity data acquisition. This involves the integration of adaptive out-of-plane posture control and in-plane palpation motion control. For anomaly localization, tissue strains are analyzed using multi-source motion data. We introduce an unsupervised tissue displacement estimation method, complemented by a strain estimator with multi-frame fusion for robust strain estimation. A 3D strain map is reconstructed to enable closed-loop control for automated robotic localization. The system has been validated through extensive experiments on two realistic phantoms and tested on human subjects. Results demonstrate that our system can perform robotic QE-based screening across subjects with varying surface conditions and lesion depths, showing improved efficiency and adaptability compared to existing systems. It achieves satisfactory accuracy in strain-based anomaly localization, with a detection rate of 0.77 and an average localization error of 1.01±0.47 mm for the abdominal phantom, and 0.73 and 3.44±0.84 mm for the more challenging thyroid phantom. By identifying and localizing suspicious anomalies in 3D space, the proposed system shows promise in providing preliminary dia
Diabetic macular edema (DME) is a leading cause of severe vision loss in the working-age population. Optical coherence tomography (OCT) is the gold standard for DME management and primary care referrals, providing retinal thickness maps (RTMs) that quantify retinal pathologies. However, its limited accessibility in resource-constrained settings necessitates more efficient solutions. While color fundus photography (C-FP) is a cost-effective screening tool, its potential for quantitative thickness evaluation remains underexplored. In this paper, we propose a novel Global-to-Local conditional Diffusion model for Retinal Thickness prediction (GLD-RT), the first attempt to predict RTM solely from C-FP. Our framework predicts thickness distributions of macular region from 2D inputs through a diffusion process guided by hierarchical global-to-local retinal features. Experimental results demonstrate that GLD-RT accurately depicts both physiological and pathological retinal morphology, achieving superior performance in thickness quantification and enabling a more detailed examination of retinal structures. Furthermore, C-FP-generated RTMs exhibit promising utility in facilitating DME diagnosis. This approach transforms conventional fundus imaging into a comprehensive and cost-effective diagnostic tool for DME screening and monitoring in resource-limited settings, thereby holding significant clinical implications.
Real-time tracking of dynamic targets amidst large-scale, high-frequency disturbances remains a critical unsolved challenge in Robotic Ultrasound Systems (RUSS), primarily due to the end-to-end latency of existing systems. This paper argues that breaking this latency barrier requires a fundamental shift towards the synergistic co-design of perception and control. We realize it in a novel framework with two tightly-coupled contributions: (1) a Decoupled Dual-Stream Perception Network that robustly estimates 3D translational state from 2D images at high frequency, and (2) a Single-Step Flow Policy that generates entire action sequences in one inference pass, bypassing the iterative bottleneck of conventional policies. This synergy enables a closed-loop control frequency exceeding 60Hz. On a dynamic phantom, our system not only tracks complex 3D trajectories with a mean error below 6.5mm but also demonstrates robust re-acquisition from over 170mm displacement. Furthermore, it can track targets at speeds of 102mm/s, achieving a terminal error below 1.7mm. Moreover, in-vivo experiments on a human volunteer validate the framework's effectiveness and robustness in a realistic clinical setting. Our work presents a RUSS holistically architected to unify high-bandwidth tracking with large-scale repositioning, a critical step towards robust autonomy in dynamic clinical environments.