Terahertz (THz) light's sensitivity to water, non-ionizing nature, and sub-millimeter resolution make it ideal for medical imaging, particularly for diagnosing and managing skin conditions such as eczema, psoriasis, and skin cancer. Traditional handheld THz probes suffer from positional errors, tremors, and inconsistent contact pressure, which worsen with prolonged use. To address this, robotic THz acquisition systems can be leveraged to segment regions of interest on the skin surface and compute the appropriate probe pose for scanning. This paper proposes a markerless method for automatic identification and estimation of the pose for the placement of the THz probe over high-curvature facial regions. An RGB-D camera is utilized to select, recognize, and track the scanning targets on the patient using an end-to-end network-based model for 3D human face mesh prediction. We then solve for the landing pose to position the probe at normal incidence to the tissue surface at the target scan regions. We evaluated scanning area localization accuracy, motion execution accuracy, and THz acquisition capability using a mannequin head and human volunteers. Results demonstrated target registration error of 1.71 mm, scan target localization accuracy of less than 3 mm and probe landing pose estimation accuracy of 2.769 +/- 3.02 mm in X, 6.62 +/- 0.6 mm in Y and 4.792 +/- 2.18 in Z and orientation error of 1.6 +/- 0.8 degrees.
Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medical robotics has been limited by a fundamental data problem: existing medical robotic datasets are small, single-embodiment, and rarely shared openly, restricting the development of foundation models that the field needs to advance. We introduce Open-H-Embodiment, the largest open dataset of medical robotic video with synchronized kinematics to date, spanning more than 49 institutions and multiple robotic platforms including the CMR Versius, Intuitive Surgical's da Vinci, da Vinci Research Kit (dVRK), Rob Surgical BiTrack, Virtual Incision's MIRA, Moon Surgical Maestro, and a variety of custom systems, spanning surgical manipulation, robotic ultrasound, and endoscopy procedures. We demonstrate the research enabled by this dataset through two foundation models. GR00T-H is the first open foundation vision-language-action model for medical robotics, which is the only evaluated model to achieve full end-to-end task completion on a structured suturing benchmark (25
Depth perception in robotic minimally invasive surgery remains a critical challenge for many downstream tasks, demanding advanced depth estimation techniques and ground truth data for their validation. Current datasets lack data with ground-truth depth information in dynamic scenarios; therefore, we present DRENDS (Depth in Robotic Endoscopy with Dynamic Scenarios)1, a novel dataset comprising sequences of high-resolution stereo images captured during the robotic laparoscopic manipulation of a human phantom and ex vivo porcine tissue, along with ground-truth point clouds for each frame and calibration data. The data were collected under three illumination conditions and across different anatomies involving tissue manipulation and non-rigid deformations. Our code for rectifying stereo images, handling camera-perspective occlusions, and obtaining depth maps per frame is open source for reproducibility and easy adaptation. Finally, we also conduct baseline evaluations using state-of-the-art depth estimation models to establish benchmark performance on our dataset. The results and data highlight the challenges and potential of metric temporally consistent depth estimation in robotic surgery, encouraging further advancements in tissue deformation prediction for medical applications. We publicly release DRENDS1 to foster innovation and collaboration in this critical field.
Robotic exploration of unknown soft objects presents significant challenges for autonomous systems due to unpredictable deformations and shape changes during manipulation. To address this, we propose a framework that integrates topology-aware 3D reconstruction with a topology-guided motion planner, enabling the discovery and reconstruction of previously hidden or concave regions. This topology-aware 3D reconstruction employs a novel representation of deformable objects by combining Cylinder & Ccaron;ech Complexes with point clouds, enabling rapid tracking of significant topology changes and detection of non-manifold boundaries. The topology analysis and canonical reconstruction guide motion planning by optimising grasp points and planning trajectories to reveal previously unseen surfaces through two actions: turning over and stretching. We validated our algorithm through simulations and experiments using the da Vinci Research Kit, demonstrating successful exploration with two or three manipulators. We showed it can fully explore surfaces of two everyday objects, a beanie and a rubber glove, and two cadaveric organs, a liver and a colon, within seven manipulations. Our method achieved a 45.6% improvement in 3D reconstruction accuracy compared to state-of-the-art point-cloud-based methods while also demonstrating the capability to detect and fix non-manifold geometry.
Morse-type tapers at the head-stem junction in total hip replacements (THRs) provide many benefits to permit a successful surgical outcome. However, with the introduction of modular tapered devices comes complications associated with fluid ingress and motion at the interface that can cause fretting corrosion, which has been implicated in clinical failure. Increased surface roughness amplitude (Ra) and angular mismatch to ensure taper contact closer to the equator of the femoral head are design features introduced for use with ceramic heads but have been adopted by metal head couples. While increased surface roughness amplitude has been found to contribute to fretting corrosion, there is a distinct lack of systematic studies investigating the interactions between angular mismatch and Ra. This study measured the fretting corrosion and motion response of clinically representative samples, in part reference to ASTM F1875, when subjected to uniaxial incremental dynamic loading. The fretting corrosion response was measured in situ with an integrated three electrode electrochemical cell. Motion at the head-neck interface was measured with a bespoke motion measurement solution based on eddy-current principles which uses four sensors to allow motion to be fully characterised in three dimensions. Key findings from this study included a 5-10-fold increase in current measured in the increased roughness amplitude samples, suggesting an increased susceptibility to fretting corrosion without a corresponding increase in motion. The distal samples engaged around the opening of the taper interface and presented the lowest current measurements but most off-axis subsidence. Findings from this study indicate that optimisation of the taper interfaces in THR, in terms of fretting corrosion and motion, can be made and can be assessed using short-term preclinical tests.
Robotic deformable object manipulation (DOM) faces critical challenges in industrial and medical applications due to under-actuation, unpredictable deformation, and partial observability. Model-free methods often suffer from unstable Jacobians arising from ill-conditioned observations, while physics-based models typically depend on precise parameters and volumetric meshing, limiting their real-time practicality. We propose a wavelet-boundary element method (BEM) framework that leverages multiscale wavelet descriptors to control 3D deformations directly from efficient feedback modalities, such as contours and curves. By coupling wavelets with BEM, we derive an analytical deformation Jacobian that functions independently of material stiffness (e.g., Young's modulus), relying solely on an online-calibrated Poisson's ratio. This mesh-free formulation significantly enhances real-time performance and robustness against sensor occlusion. Validated in simulation and on the da Vinci Research Kit (dVRK) with phantom and ex vivo animal tissue, our method achieves millimetre-level accuracy. Comparative studies against Fourier-based, model-free, and online finite element method (FEM) approaches demonstrate superior stability and computational efficiency. Notably, our framework achieves convergence speeds significantly faster than online FEM by avoiding volumetric computations, while resolving ill-conditioning through spatial-frequency localization. This work advances deformable object manipulation in unstructured environments, particularly in surgical robotics, where stability under partial observability is essential. Project page: https://junleihu.github.io/projects/dwtbem/.
Access to the small bowel remains challenging due to its length, diameter, and curvature, making conventional push endoscopes insufficient for deep exploration. Improved tools that can travel throughout the small bowel without risking mucosal trauma, maintain stable visualization, and offer fine distal manipulation are needed for diagnosis and targeted treatment of small-bowel diseases. This work presents the Inchiscope, a miniaturized modular self-propelled soft robotic endoscope developed to meet the challenges associated with small-bowel access. The presented configuration combines nine independent actuation units to deliver three anchoring balloons and two parallel bellows actuator (PBA) segments for extension/contraction and steering. The modular approach facilitates a minimized diameter (13 mm) while enabling three operation modes including full-body inchworm-inspired locomotion, independent distal segment control for traversing tight bends, and proximal anchoring for stabilized inspection. Characterization highlights stable anchoring at forces up to 12 N, a PBA elongation ratio of 300% and controlled omnidirectional bending to 110 degrees. Experiments in rigid and compliant phantoms demonstrate self-propelled locomotion at up to 11.4 mm/s and within tubular environments <25 mm in diameter and with bending radii between 30-45 mm, while the independently actuated distal tip delivered continuous 3-D control for observation and positioning in complex geometries.
Abstract Ultrasonic bone scalpels are known to offer benefits of low cutting force, high precision, low microdamage around the cut site, and tissue selectivity in surgical procedures. However, all current commercial devices are too large to be integrated with the flexible endo-wrist of a surgical robot, and therefore, there is a significant gap for innovation in miniature devices. Ultrasonic bone scalpels in use in clinical settings are all based on a bolted Langevin transducer (BLT), which consists of a pre-stressed piezoceramic ring stack, two end masses, and a cutting blade. The BLT-based device must operate in resonance to achieve sufficient displacement amplitude at the surgical tip to cut through bone, and this dictates its size. Flextensional transducers have emerged as an alternative, but these transducers generally contain a low volume of piezoelectric driving material, and hence cannot excite the required displacement amplitude, and their reliance on adhesive bonds in their fabrication means they fail at the excitation levels required for a bone surgery device. We present a flextensional configuration that forms an ultrasonic surgical device, where the vibration-amplifying metal caps are excited by a pre-stressed piezoelectric stack. In vitro ultrasonic bone cutting tests facilitated with a Kuka robot are performed for a range of cutting speeds and penetration rates. The results demonstrate effective integration with a Kuka robot and that bone cutting can be achieved with an extremely low cutting force ( < 1 N) and high precision (the width of the bone cut presents under 6% deviation from the thickness of the blade). The flextensional device overcomes both the large size of conventional ultrasonic osteotomy devices and the displacement amplitude limitations of other miniaturisation approaches. Integration with an articulated robotic endo-wrist is enabled, establishing a foundation for low-force and high-precision ultrasonic bone cutting in minimally invasive, anatomically constrained robotic surgical environments.
Computer vision-based technologies significantly enhance surgical automation by advancing tool tracking, detection, and localization. However, Current data-driven approaches are data-voracious, requiring large, high-quality labeled image datasets. Our Work introduces a novel dynamic Gaussian Splatting technique to address the data scarcity in surgical image datasets. We propose a dynamic Gaussian model to represent dynamic surgical scenes, enabling the rendering of surgical instruments from unseen viewpoints and deformations with real tissue backgrounds. We utilize a dynamic training adjustment strategy to address challenges posed by poorly calibrated camera poses from real-world scenarios. Additionally, automatically generate annotations for our synthetic data. For evaluation, we constructed a new dataset featuring seven scenes with 14,000 frames of tool and camera motion and tool jaw articulation, with a background of an exvivo porcine model. Using this dataset, we synthetically replicate the scene deformation from the ground truth data, allowing direct comparisons of synthetic image quality. Experimental results illustrate that our method generates photo-realistic labeled image datasets with the highest PSNR (29.87). We further evaluate the performance of medical-specific neural networks trained on real and synthetic images using an unseen real-world image dataset. Our results show that the performance of models trained on synthetic images generated by the proposed method outperforms those trained with state-of-the-art standard data augmentation by 10%, leading to an overall improvement in model performances by nearly 15%.
Terahertz (THz) light has the unique properties of being very sensitive to water, non-ionizing, and having sub-millimeter depth resolution, making it suitable for medical imaging. Skin conditions including eczema, psoriasis and skin cancer affect a high percentage of the population and we have been developing a THz probe to help with their diagnosis, treatment and management. Our in vivo studies have been using a handheld THz probe, but this has been prone to positional errors through sensorimotor perturbations and tremors, giving spatially imprecise measurements and significant variations in contact pressure. As the operator tires through extended device use, these errors are further exacerbated. A robotic system is therefore needed to tune the critical parameters and achieve accurate and repeatable measurements of skin. This paper proposes an autonomous robotic THz acquisition system, the PicoBot, designed for non-invasive diagnosis of healthy and diseased skin conditions, based on hydration levels in the skin. The PicoBot can 3D scan and segment out the region of interest on the skin’s surface, precisely position (± 0.5/1 mm/degrees) the probe normal to the surface, and apply a desired amount of force (± 0.1N) to maintain firm contact for the required 60 s during THz data acquisition. The robotic automation improves the stability of the acquired THz signals, reducing the standard deviation of amplitude fluctuations by over a factor of four at 1 THz compared to hand-held mode. We show THz results for skin measurements of volunteers with healthy and dry skin conditions on various parts of the body such as the volar forearm, forehead, cheeks, and hands. The tests conducted validate the preclinical feasibility of the concept along with the robustness and advantages of using the PicoBot, compared to a manual measurement setup.
Accurate instrument pose estimation is a crucial step towards the future of robotic surgery, enabling applications such as autonomous surgical task execution. Vision-based methods for surgical instrument pose estimation provide a practical approach to tool tracking, but they often require markers to be attached to the instruments. Recently, more research has focused on the development of markerless methods based on deep learning. However, acquiring realistic surgical data, with ground truth (GT) instrument poses, required for deep learning training, is challenging. To address the issues in surgical instrument pose estimation, we introduce the Surgical Robot Instrument Pose Estimation (SurgRIPE) challenge, hosted at the 26th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI) in 2023. The objectives of this challenge are: (1) to provide the surgical vision community with realistic surgical video data paired with ground truth instrument poses, and (2) to establish a benchmark for evaluating markerless pose estimation methods. The challenge led to the development of several novel algorithms that showcased improved accuracy and robustness over existing methods. The performance evaluation study on the SurgRIPE dataset highlights the potential of these advanced algorithms to be integrated into robotic surgery systems, paving the way for more precise and autonomous surgical procedures. The SurgRIPE challenge has successfully established a new benchmark for the field, encouraging further research and development in surgical robot instrument pose estimation.
We proposed a novel test-time optimisation (TTO) approach framed by a NeRF-based architecture for long-term 3D point tracking. Most current methods in point tracking struggle to obtain consistent motion or are limited to 2D motion. TTO approaches frame the solution for long-term tracking as optimising a function that aggregates correspondences from other specialised state-of-the-art methods. Unlike the state-of-the-art on TTO, we propose parametrising such a function with our new invertible Neural Radiance Field (InvNeRF) architecture to perform both 2D and 3D tracking in surgical scenarios. Our approach allows us to exploit the advantages of a rendering-based approach by supervising the reprojection of pixel correspondences. It adapts strategies from recent rendering-based methods to obtain a bidirectional deformable-canonical mapping, to efficiently handle a defined workspace, and to guide the rays' density. It also presents our multi-scale HexPlanes for fast inference and a new algorithm for efficient pixel sampling and convergence criteria. We present results in the STIR and SCARE datasets, for evaluating point tracking and testing the integration of kinematic data in our pipeline, respectively. In 2D point tracking, our approach surpasses the precision and accuracy of the TTO state-of-the-art methods by nearly 50
OBJECTIVE:To characterise and compare the effectiveness of sacral dressings at alleviating the transfer of shear forces to the skin and, therefore, to lower the risk of and prevent pressure injury (PI) development in bed-based patients. METHOD:Dressings were evaluated against a ratio of applied to transferred shear and peak shear force observed on the skin. The evaluation was undertaken using a custom benchtop in vitro experimental setup replicating the skin-bedsheet interface. Forces were applied using an automated test rig within which a pair of load cells (ATI Nano-17; ATI Industrial Automation Inc., US) were embedded, tracking the applied normal and shear forces. Normal loading regimes of 11.1N (6kPa) and 14.8N (8kPa) were applied, to match expected clinically relevant sacral pressures. Pressures were applied with an error of 2.3% and 0.6%, respectively. No significant difference was found in the loading of dressings (p>0.05). RESULTS:The performances of three dressing designs were assessed, with six samples taken for each. All dressings were found to improve significantly (p<0.001) in performance relative to the control dataset for the dynamic coefficient of friction (DCoF), peak shear at skin (Fτ) and shear ratio (τratio). The Allevyn Life dressing (AllLi; Smith+Nephew, UK) showed the lowest DCoF (0.372±0.052), which was significantly lower than all other dressings (p<0.001). There was no significant difference between Mepilex Border (MepB; Mölnlyke Health Care, Sweden) and Avarus Border Foam (AvBF; Medtrade Products Ltd, UK). AllLi recorded the lowest peak shear at the skin (1.01±0.07N), which was significantly lower (p<0.001) than that measured for MepB and AvBF. With regard to the shear ratio, AllLi (0.233±0.007 and 0.283±0.013) and MepB (0.255±0.011 and 0.275±0.013) exhibited a significantly (p<0.001) higher shear ratio at both pressures, respectively, than AvBF. CONCLUSION:Although significant differences were identified between dressing types, DCoF was not indicative of dressing performance. Dressing compression, rather than thickness, was indicative of peak shear buffering. All dressings reduced the shear observed at the skin with respect to the control. The reduction in shear ratio between the control and all dressings tested was significant, ranging from 0.528-0.567 at 6kPa and 0.510-0.560 at 8kPa. Prophylactic use of these types of dressings to reduce risk of PI (widely specified within hospital and care setting protocols) confirms that this benefit is seen in the clinical setting. However, while there was a significant difference between the three dressing types (p=0.039 at 6kPa and p=0.050 at 8kPa), further work is required to explore how this translates to a clinical setting. Shear buffering is a complex process and performance is dependent on the compound response of the constituent dressing elements, rather than being dominated by a single component.
In low-resource settings, there is a critical need for skilled surgeons. Alternative training processes that include computer-assisted surgical skill evaluation are essential to address this gap. Using tool detection, surgical videos can be leveraged to derive insights into surgical skill assessment. However, state-of-the-art laparoscopic tool detection methods usually have more complex architectures tailored for in vivo data, which suffer from challenges such as smoke, occlusion, bleeding, etc., which are absent from in vitro training contexts. Thus, this paper tests multiple anchor-based and anchor-free, convolution- and transformer-based, traditional (non-surgical domain-specific) computer vision deep learning state-of-the-art models. With various hardware configurations on a newly curated in-house laparoscopic box-trainer dataset, we emphasise real-time performance on low-cost embedded devices. Overall, the anchor-free YOLOv8-X model was the most accurate, achieving mAP 50 of 99.5% and mAP 50 : 95 of 96.6% with an inference time of 23.5 ms/ ≈ 42.6 FPS on an NVIDIA Jetson Orin Nano 8GB (comparable low-cost hardware which could be expected to run real-time skill assessment methods for surgical training boot camps in a resource-constrained environment). The most efficient model was YOLOv11-N, providing 3.1 ms/ ≈ 322.6 FPS with a performance difference of +0% mAP 50 and -2.1% mAP 50 : 95 . The results highlight the models' potential for effective real-time detection of surgical tools and are suitable for further downstream assessment of surgical skills, even in resource-constrained environments.
Premature fusion of craniofacial joints, i.e. sutures, is a major clinical condition. This condition affects children and often requires numerous invasive surgeries to correct. Minimally invasive external loading of the skull has shown some success in achieving therapeutic effects in a mouse model of this condition, promising a new non-invasive treatment approach. However, our fundamental understanding of the level of deformation that such loading has induced across the sutures, leading to the effects observed is severely limited, yet crucial for its scalability. We carried out a series of multiscale characterisations of the loading effects on normal and craniosynostotic mice, in a series of in vivo and ex vivo studies. This involved developing a custom loading setup as well as software for its control and a novel in situ CT strain estimation approach following the principles of digital volume correlation. Our findings highlight that this treatment may disrupt bone formation across the sutures through plastic deformation of the treated suture. The level of permanent deformations observed across the coronal suture after loading corresponded well with the apparent strain that was estimated. This work provides invaluable insight into the level of mechanical forces that may prevent early fusion of cranial joints during the minimally invasive treatment cycle and will help the clinical translation of the treatment approach to humans.
While technologies such as cryo- or radio-frequency ablation allow for less invasive treatment of tumors than resection, they still require needles to reach the target location with the potential risk of spreading tumor tissue around. High Intensity Focused Ultrasound (HIFU) on the other hand allows for a completely remote and concentrated delivery of energy to a target location without the need for direct access and is particularly well suited to be robotically guided and thus used as part of an autonomous system. While robotic HIFU devices have been extensively explored for extracorporeal applications, their application into a laparoscopic setting is still widely unexplored. This paper presents a novel robotic HIFU pick-up device along with an automated workflow that includes the autonomous acquisition of the tumor geometry via Ultrasound (US) imaging, trajectory planning and autonomous execution of the HIFU ablation constraint by the tissue surface. Therefore, a novel sensorised water-filled membrane is developed and evaluated, enabling hybrid force position control that allows for minimising interaction forces with the tissue surface while maintaining sufficient acoustic coupling for ablation. Experiments on a phantom with a HIFU probe dummy demonstrate the effectiveness of the approach in targeting hidden structures.
Robotic manipulation of 3-D soft objects remains challenging in the industrial and medical fields. Various methods based on mechanical modeling, data-driven approaches or explicit feature tracking have been proposed. A unifying disadvantage of these methods is the high computational cost of simultaneous imaging processing, identification of mechanical properties, and motion planning, leading to a need for less computationally intensive methods. We propose a method for autonomous robotic manipulation with 3-D surface feedback to solve these issues. First, we produce a deformation model of the manipulated object, which estimates the robots' movements by monitoring the displacement of surface points surrounding the manipulators. Then, we develop a 6-degree-of-freedom velocity controller to manipulate the grasped object to achieve a desired shape. We validate our approach through comparative simulations with existing methods and experiments using phantom and cadaveric soft tissues with the da Vinci research kit. The results demonstrate the robustness of the technique to occlusions and various materials. Compared to state-of-the-art linear and data-driven methods, our approach is more precise by 46.5% and 15.9% and saves 55.2% and 25.7% manipulation time, respectively.
Computer vision technologies markedly enhance the automation capabilities of robotic-assisted minimally invasive surgery (RAMIS) through advanced tool tracking, detection, and localization. However, the limited availability of comprehensive surgical datasets for training represents a significant challenge in this field. This research introduces a novel method that employs 3D Gaussian Splatting to generate synthetic surgical datasets. We propose a method for extracting and combining 3D Gaussian representations of surgical instruments and background operating environments, transforming and combining them to generate high-fidelity synthetic surgical scenarios. We developed a data recording system capable of acquiring images alongside tool and camera poses in a surgical scene. Using this pose data, we synthetically replicate the scene, thereby enabling direct comparisons of the synthetic image quality (27.796 ± 1.796 PSNR). As a further validation, we compared two YOLOv5 models trained on the synthetic and real data, respectively, and assessed their performance in an unseen real-world test dataset. Comparing the performances, we observe an improvement in neural network performance, with the synthetic-trained model outperforming the real-world trained model by 12
While only a limited number of procedures have image guidance available during robotically guided surgery, they still require the surgeon to manually reference the obtained scans to their projected location on the tissue surface. While the surgeon may mark the boundaries on the organ surface via electrosurgery, the precise margin around the tumor is likely to remain variable and not guaranteed before a pathological analysis. This paper presents a first attempt to autonomously extract and mark tumor boundaries with a specified margin on the tissue surface. It presents a first concept for tool-tissue interaction control via Inertial Measurement Unit (IMU) sensor fusion and contact detection from the electrical signals of the Electrosurgical Unit (ESU), requiring no force sensing. We develop and assess our approach on Ultrasound (US) phantoms with anatomical surface geometries, comparing different strategies for projecting the tumor onto the surface and assessing its accuracy in repeated trials. Finally, we demonstrate the feasibility of translating the approach to an ex-vivo porcine liver. We achieve mean true positive rates above $\mathbf {0.84}$ and false detection rates below $\mathbf {0.12}$ compared to a tracked reference for each calculation and execution of the marking trajectory for dummy and ex-vivo experiments.