Flexible ureteroscopy is a widely used surgical procedure for diagnosing and treating various urinary tract conditions, particularly kidney stones. Ensuring the complete extraction of all stones is crucial to prevent recurrence and the need for auxiliary interventions. Visual SLAM-based navigation systems have been proposed to assist surgeons by simultaneously estimating the 3D structure of the kidney and tracking the ureteroscope's tip position. However, most existing solutions assume a completely static environment, which does not account for the intraoperative situation. In this study, we extend the work of Oliva Maza et al. by incorporating real-time visual segmentation of kidney stones and surgical tools using either YOLOv7-E6E and segment anything or YOLO11m-seg. Our method discards pixels corresponding to instruments due to their inherent dynamic nature, while kidney stone pixels are incorporated into the SLAM framework but classified as potentially dynamic map points, allowing for their disappearance. This refinement enhances the robustness and the accuracy of ureteroscope position estimation for surgical navigation. To evaluate our approach, we recorded multiple datasets for both segmentation and ureteroscope pose estimation. Experimental results show an average improvement in ureteroscope pose estimation of 35.4% when using YOLOv7-E6E with SAM, and 52.49% when using YOLO11m-seg.
Monocular Metric Depth Estimation (MDE) in endoscopic images is a crucial step to improve navigation during medical procedures, as it enables the estimation of dense, real-scale 3D maps of the organs. For instance, in monocular flexible ureteroscopy (fURS), accurate navigation and real-scale information are essential for locating and removing kidney stones efficiently. Currently, the most promising approach to infer depth from single passive cameras is by supervised training of large neural networks, so-called foundation models for MDE. However, the depth output of these models is biased when the training data domain does not fit the goal domain (both camera and scene). At the same time, one of the greatest challenges in medical imaging is the lack of annotated datasets, as obtaining real ground-truth (e.g., depth data) is difficult. To overcome this, simulation has become a valuable tool in ureteroscopic imaging research. In this study, we introduce KidneyDepth, a synthetic dataset designed to reduce the gap between simulated and real-world 3D imaging. It includes a variety of shapes (e.g. mesh from CT scan, geometric primitive forms) along with different textures and lighting conditions, generated by BlenderProc2 [7]. To assess the effectiveness of KidneyDepth, we fine-tune two state-of-the-art MDE models (Depth Anything V2 and ZoeDepth) and test their performance on both simulated and real ureteroscopic images. Additionally, we evaluate the validity of their output by using the inferred depths in the context of a RGB-D SLAM system. Our results show that training models on a synthetic dataset with diverse structures and lighting conditions improves depth estimation in real endoscopic images and our simulations show that these RGB-D images enhance overall SLAM accuracy. The KidneyDepth dataset can be found at https://zenodo.org/records/14893421.
Metric depth estimation from visual sensors is crucial for robots to perceive, navigate, and interact with their environment. Traditional range imaging setups, such as stereo or structured light cameras, face hassles including calibration, occlusions, and hardware demands, with accuracy limited by the baseline between cameras. Single- and multi-view monocular depth offers a more compact alternative, but is constrained by the unobservability of the metric scale. Light field imaging provides a promising solution for estimating metric depth by using a unique lens configuration through a single device. However, its application to single-view dense metric depth is under-addressed mainly due to the technology's high cost, the lack of public benchmarks, and proprietary geometrical models and software. Our work explores the potential of focused plenoptic cameras for dense metric depth. We propose a novel pipeline that predicts metric depth from a single plenoptic camera shot by first generating a sparse metric point cloud using machine learning, which is then used to scale and align a dense relative depth map regressed by a foundation depth model, resulting in dense metric depth. To validate it, we curated the Light Field Stereo Image Dataset (LFS) of real-world light field images with stereo depth labels, filling a current gap in existing resources. Experimental results show that our pipeline produces accurate metric depth predictions, laying a solid groundwork for future research in this field.
Perceptual aliasing and weak textures pose significant challenges to the task of place recognition, hindering the performance of Simultaneous Localization and Mapping (SLAM) systems. This paper presents a novel model, called UMF (standing for Unifying Local and Global Multimodal Features) that 1) leverages multi-modality by cross-attention blocks between vision and LiDAR features, and 2) includes a re-ranking stage that re-orders based on local feature matching the top-k candidates retrieved using a global representation. Our experiments, particularly on sequences captured on a planetary-analogous environment, show that UMF outperforms significantly previous baselines in those challenging aliased environments. Since our work aims to enhance the reliability of SLAM in all situations, we also explore its performance on the widely used RobotCar dataset, for broader applicability. Code and models are available at https://github.com/DLR-RM/UMF
In ureteroscopy a flexible ureteroscope is used to inspect the kidneys and the ureters. In this context, we investigate robotic solutions that leverage recent developments from Simultaneous Localization and Mapping (SLAM) to provide a 3D map of the organ to be explored and the ureteroscope's pose. With this aid, the surgeoncan navigate through the organ more precisely. Additionally, for the final organ inspection, the risk of missing certain regions of the organ is minimized. In this paper, we propose a visual, monocular SLAM system based on ORB-SLAM3 that is able to estimate the pose of the ureteroscope's tip. In order to fulfill this task, we introduce two pre-processing steps. The first one aims at increasing the contrast of the image and the second one helps to avoid detecting features in non-desired regions (e.g. reflections). Additionally, we extend ORB-SLAM3 with A-KAZE and SuperPoint features and compare their performances to ORB features. The proposed method is evaluated in two experiments. The first experiment shows that we are able to estimate the trajectory of the ureteroscope and the map of a synthetic kidney with low errors. The second experiment shows that the method also gives promising results in cystoscopy.
Region-based methods have become increasingly popular for model-based, monocular 3D tracking of texture-less objects in cluttered scenes. However, while they achieve state-of-the-art results, most methods are computationally expensive, requiring significant resources to run in real-time. In the following, we build on our previous work and develop SRT3D , a sparse region-based approach to 3D object tracking that bridges this gap in efficiency. Our method considers image information sparsely along so-called correspondence lines that model the probability of the object’s contour location. We thereby improve on the current state of the art and introduce smoothed step functions that consider a defined global and local uncertainty. For the resulting probabilistic formulation, a thorough analysis is provided. Finally, we use a pre-rendered sparse viewpoint model to create a joint posterior probability for the object pose. The function is maximized using second-order Newton optimization with Tikhonov regularization. During the pose estimation, we differentiate between global and local optimization, using a novel approximation for the first-order derivative employed in the Newton method. In multiple experiments, we demonstrate that the resulting algorithm improves the current state of the art both in terms of runtime and quality, performing particularly well for noisy and cluttered images encountered in the real world.
Robots with elasticity in structural components can suffer from undesired end-effector positioning imprecision, which exceeds the accuracy requirements for successful manipulation. We present the Probabilistic-Product-Of-Exponentials robot model, a novel approach for kinematic modeling of robots. It does not only consider the robot's deterministic geometry but additionally models time-varying and configuration-dependent errors in a probabilistic way. Our robot model allows to propagate the errors along the kinematic chain and to compute their influence on the end-effector pose. We apply this model in the context of sensor fusion for manipulator pose correction for two different robotic systems. The results of a simulation study, as well as of an experiment, demonstrate that probabilistic configuration-dependent error modeling of the robot kinematics is crucial in improving pose estimation results.
We present a first-of-its-kind end-to-end tele-robotic VR system where the user operates a robot arm remotely, while being virtually immersed into the scene through force feedback and holographic vision. In contrast to stereoscopic head mounted displays that only provide depth perception to the user, the holographic vision device projects a light field, additionally allowing the user to correctly accommodate his/her eyes to the perceived depth of the scene's objects. The highly improved immersive user experience results in less fatigue in the tele-operator's daily work, creating safer and/or longer working conditions. The core technology relies on recent advances in immersive video coding for audio-visual transmission developed within the MPEG standardization committee. Virtual viewpoints are synthesized for the tele-operator's viewing direction from a couple of colour and depth fixed video feeds. Besides of the display hardware and its GPU-enabled view synthesis driver, the biggest challenge hides in obtaining high-quality and reliable depth images from low-cost depth sensing devices. Specialized depth refinement tools have been developed for running in realtime at zero delay within the end-to-end tele-robotic immersive video pipeline, which must remain interactive by essence. Various modules work asynchronously and efficiently at their own pace, with the acquisition devices typically being limited to 30 frames per second (fps), while the holographic headset updates its projected light field at up to 240 fps. Such modular approach ensures high genericity over a wide range of free navigation VR/XR applications, also beyond the tele-robotic one presented in this paper.
We propose a novel, highly efficient sparse approach to region-based 6DoF object tracking that requires only a monocular RGB camera and the 3D object model. The key contribution of our work is a probabilistic model that considers image information sparsely along correspondence lines. For the implementation, we provide a highly efficient discrete scale-space formulation. In addition, we derive a novel mathematical proof that shows that our proposed likelihood function follows a Gaussian distribution. Based on this information, we develop robust approximations for the derivatives of the log-likelihood that are used in a regularized Newton optimization. In multiple experiments, we show that our approach outperforms state-of-the-art region-based methods in terms of tracking success while being about one order of magnitude faster. The source code of our tracker is publicly available (https://github.com/DLR-RM/RBGT).
In this paper we present DOT (Dynamic Object Tracking), a front-end that added to existing SLAM systems can significantly improve their robustness and accuracy in highly dynamic environments. DOT combines instance segmentation and multi-view geometry to generate masks for dynamic objects in order to allow SLAM systems based on rigid scene models to avoid such image areas in their optimizations. To determine which objects are actually moving, DOT segments first instances of potentially dynamic objects and then, with the estimated camera motion, tracks such objects by minimizing the photometric reprojection error. This short-term tracking improves the accuracy of the segmentation with respect to other approaches. In the end, only actually dynamic masks are generated. We have evaluated DOT with ORB-SLAM 2 [1] in three public datasets. Our results show that our approach improves significantly the accuracy and robustness of ORB-SLAM 2, especially in highly dynamic scenes.
The MMX - Martian Moons eXploration - mission, as the name already suggests, aims to explore the two moons of Mars, Phobos and Deimos. The goal of this space mission led by JAXA is to acquire the scientific data necessary to understand the composition, structure, and history of these peculiar celestial bodies. The first man-made object to ever land on the larger and closer of these two moons, Phobos, shall be the small and lightweight MMX rover. The rover will be designed, manufactured and operated jointly by CNES and DLR. After separation from the carrier spacecraft, landing, uprighting, and deployment, this rover will start to drive in the low-gravity environment of Phobos surface and perform scientific operations. The rover will have a length of 44cm and a weight below 30kg. It will be solar-powered to operate for an intended mission duration of 90 days [1]. Communication round-trip times between Earth and Mars are already eight to forty minutes. For Phobos, we expect significantly higher values due to the need for relay satellites and limited communication windows. This leads to a requirement for a high level of autonomy for the robot, particularly its navigation capabilities, in order to maximize the scientific output of the mission. We will thus develop a navigation solution and integrate it as a software component running on the MMX rovers on-board computer. This navigation software will be verified in the introductory commissioning phase and is intended to be useful in the subsequent main operations phase. On the one hand, our design is inspired by the previous successful NASA planetary rovers, which have been the pioneers of this software technology. On the other hand, the design of the MMX rover brings its own sets of specifics and limitations to the table, and the uncharted celestial body Phobos itself is a source of several major and unprecedented challenges. For some of these, we can build upon the experience gained with the MASCOT mobile asteroid lander [2], which was deployed on the asteroid Ryugu in 2018 and successfully performed jumps for relocalization in the micro-gravity environment. The most notable challenges provided by the MMX rover design are the navigation cameras located at a fixed position and orientation w.r.t. the rovers body at a height of only 30cm above ground; the skid-steered locomotion of the rover limiting its turning speed; and restrictions on weight and power consumption further limiting the range of permissible operations. The most notable challenges provided by Phobos include not knowing the map of the terrain at the operational scale beforehand - the best maps available are at a resolution of 5m/px; the unknown soil composition; and the unknown local gravity. The combination of the rover and the celestial body is also a source of challenges: the behavior of wheels in contact with the soil is impossible to investigate beforehand, apart from software simulations, and Phobos very fast rotation - a Phobos day only lasts eight Earth hours - in the combination with the very slow maximum rover speed of approximately 4mm/s, makes shadows move relatively quickly, which could confuse visual odometry approaches. The on-board computer provides its own set of limitations and pre-requisites such as memory allocation and orchestration of concurrently running software processes. In our workshop contribution, we will identify and categorize challenges for navigation on Phobos and sketch our planned solutions to tackle them. Our planned navigation architecture will contain several FPGA and CPU-based modules: Dense depth data will be computed via Semi-Global Matching [3] on a FPGA. This is the basis for a stereo visual odometry such as [4] used to estimate the robot's trajectory. Further modules include an obstacle classification on individual depth images similar to [5] and possibly further mapping modules to create maps in compact representations to be sent to operators on Earth. Such obstacle and map information can then be used to realize autonomous emergency stop behavior up to future reactive obstacle avoidance or path planning modules to support (semi-)autonomous operation. These developments are based on experience we gained developing complex autonomous robotic navigation systems [6, 7] that we tested and evaluated in several field tests at Moon-analogue environments on the volcano Mt. Etna, Sicily, Italy [8, 9]. Some of the greatest challenges arise from the daring ambition to bring a rover into an environment where mankind has never been before, and expecting it to drive there to some extent autonomously. But that is also what makes this mission interesting in the first place. The scientific and technological discussions at the workshop may both help us to steer our decision making and enrich the scientific community with our findings. References: [1] J. Bertrand, et al., Roving on Phobos: Challenges of the MMX Rover for Space Robotics, ASTRA (2019) [2] J. Reill et al., MASCOT - Asteroid Lander with Innovative Mobility Mechanism, ASTRA (2015) [3] H. Hirschmuller, Stereo processing by semiglobal matching and mutual information, TPAMI (2007) [4] H. Hirschmuller, et al., Fast, unconstrained camera motion estimation from stereo without tracking and robust statistics, ICARCV (2002) [5] C. Brand, et al., Stereo-Vision Based Obstacle Mapping for Indoor / Outdoor SLAM, IROS (2014) [6] M. J. Schuster, et al., Towards Autonomous Planetary Exploration: The Lightweight Rover Unit (LRU), its Success in the SpaceBotCamp Challenge, and Beyond, JINT (2017) [7] M. J. Schuster, et al., Distributed stereo vision-based 6D localization and mapping for multi-robot teams, JFR (2018) [8] M. Vayugundla, et al., Datasets of Long Range Navigation Experiments in a Moon Analogue Environment on Mount Etna, ISR (2018) [9] A. Wedler, et al., First Results of the ROBEX Analogue Mission Campaign: Robotic Deployment of Seismic Networks for Future Lunar Missions, IAC (2017)
This work proves that semantic segmentation on minimally invasive surgical instruments can be improved by using training data that has been augmented through domain adaptation. The benefit of this method is twofold. Firstly, it suppresses the need of manually labeling thousands of images by transforming synthetic data into realistic-looking data. To achieve this, a CycleGAN model is used, which transforms a source dataset to approximate the domain distribution of a target dataset. Secondly, this newly generated data with perfect labels is utilized to train a semantic segmentation neural network, U-Net. This method shows generalization capabilities on data with variability regarding its rotation- position- and lighting conditions. Nevertheless, one of the caveats of this approach is that the model is unable to generalize well to other surgical instruments with a different shape from the one used for training. This is driven by the lack of a high variance in the geometric distribution of the training data. Future work will focus on making the model more scale-invariant and able to adapt to other types of surgical instruments previously unseen by the training.
This paper discusses the concept of plenoptic hand lens imagers for in-situ close-range imaging during planetary exploration missions. Hand lens imagers, such as the Mars Hand Lens Imager on-board the Mars rover Curiosity, are important cameras for in-situ investigations, e.g. of rock layer, minerals or dust. They are also important for the preparation and documentation of other instrument operations and for rover health assessment. Due to the small working distance between object and the camera's main lens, significant physical limitations affect the imaging performance. Most evident is the limited depth of field of a few millimeters for working distances of a few centimeters. This requires a highly accurate positioning of the camera and also limits the in-focus content of an image significantly. Hence, in order to have an extended object completely in focus, a sequence of images, each being focused to a different distance, is required. A single, passive camera is insufficient to compute depth from a single shot; only the combination of multiple images, either taken from different vantage points or at different focal settings, allows this. To overcome those limitations, we propose the use of plenoptic cameras as hand lens imagers. From a single exposure, they allow to create an extended depth of field image and at the same time a metric depth map while maintaining a more open aperture. These and other advantages might make it possible to omit space grade focus mechanisms in the future. A plenoptic camera is achieved by adding an additional matrix of lenslets shortly in front of the image sensor of a conventional camera. Hence, available space camera hardware can be used to form a new type of sensor. Due to its recording concept, a plenoptic camera maintains the depth of the scene as it is projected into the camera by the main lens. Thanks to the parallax between the lenslets, it is possible to compute depth via triangulation for each image point as well as a high resolution 2-D extended depth of field image. This paper provides an overview of the state of the art of hand lens imaging from which we derive a set of common requirements for future devices. We briefly introduce the plenoptic camera technology and provide first experimental results on the imaging performance based on samples of test targets and rocks. The results show that our preliminary plenoptic camera setup can comply with the requirements for in-situ hand lens imaging in terms of image quality, depth estimation and the usability for planetology.
This paper provides a detailed system theoretical model of a plenoptic camera with the aim to provide in-depth understanding of the plenoptic data recording concept and its effects. Plenoptic cameras, also known as light field cameras, were firstly thought of in the beginning of the 20th century and became recently possible thanks to rapid development of processing hardware and the increase of camera sensor resolution. Despite being a new type of sensor, they are operated in the same way as conventional cameras, but offer several advantages. A plenoptic camera consists of a main lens and a lenslet array (microlens array) right in front of the detector. The microlens array causes not only the recording of the incident location of a light ray on the sensor, as it is done by a conventional camera, but also the incident direction. Such a record can be represented by a 4-D data set known as the light field. In fact, by inserting a microlens array any conventional camera can be transformed into a plenoptic camera. The plenoptic recording concept and the 4-D light field provide multiple advantages over conventional cameras. For example, a single recorded light field allows first, to reconstruct novel views with small changes in viewpoint, second, to create a depth map, and third, to refocus images after the data capture. Hence, the process of focusing is shifted from hardware to software. Last, but not least, plenoptic cameras allow an extended depth of field in comparison to a conventional camera and the use of a bigger camera aperture. Most of the mentioned advantages become particularly effective at close-range to an object. The German Aerospace Center performs research on plenoptic cameras for close-range imaging in space. Possible applications are for example robot vision with plenoptic cameras for robotic arm operations during on-orbit servicing missions or the use of plenoptic cameras on rovers in the course of exploration missions to other planets. Those application scenarios and the demanding conditions in space require thorough comprehension of plenoptic cameras. For this purpose, this paper shall provide a detailed model of plenoptic cameras, which allows to derive camera parameters and optimize them with particular attention to the user requirements and to generate synthetic data. The latter can be utilized to assess the evaluation algorithms, which are not mentioned in detail in this paper. The modeling of the plenoptic camera is mainly based on the theory of geometric optics expanded by elements of diffraction optics.
This work deals with the passive tracking of the pose of a close-range 3-D modeling device using its own high-rate images in realtime, concurrently with customary 3-D modeling of the scene. This novel development makes it possible to abandon using inconvenient, expensive external trackers, achieving a portable and inexpensive solution. The approach comprises efficient tracking of natural features following the Active Matching paradigm, a frugal use of interleaved feature-based stereo triangulation, visual odometry using the robustified V-GPS algorithm, graph optimization by local bundle adjustment, appearance-based relocalization using a bank of parallel three-point-perspective pose solvers on SURF features, and online reconstruction of the scene in the form of textured triangle meshes to provide visual feedback to the user. Ideally, objects are completely digitized by browsing around the scene; in the event of closing the motion loop, a hybrid graph optimization takes place, which delivers highly accurate motion history to refine the whole 3-D model within a second. The method has been implemented on the DLR 3D-Modeler; demonstrations and abundant video material validate the approach. These types of low-cost systems have the potential to enhance traditional 3-D modeling and conquer new markets owing to their mobility, passivity, and accuracy. (C) 2018 Elsevier B.V. All rights reserved.
This work discusses the benefits of plenoptic cameras for future hand lens imagers for in-situ planetology. Such cameras offer advantages over conventional cameras, especially at the small working distance that are common for in-situ micro imaging. For example, the extension of the depth of field without a focus mechanism while maintaining a more open aperture at the same time. Additionally, textured depth maps as well as views with small perspective changes are possible, all from a single recorded image. We present a brief introduction of the plenoptic camera technology and examples from laboratory experiments.
In this paper we present the ability to intrinsically and extrinsically calibrate a Time-of-Flight sensor, namely, a Photonic Mixer Device (PMD) camera, using the DLR CalDe and DLR CalLab camera calibration toolbox. This camera is intended as a visual sensor for pose estimation in the close rendezvous phase during future On-Orbit servicing. In order to test and verify the pose estimation algorithms on the ground, we conduct different rendezvous scenarios using the European Proximity Operation Simulator. It is necessary to accurately know intrinsic parameters like the focal length, the principal point, and the distortion parameters, as well as the extrinsic parameters, i.e., the position and orientation of the PMD camera relating to the mounting board, whenever it is fixed on the robot and involved in the process of target pose estimation. In this work we differentiate from state-of-the-art approaches for the calibration of PMD cameras in this context by making use of the motion of the mounting robotic manipulator alone, i.e., without the need for accurate positioning of the target calibration plate by a second robotic manipulator.
This paper discusses the potential benefits of plenoptic cameras for robot vision during on-orbit servicing missions. Robot vision is essential for the accurate and reliable positioning of a robotic arm with millimeter accuracy during tasks such as grasping, inspection or repair that are performed in close range to a client satellite. Our discussion of the plenoptic camera technology provides an overview of the conceptional advantages for robot vision with regard to the conditions during an on-orbit servicing mission. A plenoptic camera, also known as light field camera, is basically a conventional camera system equipped with an additional array of lenslets, the micro lens array, at a distance of a few micrometers in front of the camera sensor. Due to the micro lens array it is possible to record not only the incidence location of a light ray but also its incidence direction on the sensor, resulting in a 4-D data set known as a light field. The 4-D light field allows to derive regular 2-D intensity images with a significantly extended depth of field compared to a conventional camera. This results in a set of advantages, such as software based refocusing or increased image quality in low light conditions due to recording with an optimal aperture while maintaining an extended depth of field. Additionally, the parallax between corresponding lenslets allows to derive 3-D depth images from the same light field and therefore to substitute a stereo vision system with a single camera. Given the conceptual advantages, we investigate what can be expected from plenoptic cameras during close range robotic operations in the course of an on-orbit servicing mission. This includes topics such as image quality, extension of the depth of field, 3-D depth map generation and low light capabilities. Our discussion is backed by image sequences for an on-orbit servicing scenario that were recorded in a representative laboratory environment with simulated in-orbit illumination conditions. We mounted a plenoptic camera on a robot arm and performed an approach trajectory from up to 2 m towards a full-scale satellite mockup. Using these images, we investigated how the light field processing performs, e.g. in terms of depth of field extension, image quality and depth estimation. We were also able to show the applicability of images derived from light fields for the purpose of the visual based pose estimation of a target point.