Spinal ablation and vertebroplasty are commonly used to treat vertebral tumours and any bone fractures they might cause. In this procedure, a hollow channel is created transpedicularly through which surgical tools of varying lengths are passed under fluoroscopy guidance. Electromagnetic (EM) tracking of the cannula that creates the channel is a helpful adjunct to fluoroscopy by providing real-time 3D spatial information, leading to the potential of reduced radiation exposure and surgical time. Although integration of EM tracked tools into spine procedures has been explored, calibration of variable-length tools, such as a bone access kit, remains challenging. We present a standardized calibration framework for "variable-length" needle-like tools, with the objective of integration into a co-registered fluoroscopy-EM system. We devised a template-based calibration process by registering a tracked tool shaft to the centerline of the indented line of a custom calibration block. The calibration was validated in three experiments against a commercial pre-calibrated needle functioning as ground truth. The cannula tip final position was < 1 mm from the ground truth needle tip across experiments. Further validation involved using our group's previously developed fluoroscopy-EM registration to reproject the EM-tracked calibrated cannula tip onto corresponding fluoroscopic images, producing a mean reprojection error of 5.8 px or 1.6 mm. Future work will involve improving the projection matrix calculation and assessing the variable-length tool reprojection in multiple views during a simulated spine intervention.
Renal ultrasound is routinely used in pediatric care due to its safety, accessibility, real-time visualization and effectiveness in assessing kidney anatomy. In particular, accurate segmentation of the kidney and its internal structures is essential for evaluating conditions such as hydronephrosis, a common pediatric disorder involving kidney swelling due to urine buildup. Automated segmentation provides a foundation for objective kidney volume estimation, which is critical for diagnosis, monitoring, and treatment planning. In this study, we benchmark a diverse suite of deep learning based segmentation models, including promptable models: a pretrained SAM2 model, a fine-tuned SAM2 model, and three pretrained MedSAM2 variants, as well as fully supervised models: nnU-Net, U-Net, DeepLabV3+, and SegResNet. We performed segmentation of key kidney structures: Capsule, Central Echo Complex, Renal Cortex, and Renal Medulla. We evaluate the models using Dice score, Hausdorff Distance and Average Symmetric Surface Distance to quantify accuracy and boundary precision. Our initial results show that SAM2 performs well for Capsule segmentation with our finetuned model having an average Dice score of 0.96, and nnU-Net demonstrates consistent, strong performance across all classes with average Dice scores of 0.83, 0.69 and 0.77 for Central Echo Complex, Cortex and Medulla, respectively. Clinical Relevance: This work demonstrates the utility of modern deep learning methods for detailed kidney structure segmentation in pediatric ultrasound. The results support their use in enabling automated kidney volume estimation, which could improve the assessment and management of hydronephrosis and enhance overall diagnostic efficiency in pediatric nephrology.
Accurate and robust three-dimensional 3D to two-dimensional 2D registration is crucial for augmented reality and image guidance in minimally invasive surgery. Conventional methods require precise camera calibration, which is challenging to achieve, especially for cameras with high optical zoom, significantly impacting the accuracy of 3D-2D registration algorithms. This paper proposes a novel method for 3D-2D surface registration that eliminates the need for camera calibration. Under small field of view, the method transforms the 3D-2D registration into an equivalent 2.5D-2.5D process, simplifying registration and improving efficiency. Moreover, by introducing a bias-correction technique to pseudo depth map estimation, the multimodal registration problem is converted into an approximate unimodal registration problem, enhancing robustness across various initial conditions and camera types. The method is validated through experiments on both a high-zoom surgical exoscope dataset with unknown camera parameters and a benchmark endoscopic dataset with known camera parameters. For exoscope experiments, the proposed method achieves a highly accurate registration with mean target registration error of 1.56 mm for rigid phantoms and 1.53 mm for deformable cadaver brains. In endoscope experiments, it achieves a 2D fiducial distance error of 1.95 mm, demonstrating efficacy across different scenarios and camera models.
Hydronephrosis is a painful condition characterized by dilation of the renal pelvis and calyces, commonly diagnosed with Ultrasound (US) or Computed Tomography (CT). Percutaneous nephrostomy (PN) is one of the standard treatments for this condition. PN, by nature, requires image-guided training, which is limited by the availability of realistic models. We present the design and development of an anthropomorphic kidney phantom with simulated hydronephrosis for multimodal imaging and puncture training. We developed a kidney phantom using polyvinyl alcohol cryogel (PVA-c) for tissue-mimicking material for its multimodality compatibility and realistic needle insertion tactile sensation. Anatomical geometry was derived from an open CT dataset, where segmentation of the kidney cortex and collective system was used to create 3D-printed moulds. A swollen collective system with refilling capabilities was simulated by adding a silicone balloon around the collective system dissolvable mould. Multimodality was validated using CT and US, confirming clear visualization of internal structures, including the kidney cortex and an anechoic swollen collective system, consistent with clinical hydronephrosis imaging. We present an anthropomorphic, multimodal, and refillable kidney phantom with hydronephrosis characteristics suitable for US-guided needle puncture training.
Cardiovascular diseases (CVDs) are the leading cause of death worldwide. While the most common technique for diagnosing CVDs is X-ray Angiography (XA), the presence of medical devices in angiograms can significantly obstruct the visibility of arterial abnormalities, reducing the efficacy and accuracy of diagnosis. Pacemakers are among the most common of these obstructive devices. Therefore, we propose a state of-the-art deep learning model capable of automatic digital removal of pacemakers from XA images. First, we generate a synthetic dataset of 25 000 training images with and without a pacemaker, using annotated images manually collected and then combined through a derived variation of the Beer-Lambert law. Next, we train our model on the synthetic dataset. To validate our approach, we evaluate our model on synthetic data as well as real data captured from an FDA-approved chest phantom. Despite being trained solely on synthetic data, our model yields an average structural similarity index of 0.98 on our synthetic test set, and 0.95 on our real-world test set. This approach marks the first fully automated solution to pacemaker removal from coronary angiograms. Our work has the potential to enhance the accuracy of CVD diagnoses by providing clearer coronary images, ultimately improving patient outcomes and supporting more effective clinical decision-making.
Percutaneous tumor cryoablation is an effective treatment for small renal cell carcinoma masses (<4.0cm). This procedure is conventionally performed under 2D ultrasound (US) guidance with pre- and intra-operative computed tomography (CT) for treatment planning and monitoring. 3D US systems can supplement the volumetric information provided by CT, allowing for accurate localization of structures without ionizing radiation. In 2D US images, the ice-ball generated during a cryoablation obscures structures within and beyond it. 3D US localization of the ice-ball margin could therefore benefit from merging of images at several positions and angles, necessitating accurate spatial tracking of the 3D images. Previous work has demonstrated the feasibility of 3D US in guidance, localization, and tumor coverage for percutaneous ablation of liver tumors. We implemented a two-part calibration framework to track an US transducer attached to a mechatronic arm. Using arm joint encoder readings collected in parallel with 3D points from an optical sensor attached to the transducer, we determined a transformation that maps the US transducer to the base of the arm. We then performed a calibration to localize the US image with respect to the optical sensor. This image calibration was formulated as a registration between paired image points and lines from a custom oriented line phantom. The arm calibration achieved a mean absolute error of 2.4 mm, and the image calibration produced a mean fiducial registration error of 1.4 mm. We successfully merged several combinations of US images, demonstrating the utility of the system for renal tumor cryoablation.
Accurate tracking of ultrasound (US) probes is essential in both clinical interventions and simulation based medical training. Applications such as freehand 3D US reconstruction, phantom scanning, and needle guidance all rely on precise 6D pose estimation of the US probe. Conventional tracking such as electromagnetic (EM) and optical systems offer high accuracy but are often expensive, sensitive to environmental interference, and require bulky fiducial markers that can obstruct clinical workflows. These limitations make them impractical for surgical settings and use in educational settings or low-resource environments. This study evaluates the use of deep learning based camera tracking for estimating the 6D pose of US probes in simulated surgical environment. Four 3D printed US probes were scanned under varying camera exposures and phantom studies, and their poses were estimated using two commodity RGB-D cameras, and data were validated against ground truth from magnetic tracking. Tracking methods from deep learning model MRC-Net was evaluated and is compared to assess whether camera-based tracking can meet clinical accuracy threshold. The fine-tuned MRC-Net model achieved a mean Euclidean error of 3.7 +/- 4.5 mm, with translation errors of 1.6 +/- 2.9 mm, 1.6 +/- 2.1 mm, and 2.3 +/- 3.3 mm along the x, y, and z axes, respectively. Inference rates of 6-24 frames per second across tested hardware demonstrate real-time feasibility. These results establish markerless deep learning tracking as a clinically relevant, scalable alternative for ultrasound simulation and training applications.
The 2D projective nature of X-ray radiography presents significant limitations in fluoroscopy-guided interventions, particularly the loss of depth perception and prolonged radiation exposure. Integrating magnetic trackers into these workflows is promising; however, it remains challenging and under-explored in current research and practice. To address this, we employed a radiolucent magnetic field generator (FG) prototype as a foundational step towards seamless magnetic tracking (MT) integration. A two-layer FG mounting frame was designed for compatibility with various C-arm X-ray systems, ensuring smooth installation and optimal tracking accuracy. To overcome technical challenges, including accurate C-arm pose estimation, robust fluoro-CT registration, and 3D navigation, we proposed the incorporation of external aluminum fiducials without disrupting conventional workflows. Experimental evaluation showed no clinically significant impact of the aluminum fiducials and the C-arm on MT accuracy. Our fluoro-CT registration demonstrated high accuracy (mean projection distance approxiamtely 0.7 mm, robustness (wide capture range), and generalizability across local and public datasets. In a phantom targeting experiment, needle insertion error was between 2 mm and 3 mm, with real-time guidance using enhanced 2D and 3D navigation. Overall, our results demonstrated the efficacy and clinical applicability of the MT-assisted approach. To the best of our knowledge, this is the first study to integrate a radiolucent FG into a fluoroscopy-guided workflow.
Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, producing misleading, unstable, or clinically irrelevant results. In a pioneering, consensus-driven effort, we introduce a comprehensive catalog of validation pitfalls in AI-based surgical video analysis that was derived from a multi-stage Delphi process with 92 international experts. The collected pitfalls span three categories: (1) data (e.g., incomplete annotation, spurious correlations), (2) metric selection and configuration (e.g., neglect of temporal stability, mismatch with clinical needs), and (3) aggregation and reporting (e.g., clinically uninformative aggregation, failure to account for frame dependencies in hierarchical data structures). A systematic review of surgical AI papers reveals that these pitfalls are widespread in current practice, with the majority of studies failing to account for temporal dynamics or hierarchical data structure, or relying on clinically uninformative metrics. Experiments on real surgical video datasets provide empirical evidence that ignoring temporal and hierarchical data structures can substantially understate uncertainty, obscure critical failure modes, and even alter algorithm rankings. To address these shortcomings, we provide a catalogue of best practices compiled in a multi-stage Delphi process. Together, this work provides an evidence-based framework to inform more rigorous validation of surgical video analysis algorithms and to guide future efforts in benchmarking, reporting, regulatory review, and clinical translation.
With the growing adoption of Competency-Based Medical Education (CBME) in surgical training, there is increasing demand for robust simulation platforms that provide realistic visualization for mitral valve repair procedures. While silicone-based surgical phantoms offer cost-effective training solutions, they lack the visual fidelity of real tissue. This paper presents a comprehensive framework for evaluating and improving the realism of phantom-based surgical simulators through neural style transfer. We introduce a novel patch-based consistency metric for assessing temporal stability in generated images and evaluate multiple GAN architectures including CycleGAN, MUNIT, and DRIT. Our approach is supported by a large-scale dataset comprising 59,411 frames constructed from in vivo surgical procedures and 6,625 phantom-based training images. Results demonstrate that the MUNIT architecture with ImageNet pre-training achieved superior perceptual similarity (FID: 16.3 +/- 3.2) compared to baseline simulator images (FID: 79.7 +/- 17.9), while our patch-based training methodology improved temporal consistency across architectures, reducing mean pixel differences by up to 23.5 %. These findings establish quantitative benchmarks for surgical simulation hyperrealism and provide insights into the fundamental trade-offs between perceptual quality and temporal consistency. Our framework and dataset provide valuable resources for advancing the development of more effective surgical training platforms.
Introduction Levator ani muscle (LAM) avulsion is a common traumatic injury of the pelvic floor muscle occurring during vaginal childbirth and is linked to the development of pelvic organ prolapse (POP). POP is a pelvic floor disorder that affects up to 40% of women during their lifetime. Pelvic floor ultrasound imaging is used to diagnose LAM avulsion, but it requires trained experts and is time-consuming, leading to weeks-long delays in receiving diagnostic results and treatment. The purpose of this study is to demonstrate the feasibility of a deep learning system to automatically classify the degree of LAM avulsion from 3D transperineal ultrasound images (TPUS). Methods 3D TPUS images of the pelvic floor from 150 patients with and without POP-related LAM avulsion were collected. Out of these, 113 patients were included in the study. Over 650 key slices were extracted from the ultrasound volumes and cropped to a region of interest. A two-stage cascading ensemble architecture was developed, combining three convolutional neural networks (MobileNetV3-Small, EfficientNet-B0, and RegNetY-800MF) with a final decision layer. The system performs hierarchical classification: first distinguishing between normal and avulsion cases, then determining unilateral versus bilateral involvement, and finally classifying the degree of avulsion. Results In 5-fold cross-validation, the ensemble model demonstrated strong performance in both binary classification tasks. For avulsion detection, it achieved 86% accuracy, 88% sensitivity, 85% specificity, and an AUC of 0.94, consistently outperforming individual base classifiers (which achieved AUCs of 0.71 to 0.75). For bilateral/unilateral classification, the model achieved 80% accuracy, 82% sensitivity, 78% specificity, and an AUC of 0.87. When evaluated on a test set for final patient-level classification across five classes (normal, complete bilateral avulsion, complete unilateral avulsion, partial bilateral avulsion, and partial unilateral avulsion), the system achieved 46% accuracy. Conclusion This study demonstrates the feasibility of a deep learning classification system to automatically classify the degree of LAM avulsion from 3D TPUS images. While the system showed promising performance in binary classification tasks, its sequential decision-making design means that a single slice misclassification in the first classification stage can impact the final patient-level accuracy. Additionally, the limited dataset size can hinder the model's ability to generalize effectively to unseen cases. Despite these limitations, the developed system shows the potential to expedite LAM avulsion diagnosis, overcoming the time constraints of manual diagnosis. This approach can broaden screening access, benefiting areas with limited healthcare resources, by reducing expert reliance and enabling timely treatment. Future work with larger and more diverse datasets could help address current limitations and further improve classification accuracy.
Many spinal operations are performed using fluoroscopic guidance due to its excellent visualization of osseous structures and surgical instrumentation in real-time, however, its efficacy is conditional on accurate needle placement. Image-guided surgical navigation systems allow for intraoperative and continuous localization of surgical tools with respect to patient anatomy, leading to significantly improved needle placement accuracy. Magnetic navigation systems require a field generator (FG) whose placement must be near the patient and may partially obstruct the x-ray beam, causing image artifacts and degraded image quality. Northern Digital Inc. has developed a radiolucent FG (RLFG) prototype to reduce image artifacts, however, the X-ray photon scatter interactions from the RLFG may reduce image contrast, add noise and decrease spatial resolution. These scatter interactions can be assessed in terms of the scatter-to-primary ratio (SPR) and its effect on image quality can be described using the modulation transfer function (MTF) and the generalized detective quantum efficiency (DQE). SPR measurements of a 20 cm water phantom and surgical table were taken with and without the RLFG using a slanted-edge technique as described by Garland and Cunningham, as well as the SPR measurements of the isolated RLFG and isolated water phantom. MTF and generalized DQE measurements of the imaging system were taken with and without the RLFG using the commercially available DQEPro (DQE Instruments, Ontario, Canada). SPR measurments demonstrated an 8% average increase when the RLFG was added underneath the surgical table, and the SPR of the water phantom was on average 5 times larger than the SPR of the RLFG. Therefore, the photon scatter interactions within the RLFG would likely cause minimal image quality deterioration, especially in comparison to a patient-representing water phantom. Introducing the RLFG in the imaging system demonstrates no practically significant difference in MTF, and a 9% average decrease in generalized DQE. The decreased DQE may be due in part to increased scatter on the exposure sensor relative to the image detector, and further experimentation is needed to validate this hypothesis. This work demonstrates the minimal effects on radiograph image quality with the introduction of a RLFG into a fluoroscopic imaging system, moving towards the seamless integration of magnetic tracking systems for fluoroscopy-guided interventions.
Percutaneous liver tumour ablation is becoming the preferred treatment option for patients ineligible for surgery. Percutaneous insertion of ablation applicator is often assisted by ultrasound (US) as it provides real-time visualization of subcutaneous anatomy and needle advancement. However, US-guided approach often requires multiple needle repositioning, leading to increased risk of tumour seeding and patient harm. Surgical Navigation Systems (SNS) have the potential to mitigate these limitations of the US-guided approach by providing additional visual and mechanical support. In this paper, we present a SNS comprising a mini stereotactic, patient-attached, mechanical needle guider with virtual reality (VR) visualization to facilitate needle positioning. We present a preliminary user trial comparing traditional US needle guidance puncture against our Mini-SNS. Nineteen (19) non-experts performed two randomized needle insertions into a designated tumour in a custom liver phantom. Preparation time, insertion time, total procedural time, needle repositioning, and final needle tip (tracked and virtual) position were recorded and analyzed. The Mini-SNS results show a reduction in insertion time compared to the traditional approach, without significant difference in the total procedure time. Compared to the traditional manual US-guided method, our SNS improves needle tip placement accuracy from 17mm to less than 10 mm. The number of repositioning was significantly reduced using the Mini-SNS, being just 3 in times total compared to an average of 8 times per user using the traditional approach. These initial findings suggest that the Mini-SNS enhances efficiency in needle punctures, with reduced insertion time and repositioning, supporting its potential to minimize patient harm.
Coronary artery bypass grafting (CABG) is one of the most commonly performed major cardiac surgeries in North America.1 The standard CABG surgical approach is performed on pump; however, advances in medical mechatronics and surgical techniques have made the option of performing CABG without putting the patient on bypass and stopping the heart to be a viable strategy. The off-pump method (OP-CABG) provides patients with superior long-term results and is now being preferred.(2) Due to the minimally-invasive and dynamic nature of the off-pump method, appropriate surgical training for OP-CABG is becoming increasingly important. However, current training apparatuses used by the medical industry and academia, along with various ethical and sourcing concerns, are lacking in the ability to replicate the motion of a beating heart, a clear requirement for an effective OP-CABG training environment. In this work, we propose the development of a mechanical OP-CABG surgical training device with disposable (i.e. replaceable) synthetic coronary arteries resting on a silicone heart surface. Cardiac movement is simulated via targeted and timed tensioning and releasing of multiple pull cords operated by a servo-motor-driven crankshaft. Using gated 4D CT, we demonstrate that our beating heart model mimics 3D heart motion. To our knowledge, this is the first mechanical OP-CABG simulator that replicates cardiac motion, providing a novel and effective tool for enhancing surgical training.
Liver tumour ablation procedures require accurate placement of the needle applicator at the tumour centroid. The lower-cost and real-time nature of ultrasound (US) has advantages over computed tomography for applicator guidance, however, in some patients, liver tumours may be occult on US and tumour mimics can make lesion identification challenging. Image registration techniques can aid in interpreting anatomical details and identifying tumours, but their clinical application has been hindered by the tradeoff between alignment accuracy and runtime performance, particularly when compensating for liver motion due to patient breathing or movement. Therefore, we propose a 2D-3D US registration approach to enable intra-procedural alignment that mitigates errors caused by liver motion. Specifically, our approach can correlate imbalanced 2D and 3D US image features and use continuous 6D rotation representations to enhance the model's training stability. The dataset was divided into 2388, 196, and 193 image pairs for training, validation and testing, respectively. Our approach achieved a mean Euclidean distance error of 2.28 m m ± 1.81 m m and a mean geodesic angular error of ± , with a runtime of 0.22 s per 2D-3D US image pair. These results demonstrate that our approach can achieve accurate alignment and clinically acceptable runtime, indicating potential for clinical translation.
3D ultrasound (US) imaging has shown significant benefits in enhancing the outcomes of percutaneous liver tumour ablation. Its clinical integration is crucial for transitioning 3D US into the therapeutic domain. However, challenges of tumour identification in US images continue to hinder its broader adoption. In this work, we propose a novel framework for integrating 3D US into the standard ablation workflow. We present a key component, a clinically viable 2D US–CT/MRI registration approach, leveraging 3D US as an intermediary to reduce registration complexity. To facilitate efficient verification of the registration workflow, we also propose an intuitive multimodal image visualization technique. In our study, 2D US–CT/MRI registration achieved a landmark distance error of ∼ 2–4 mm with a runtime of 0.22 s per image pair. Additionally, non-rigid registration reduced the mean alignment error by ∼ 40
The development of effective algorithms for removing surgical smoke in laparoscopic surgery has been hindered by the absence of a paired dataset containing real smoky and smoke-free surgical scenes. As a result, existing de-smoking methods have been primarily based on synthetic datasets and non-reference image enhancement metrics, which fail to fully capture the complexity of in vivo surgical scenes. To address this gap, we present a novel paired dataset derived from laparoscopic surgical recordings by identifying video sequences with relatively stationary scenes where smoke emerges. Our approach includes a robust motion-tracking technique that compensates for involuntary patient movements, ensuring reliable pairing of smoky images and their corresponding smoke-free ground truths. From 132 laparoscopic prostatectomy recordings, we curated 41 video sequences, resulting in a dataset of 2000 smoky-to-smoke-free image pairs. From 45 cholecystectomy recordings, we extracted 68 video sequences, resulting in an additional dataset of 1000 image pairs. Using this unique dataset, we evaluated a representative selection of current de-smoking methods, confirming their effectiveness while also highlighting their limitations. Furthermore, we critically revisited the commonly used atmospheric scattering model, atmospheric colour assumptions, and the dark channel prior. Our analysis demonstrated that the traditional atmospheric scattering model with “gray smoke” assumption introduces significant residual errors in the green and blue channels, while the dark channel prior maintains a strong correlation with smoke intensity. These observations suggest that, while less effective for direct smoke separation, the dark channel prior has potential to serve as a useful attention map for deep learning-based de-smoking approaches.
Purpose: In conventional fluoroscopy-guided interventions, the 2D projective nature of X-ray imaging limits depth perception and leads to prolonged radiation exposure. Virtual fluoroscopy, combined with spatially tracked surgical instruments, is a promising strategy to mitigate these limitations. While magnetic tracking shows unique advantages, particularly in tracking flexible instruments, it remains under-explored due to interference from ferromagnetic materials in the C-arm room. This work proposes a virtual fluoroscopy workflow by effectively integrating magnetic tracking, and demonstrates its clinical efficacy. Methods: An automatic virtual fluoroscopy workflow was developed using a radiolucent tabletop field generator prototype. Specifically, we developed a fluoro-CT registration approach with automatic 2D-3D shared landmark correspondence to establish the C-arm-patient relationship, along with a general C-arm modelling approach to calculate desired poses and generate corresponding virtual fluoroscopic images. Results: Testing on a dataset with views ranging from RAO 90 degrees to LAO 90 degrees, simulated fluoroscopic images showed visually imperceptible differences from the real ones, achieving a mean target projection distance error of 1.55 mm. An endoleak phantom insertion experiment highlighted the effectiveness of simulating multiplanar views with real-time instrument overlays, achieving a mean needle tip error of 3.42 mm. Conclusions: Results demonstrated the efficacy of virtual fluoroscopy integrated with magnetic tracking, improving depth perception during navigation. The broad capture range of virtual fluoroscopy showed promise in improving the users understanding of X-ray imaging principles, facilitating more efficient image acquisition.
Accurate 3D shape measurement is crucial for surgical support and alignment in robotic surgery systems. Stereo cameras in laparoscopes offer a potential solution; however, their accuracy in stereo image matching diminishes when the target image has few textures. Although stereo matching with deep learning has gained significant attention, supervised learning requires a large dataset of images with depth annotations, which are scarce for laparoscopes. Thus, there is a strong demand to explore alternative methods for depth reconstruction or annotation for laparoscopes. Active stereo techniques are a promising approach for achieving 3D reconstruction without textures. In this study, a 3D shape reconstruction method is proposed using an ultra-small patterned projector attached to a laparoscopic arm to address these issues. The pattern projector emits a structured light with a grid-like pattern that features node-wise modulation for positional encoding. To scan the target object, multiple images are taken while the projector is in motion, and the relative poses of the projector and a camera are auto-calibrated using a differential rendering technique. In the experiment, the proposed method is evaluated by performing 3D reconstruction using images obtained from a surgical robot and comparing the results with a ground-truth shape obtained from X-ray CT.