Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional fringe projection profilometry can achieve millimeter-scale accuracy but often requires multi-shot acquisition, digital-micromirror-device projection, and projector-camera synchronization, complicating integration into compact laparoscopic systems. Aim. To develop a synchronization-free, single-shot depth-sensing platform using a passive LED-illuminated binary mask and a VQ-VAE prior with a custom U-Net depth head. Approach. A compact projection module was coupled to one channel of a dual-channel laparoscope, while the second channel imaged the fringe-illuminated target. A Zivid 3D camera acquired reference depth for 722 paired phantom images. Zivid depth maps were reprojected into the SSLE image frame for supervised training and evaluation. The VQ-VAE encoded each input into a discrete latent representation, and a latent-space U-Net predicted depth without a separate mask-prediction branch. Results. Using a fixed train/validation/test split, the proposed model achieved an MAE of 3.70 mm, AbsRel of 0.0326, delta=1.1 accuracy of 0.962, and delta=1.1^2 accuracy of 0.970. It achieved lower MAE than the dual U-Net MaskNet + DepthNet baseline and outperformed off-the-shelf monocular depth models in MAE, AbsRel, and threshold accuracy. The pipeline operated at 26.0 Hz over 301 consecutive frames on an NVIDIA A100 GPU. Conclusions. The LED-illuminated binary-pattern platform with latent-space depth reconstruction enables synchronization-free, video-rate endoscopic depth estimation. Results demonstrate Zivid-referenced phantom reconstruction without an explicit segmentation stage, while emphasizing the importance of dataset size and SSLE-Zivid calibration accuracy.
Tagged MRI enables tracking internal tissue motion non-invasively. It encodes motion by modulating anatomy with periodic tags, which deforms along with tissue. However, the entanglement between anatomy, tags and motion poses significant challenges on post processing. Existence of tags and imaging blur hinders downstream tasks such as segmenting anatomy. Tag fading, due to T1-relaxation, disrupts brightness constancy assumption for motion tracking. For decades, these challenges are handled in isolation and sub-optimally. In contrast, we introduce a blind and nonlinear inverse framework for tagged MRI that, for the first time, unifies these tasks: anatomical image recovery, high-resolution cine image synthesis, and motion estimation. At its core, the synergy of MR physics and generative priors enables us to blindly estimate the unknown forward imaging models, high-resolution underlying anatomy, while simultaneously tracking 3D diffeomorphic Lagrangian motion over time. Experiments on tagged brain MRI demonstrate that our approach yields high-resolution anatomy images, cine images, and more accurate motion than specialized methods.
Normal Pressure Hydrocephalus (NPH) is a neurological disorder characterized by significant ventricular enlargement, potentially leading to cognitive impairment with gait and incontinence issues. Shunt implants can be used to drain excess cerebrospinal fluid (CSF) from the ventricles and can be an effective treatment in the correct cohorts. However, the presence of the shunt valve introduces artifacts in magnetic resonance images (MRI), which can bias downstream image processing pipelines. This bias hinders the accurate computation of biomarkers in evaluating the efficacy of shunt treatment. To address this, we propose VAID (Valve Artifact Inpainting with Diffusion models), a valve artifact removal method based on a fine-tuned 3D MRI diffusion model. The probabilistic representation learned by the diffusion model serves to provide realistic inpainting in the valve region of the shunt. Experiments with real pre- and post-shunt surgery MRI data demonstrate that our inpainting approach improves the accuracy of biomarker quantification in downstream neuroimaging pipelines such as SLANT and FreeSurfer.
Purpose:In clinical imaging, magnetic resonance (MR) image volumes are often acquired as stacks of 2D slices with decreased scan times, improved signal-to-noise ratio, and image contrasts unique to 2D MR pulse sequences. Although this is sufficient for clinical evaluation, automated algorithms designed for 3D analysis perform poorly on multislice 2D MR volumes, especially those with thick slices and gaps between slices. Superresolution (SR) methods aim to address this problem, but previous methods do not address all of the following: slice profile shape estimation, slice gap, domain shift, and noninteger or arbitrary upsampling factors. Approach:We propose ECLARE (Efficient Cross-planar Learning for Anisotropic Resolution Enhancement), a self-SR method that addresses each of these factors. ECLARE uses a slice profile estimated from the multislice 2D MR volume, trains a network to learn the mapping from low-resolution to high-resolution in-plane patches from the same volume, performs SR with antialiasing, and respects the image FOV during resampling. We compared ECLARE with cubic B-spline interpolation, SMORE, and other contemporary SR methods. We used realistic and representative simulations on human head MR volumes so that quantitative performance against ground truth can be computed. Specifically, healthy T 1 -w and people with MS T 2 -w FLAIR datasets were used for evaluations. We used the peak signal-to-noise ratio and structural similarity index measure as signal recovery metrics. We additionally used two independent brain parcellation algorithms, SLANT and SynthSeg, to compute the consistency Dice similarity coefficient and the R 2 coefficient of determination, respectively, as comparison metrics. Results:For images with up to 5 mm of slice thickness and up to 1.5 mm of gap, ECLARE achieves greater mean PSNR and SSIM compared with other methods. In representative regions of interest, such as the ventricles, caudate, cerebral white matter, and cerebellar white matter, ECLARE performs comparably or better than other approaches. These trends are similar for both investigated datasets. Conclusions:The use of slice profile estimation, FOV-aware resampling, and self-SR allowed ECLARE to robustly superresolve anisotropic images without the need for external training data. Future work will investigate the utility of ECLARE on other organs, species, modalities, and resolutions. Our code is open-source and available at https://www.github.com/sremedios/eclare.
Medical image challenges have played a transformative role in advancing the field, catalyzing innovation and establishing new performance benchmarks. Image registration, a foundational task in neuroimaging, has similarly advanced through the Learn2Reg initiative. Building on this, we introduce the Large-scale Unsupervised Brain MRI Image Registration (LUMIR) challenge, a next-generation benchmark for unsupervised brain MRI registration. Previous challenges relied upon anatomical label maps, however LUMIR provides 4,014 unlabeled T1-weighted MRIs for training, encouraging biologically plausible deformation modeling through self-supervision. Evaluation includes 590 in-domain test subjects and extensive zero-shot tasks across disease populations, imaging protocols, and species. Deep learning methods consistently achieved state-of-the-art performance and produced anatomically plausible, diffeomorphic deformation fields. They outperformed several leading optimization-based methods and remained robust to most domain shifts. These findings highlight the growing maturity of deep learning in neuroimaging registration and its potential to serve as a foundation model for general-purpose medical image registration.
Image harmonization (IH) makes images from different domains consistent, enabling multi-domain comparison with reliable quantitative measurements. Optical coherence tomography (OCT) could benefit from IH, as variations in OCT systems cause significant discrepancies in image quality due to speckle and noise. However, IH approaches typically rely on paired images during the training phase, and such images are not available in OCT data due to sparse image acquisition. A Schrödinger bridge (SB) finds an optimal coupling between arbitrary probability spaces with respect to a reference path measure and can be applied to unpaired IH. SBs are the core of many generative diffusion models, when one of the probability spaces is easy to sample from (i.e., Gaussian). SBs have been applied to natural images for image-to-image translation; however, they have not been applied to OCT IH because a SB does not guarantee anatomical consistency. In this paper, we use a dual diffusion implicit bridge (DDIB) for OCT IH, which finds independent SBs between arbitrary domains and a common Gaussian probability space. We generalize the assumption of DDIB and use segmentation masks as a condition to improve anatomical consistency. We conduct DDIB experiments with and without segmentation masks as a condition and analyze performance in terms of anatomical consistency and harmonization quality. The method does not require any paired training data and can in principle be quickly adapted to new domains.
Diffusion models are the current state-of-the-art for solving inverse problems in imaging. Their impressive generative capability allows them to approximate sampling from a prior distribution, which alongside a known likelihood function permits posterior sampling without retraining the model. While recent methods have made strides in advancing the accuracy of posterior sampling, the majority focuses on single-image inverse problems. However, for modalities such as magnetic resonance imaging (MRI), it is common to acquire multiple complementary measurements, each low-resolution along a different axis. In this work, we generalize common diffusion-based inverse single-image problem solvers for multi-image super-resolution (MISR) MRI. We show that the DPS likelihood correction allows an exactly-separable gradient decomposition across independently acquired measurements, enabling MISR without constructing a joint operator, modifying the diffusion model, or increasing network function evaluations. We derive MISR versions of DPS, DMAP, DPPS, and diffusion-based PnP/ADMM, and demonstrate substantial gains over SISR across 4×/8×/16× anisotropic degradations. Our results achieve state-of-the-art super-resolution of anisotropic MRI volumes and, critically, enable reconstruction of near-isotropic anatomy from routine 2D multi-slice acquisitions, which are otherwise highly degraded in orthogonal views.
Optical coherence tomography (OCT) is a non-invasive imaging technique that can visualize the retinal layers in the human macula. Deep learning algorithms segment these layers, from which mean thickness values are computed for each retinal layer in different regions of the macula. However, artifacts in OCT volumes can mislead segmentation results, in turn corrupting corresponding thickness measures. To establish an accurate database of retinal layer thicknesses, we propose statistical iterative data truncation (SIDT), an algorithm to perform automated data cleaning to generate normative measures on large datasets. We apply this method to more than 170, 000 OCT volumes from the UK Biobank database, inspecting thickness values for nine retinal layers in 14 macular regions of interest. The constructed database of normative thickness values is distributable.
Tagged magnetic resonance imaging (tMRI) is a valuable tool for visualizing and quantifying tissue deformation in vivo. Its use is often hampered, however, by tag fading, long computation times, and the challenge of ensuring diffeomorphic, incompressible motion fields. In this paper, we describe a novel integration of the harmonic phase (HARP) approach to tMRI analysis with an unsupervised deep learning-based registration framework to estimate 2D and 3D motion fields that are diffeomorphic and nearly incompressible. The resulting method, called deep sinusoidally transformed HARP, or DSHARP, enables end-to-end network training by implementing a transformation of the harmonic phase to remove phase-wrapping discontinuities. It produces diffeomorphic motion by estimating a stationary velocity field from which motion is computed using the scaling and squaring technique. Finally, it encourages incompressibility using a novel Jacobian determinant loss term during network training. We evaluated DSHARP on 2D and 3D phantom data with simulated incompressible motions, real 3D human tongue data acquired during speech from both healthy and glossectomy subjects, and cardiac tagged MRI from the public STACOM 2011 benchmark. Our approach outperforms HARP, SinMod, SyN, PVIRA, VoxelMorph, and DeepTag in tracking accuracy, computation speed, and preservation of incompressibility.
Fringe projection profilometry (FPP) traditionally requires multi-frequency and multi-phase sinusoidal patterns and precise projector-camera synchronization, limiting real-time performance in dynamic surgical scenes. We present a compact single-shot FPP system that employs a single static binary pattern projected through the illumination channel of a dual-channel laparoscope. Depth is reconstructed from a single captured frame using a two-stage deep learning pipeline consisting of MaskNet for tissue segmentation and DepthNet for dense depth estimation. Training leverages a combination of FPP-derived ground truth and a large-scale physically based synthetic dataset generated in Blender, modeling tissue geometry, illumination, and sensor noise. Prior studies have demonstrated millimeter-level depth accuracy and improved robustness of binary patterns to motion blur and optical misalignment compared to multi-shot sinusoidal methods. In this work, we focus on the system-level design and translation of binary-pattern single-shot FPP into a compact, synchronization-free endoscopic platform, highlighting its feasibility for real-time robotic surgical guidance.
Unique identification of multiple sclerosis (MS) white matter lesions (WMLs) is important to help characterize MS progression. WMLs are routinely identified from magnetic resonance images (MRIs) but the resultant total lesion load does not correlate well with EDSS; whereas mean unique lesion volume has been shown to correlate with EDSS. Our approach builds on prior work by incorporating Hessian matrix computation from lesion probability maps before using the random walker algorithm to estimate the volume of each unique lesion. Synthetic images demonstrate our ability to accurately count the number of lesions present. The takeaways, are: 1) that our method correctly identifies all lesions including many that are missed by previous methods; 2) we can better separate confluent lesions; and 3) we can accurately capture the total volume of WMLs in a given probability map. This work will allow new more meaningful statistics to be computed from WMLs in brain MRIs
Magnetic resonance (MR) tagging is an imaging technique for noninvasively tracking tissue motion in vivo by creating a visible pattern of magnetization saturation (tags) that deforms with the tissue. Due to longitudinal relaxation and progression to steady-state, the tags and tissue brightnesses change over time, which makes tracking with optical flow methods error-prone. Although Fourier methods can alleviate these problems, they are also sensitive to brightness changes as well as spectral spreading due to motion. To address these problems, we introduce the brightness-invariant tracking estimation (BRITE) technique for tagged MRI. BRITE disentangles the anatomy from the tag pattern in the observed tagged image sequence and simultaneously estimates the Lagrangian motion. The inherent ill-posedness of this problem is addressed by leveraging the expressive power of denoising diffusion probabilistic models to represent the probabilistic distribution of the underlying anatomy and the flexibility of physics-informed neural networks to estimate biologically-plausible motion. A set of tagged MR images of a gel phantom was acquired with various tag periods and imaging flip angles to demonstrate the impact of brightness variations and to validate our method. The results show that BRITE achieves more accurate motion and strain estimates as compared to other state of the art methods, while also being resistant to tag fading.
Magnetic resonance (MR) images are often acquired as anisotropic volumes in clinical settings. Such volumes have a worse throughplane resolution than in-plane resolution, hampering results in many processing pipelines that expect isotropic resolutions. Super-resolution (SR) is a promising methodology to address this problem, but there is concern whether the estimated high-resolution (HR) image suffers from egregious hallucinations, especially with deep learning methods that produce aesthetically pleasing results. One approach to restrict the impact of hallucinations is to guarantee that the estimated HR image is exactly cycle-consistent with the low-resolution observation. The denoising diffusion null space model (DDNM) achieves this through a range null space decomposition, but the specific design of the forward map is left to the application. In this work, we analyze the forward problem in 2D MR acquisition and construct an appropriate linear map A. We train a denoising diffusion probabilistic model on T1-weighted (T1-w) head MR images from multiple datasets and implement DDNM using A for the SR task. We show that the approach yields exact cycle-consistent solutions that are also realistic. We evaluated the approach in a wide variety of T1w MR datasets, including withheld subjects from training sites and two sites outside of the training domain. We achieve excellent qualitative and quantitative results according to both distortion and perceptual metrics.
Medical image harmonization aims to reduce the differences in appearance caused by scanner hardware variations to allow for consistent and reliable comparisons across devices. Harmonization based on paired images from different devices has limited applicability in real-world clinical settings. On the other hand, unpaired harmonization typically does not guarantee anatomy consistency, which is problematic because anatomical information preservation is paramount. The Schrödinger bridge framework has achieved state-of-the-art style transfer performance with natural images by matching distributions of unpaired images, but this approach can also introduce anatomy changes when applied to medical images. We show that such changes occur because the Schrödinger bridge uses the square of the Euclidean distance between images as the transport cost in an entropy-regularized optimal transport problem. Such a transport cost is not appropriate for measuring anatomical distances, as medical images with the same anatomy need not have a small Euclidean distance between them. In this paper, we propose a latent metric Schrödinger bridge (LMSB) framework to improve the anatomical consistency for the harmonization of medical images. We develop an invertible network that maps medical images into a latent Euclidean metric space where the distances among images with the same anatomy are minimized using the pullback latent metric. Within this latent space, we train a Schrödinger bridge to match distributions. We show that the proposed LMSB is superior to the direct application of a Schrödinger bridge to harmonize optical coherence tomography (OCT) images.
Spatially varying regularization accommodates the deformation variations that may be necessary for different anatomical regions during deformable image registration. Historically, optimization-based registration models have harnessed spatially varying regularization to address anatomical subtleties. However, most modern deep learning-based models tend to gravitate towards spatially invariant regularization, wherein a homogenous regularization strength is applied across the entire image, potentially disregarding localized variations. In this paper, we propose a hierarchical probabilistic model that integrates a prior distribution on the deformation regularization strength, enabling the end-to-end learning of a spatially varying deformation regularizer directly from the data. The proposed method is straightforward to implement and easily integrates with various registration network architectures. Additionally, automatic tuning of hyperparameters is achieved through Bayesian optimization, allowing efficient identification of optimal hyperparameters for any given registration task. Comprehensive evaluations on publicly available datasets demonstrate that the proposed method significantly improves registration performance and enhances the interpretability of deep learning-based registration, all while maintaining smooth deformations. Our code is freely available at http://bit.ly/3BrXGxz.
Recent advances in deep learning-based medical image registration have shown that training deep neural networks~(DNNs) does not necessarily require medical images. Previous work showed that DNNs trained on randomly generated images with carefully designed noise and contrast properties can still generalize well to unseen medical data. Building on this insight, we propose using registration between random images as a proxy task for pretraining a foundation model for image registration. Empirical results show that our pretraining strategy improves registration accuracy, reduces the amount of domain-specific data needed to achieve competitive performance, and accelerates convergence during downstream training, thereby enhancing computational efficiency.
In recent years, unsupervised learning for deformable image registration has been a major research focus. This approach involves training a registration network using pairs of moving and fixed images, along with a loss function that combines an image similarity measure and deformation regularization. For multi-modal image registration tasks, the correlation ratio has been a widely-used image similarity measure historically, yet it has been underexplored in current deep learning methods. Here, we propose a differentiable correlation ratio to use as a loss function for learning-based multi-modal deformable image registration. This approach extends the traditionally non-differentiable implementation of the correlation ratio by using the Parzen windowing approximation, enabling backpropagation with deep neural networks. We validated the proposed correlation ratio on a multi-modal neuroimaging dataset. In addition, we established a Bayesian training framework to study how the trade-off between the deformation regularizer and similarity measures, including mutual information and our proposed correlation ratio, affects the registration performance.
Optical coherence tomography (OCT) images are often acquired as highly anisotropic volumes, where the scanning step is dense along the fast axis but sparse along the slow axis. This affects image analysis, such as image registration for longitudinal alignment. To create more isotropic volumes, bicubic interpolation can be used along the slow axis, but it generally produces blurry features. Registration-based interpolation can reduce blurriness, but often fails to generate realistic OCT images. Deep generative models can sample realistic images, but lack the structural consistency constraints required for interpolation. In this paper, we propose an unsupervised image interpolation method that combines registration-based interpolation with a deep generative model to overcome their individual limitations and improve the structural accuracy and realism of interpolated OCT images. We compare the proposed method with both bicubic and registration-based interpolation on real OCT datasets, and show that it achieves the best interpolation performance.
Purpose:Regenerative therapies for retinal diseases include cell and gene therapy modalities that are targeted to the subretinal space. Several recent clinical trials have shown that the morbidity of surgical access is the major limitation of safe subretinal space delivery. We aimed to develop an image-guided procedure for minimally invasive subretinal access (MISA) as a platform to deliver therapeutic agents for the treatment of degenerative retinal diseases. Methods:We engineered prototypes of a novel common-path swept source optical coherence tomography (CP-SSOCT)-enabled needle, coaxial guide (COG), and subretinal access cannula (SAC). We pilot tested the MISA procedure in ex vivo bovine eyes and in vivo porcine ocular surgery. Results:A- and M-mode scan recordings of ex vivo and in vivo animal eye models demonstrated that CP-SSOCT imaging from the scleral side ( ab externo ) was capable of identifying the retinal laminae and the sub-retinal space. We show results from in vivo porcine MISA surgeries (N=4) using the novel CP-SSOCT-enabled sub-retinal injection needle, COG, and SAC through the transscleral approach. The MISA approach enabled subretinal device placement in the posterior pole, however, cases of retinal incarceration and retinal perforation were encountered. Conclusions:We describe a novel CP-SSOCT-guided subretinal access approach that, with further optimization, may be useful in regenerative retinal surgery.
Objective: Deep learning-based deformable image registration has achieved strong accuracy, but remains sensitive to variations in input image characteristics such as artifacts, field-of-view mismatch, or modality difference. We aim to develop a general training paradigm that improves the robustness and generalizability of registration networks. Methods: We introduce surrogate supervision, which decouples the input domain from the supervision domain by applying estimated spatial transformations to surrogate images. This allows training on heterogeneous inputs while ensuring supervision is computed in domains where similarity is well defined. We evaluate the framework through three representative applications: artifact-robust brain MR registration, mask-agnostic lung CT registration, and multi-modal MR registration. Results: Across tasks, surrogate supervision demonstrated strong resilience to input variations including inhomogeneity field, inconsistent field-of-view, and modality differences, while maintaining high performance on well-curated data. Conclusions: Surrogate supervision provides a principled framework for training robust and generalizable deep learning-based registration models without increasing complexity. Significance: Surrogate supervision offers a practical pathway to more robust and generalizable medical image registration, enabling broader applicability in diverse biomedical imaging scenarios.