Recent advances in volumetric super-resolution (SR) have demonstrated great performance in medical and scientific imaging, with transformer- and CNN-based approaches achieving impressive results even at extreme scaling factors. We show that this impressive performance largely stems from training on downsampled data rather than real low-resolution scans. Such a training setup arises partly from the limited availability of paired high- and low-resolution volumetric datasets. To address this gap, we introduce VoDaSuRe, a large-scale volumetric dataset containing paired high- and low-resolution scans. When training models on VoDaSuRe, we reveal a significant discrepancy: models trained on downscaled data produce substantially sharper predictions than those trained on real low-resolution scans, which smooth fine structures. Conversely, applying downscaled trained models to real scans preserves more structure but is inaccurate. Our findings suggest that current SR methods are overstated - when applied to real data, they do not recover structures lost in low-resolution scans but instead predict a smoothed average. We argue that progress in deep learning-based volumetric SR requires datasets with paired real scans of high complexity, such as VoDaSuRe. Our dataset and code are publicly available at link when published.
The accurate characterization of the microstructure of fiber-reinforced polymer composites is crucial for quality control in manufacturing, material design, and robust performance prediction. Most of the characterization methods rely on the analysis of 2D or 3D images. However, the composites community lacks commonly shared best practices and a clear quantitative understanding of how different image processing approaches compare. This work presents a benchmarking exercise on image processing of fiber-reinforced composite materials to address this challenge. Processing three cross-section micrographs, acquired from a single unidirectional composite sample using typical yet distinct imaging protocols, 11 participants from 8 research institutions extracted fiber centroids and radii. The estimated fiber volume fraction (Vf) values for a single cross-section varied between participants from 0.42 to 0.65. The results highlight the sensitivity of different methods to factors like illumination inhomogeneities, pixel density, polishing-related surface defects, and fiber packing. Considering these observations, several best practices for image analysis are identified, providing critical insight into the methods’ sensitivities. To promote transparency and community uptake, the benchmark dataset, evaluation scripts, documentation, and algorithms that were open-sourced are made openly available through a dedicated website, paving the way toward more reliable microstructural analysis in composite materials and ultimately more standardized protocols.
Geo-spatial analysis of our world benefit from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, and geographic coordinates). Current geo-spatial benchmarks have limited coverage across modalities, considerably restricting progress in the field, as current approaches cannot integrate all relevant modalities within a unified framework. We introduce the Multi-Modal Landmark dataset (MMLandmarks), a benchmark composed of four modalities: 197k high-resolution aerial images, 329k ground-view images, textual information, and geographic coordinates from 18,557 distinct landmarks in the United States. The MMLandmarks dataset has a one-to-one correspondence across every modality, which enables training and benchmarking models on various geo-spatial tasks, including cross-view Ground-to-Satellite retrieval, ground and satellite geolocalization, Text-to-Image, and Text-to-GPS retrieval.We demonstrate broad generalization and competitive performance against off-the-shelf foundational models and specialized state-of-the-art models across different tasks by employing a simple CLIP-inspired baseline, illustrating the necessity for multimodal datasets to achieve broad geo-spatial understanding. The dataset, labels, and code will be released upon acceptance.
Retinal vessel segmentation is a crucial yet challenging task for early diagnosis of diseases such as diabetes and hypertension. SA-UNet achieves promising performance by introducing spatial attention in the bottleneck, but it underuses attention in skip connections and fails to address the severe foreground-background imbalance. We propose SA-UNetv2, a lightweight model that injects cross-scale spatial attention into all skip connections to strengthen multi-scale feature fusion and adopts a weighted Binary Cross-Entropy (BCE) + Matthews Correlation Coefficient (MCC) loss to improve robustness to class imbalance. On the public DRIVE and STARE datasets, SA-UNetv2 achieves state-of-the-art performance with only $\mathbf{1. 2 M B}$ memory and $\mathbf{0. 2 6 M}$ parameters (less than 50% of SA-UNet), and $\mathbf{1}$ second CPU inference on $592 \times 592 \times 3$ images, demonstrating strong efficiency and deployability in resource-constrained, CPU-only settings. The code is available at github.com/clguo/SA-UNetv2.
Computing a minimum s-t cut in a graph is a solution to a wide range of computer vision problems, and is often done using the Boykov-Kolmogorov (BK) algorithm. In this paper, we revisit the BK algorithm from both a theoretical and practical point of view. We improve the analysis of the time complexity of the BK algorithm to O(mn|C|) and propose a new algorithm, the fast and compact BK (fcBK) algorithm, with a time complexity of O(m|C|), where m, n, and |C| are the number of edges, number of vertices, and the capacity of the cut, respectively. We additionally propose a compact graph representation that allows our implementation to find a minimum s-t cut in a graph with upwards of 10^9 vertices and 10^10 edges on a machine with 128 GB of memory. We find our implementation of the BK algorithm to be the fastest available implementation of the BK algorithm when evaluating on a comprehensive set of benchmark datasets, highlighting the importance of memory-efficient implementations. We make our implementations publicly available for further research and implementation development within minimum s-t cut algorithms.
Microscopic imaging techniques (2D and 3D) are widely employed to investigate structured food systems. However, the practical limitations of these imaging methods often restrict analyses to a small number of samples, which may be acquired under loosely defined imaging conditions and parameters. Assessing the representativeness of these measurements is further complicated by the fundamental properties of food matrices, such as their inherent heterogeneity, disordered nature, and microstructural complexity. To address this in a practical setting, a comprehensive dataset of high-resolution synchrotron X-ray tomography scans of Mozzarella cheese is analyzed. Representative Elementary Volume (REV) analysis is applied to key structural descriptors — such as anisotropy, width, and orientation — to determine the volume and resolution thresholds required for reliable local characterization. Additionally, macroscale heterogeneity is quantified by evaluating descriptor variability in samples from the same Mozzarella cheese formulation, followed by comparison to inter-cheese distances in descriptor space. These findings offer methodological guidance for designing reliable imaging protocols not only for Mozzarella but also for other structurally similar food matrices, supporting broader adoption of image-based structural measurements in both research and industrial applications.
Wall shear stress (WSS) plays a crucial role in the initiation and progression of atherosclerosis. However, its noninvasive quantification remains challenging due to the limited spatiotemporal resolution and scalability of MRI, as well as the limited precision (±30% error range) and image quality trade-off of existing ultrasound-based approaches. This study presents an ultrasound-based wall shear stress imaging (WASHI) framework that simultaneously provides high-quality B-mode images and spatiotemporally resolved WSS maps, and evaluates its accuracy and precision. WASHI derives WSS directly from velocity gradients obtained using transverse oscillation vector flow imaging based on an interleaved synthetic aperture imaging sequence implemented on a Verasonics Vantage 256. Accuracy was assessed using a flow-rig setup under controlled conditions, and in vivo precision was evaluated in ten healthy volunteers through bilateral scans of the common carotid arteries (19 CCAs in total). In the flow-rig experiments,WASHI produced consistent WSS estimates across 0° and 20° tilting angles with accuracy comparable to a velocity-based estimator (bias: -1 mPa vs. -3 mPa). In in vivo measurements, WASHI successfully tracked vessel wall motion over multiple cardiac cycles and resolved both spatial and temporal variations in WSS along both vessel walls. The median coefficient of variation (CV) across the 19 CCAs was 10.2%, demonstrating high measurement precision. Although this precision was slightly lower than that of the velocity-based estimator (CV: 6.2%), WASHI produced WSS magnitudes (2.7 & 2.3 Pa) that closely reflected the captured flow profile (2.9 & 2.4 Pa). In addition, the intrasubject variability of time-averaged WSS was the same across participants (p = 0.45), indicating reproducible performance. These results demonstrate the feasibility of WASHI for precise and reproducible WSS imaging, enabling future longitudinal studies and large-cohort investigations of vascular hemodynamics.
Standard deep learning models for image segmentation cannot guarantee topology accuracy, failing to preserve the correct number of connected components or structures. This, in turn, affects the quality of the segmentations and compromises the reliability of the subsequent quantification analyses. Previous works have proposed to enhance topology accuracy with specialized frameworks, architectures, and loss functions. However, these methods are often cumbersome to integrate into existing training pipelines, they are computationally very expensive, or they are restricted to structures with tubular morphology. We present SCNP, an efficient method that improves topology accuracy by penalizing the logits with their poorest-classified neighbor, forcing the model to improve the prediction at the pixels' neighbors before allowing it to improve the pixels themselves. We show the effectiveness of SCNP across 13 datasets, covering different structure morphologies and image modalities, and integrate it into three frameworks for semantic and instance segmentation. Additionally, we show that SCNP can be integrated into several loss functions, making them improve topology accuracy. Our code can be found at https://github.com/Anonymous.
Achieving physically consistent image editing remains a significant challenge in computer vision. Existing image editing methods typically rely on neural networks, which struggle to accurately handle shadows and refractions. Conversely, physics-based inverse rendering often requires multi-view optimization, limiting its practicality in single-image scenarios. In this paper, we propose Materialist, a neural-initialized physically based rendering pipeline for single-image inverse rendering. Unlike previous hybrid methods that use physics to guide neural generation, our method leverages neural networks to predict initial material properties, which are then rigorously optimized via progressive differentiable rendering. Our approach enables a range of applications, including material editing, object insertion, and relighting, while also introducing an effective method for editing material transparency via ray-traced refraction without requiring full scene geometry. Furthermore, our envmap estimation method also achieves competitive performance, further enhancing the accuracy of image editing task. Experiments demonstrate strong performance across synthetic and real-world datasets, excelling even on challenging out-of-domain images.
Visual counterfactual explanations aim to change classifier decisions through realistic and localized edits while preserving decision-irrelevant content. Existing DDPM-based methods typically perform classifier-guided editing along a long reverse denoising trajectory. The changing noise levels make semantic editability and spatial control difficult to balance, and the editable state is noisy, whereas the target classifier is trained on clean images. As a result, these methods require either costly recursive denoising or low-quality one-step estimates to obtain classifier-facing clean images. We propose FiRe, a Fixed-noise Refinement framework for visual counterfactual explanations. Rather than following a reverse denoising trajectory, FiRe maps the input to a fixed noise level and iteratively refines the noisy state at that level. To provide clean images for classifier guidance, FiRe first adapts Pixel Mean Flow to visual counterfactual explanation, enabling direct clean-image prediction from noisy states. To make fixed-noise refinement produce minimal and localized counterfactual edits, FiRe introduces three FiRe-specific controls: a dynamic dual-mask strategy, adaptive guidance, and early stopping, which determine where edits accumulate, which changes become visible, and when refinement stops. Experiments on five tasks across three datasets show that, compared with the strongest recent baseline, FiRe achieves about 3× faster online inference and 8× fewer FLOPs while obtaining comparable or state-of-the-art counterfactual quality.
This paper introduces a diffusion-based framework for universal image segmentation, making agnostic segmentation possible without depending on mask-based frameworks and instead predicting the full segmentation in a holistic manner. We present several key adaptations to diffusion models, which are important in this discrete setting. Notably, we show that a location-aware palette with our 2D gray code ordering improves performance. Adding a final tanh activation function is crucial for discrete data. On optimizing diffusion parameters, the sigmoid loss weighting consistently outperforms alternatives, regardless of the prediction type used, and we settle on x-prediction. While our current model does not yet surpass leading mask-based architectures, it narrows the performance gap and introduces unique capabilities, such as principled ambiguity modeling, that these models lack. All models were trained from scratch, and we believe that combining our proposed improvements with large-scale pretraining or promptable conditioning could lead to competitive models.
Uncertainty quantification (UQ) is crucial in safety-critical applications such as medical image segmentation. Total uncertainty is typically decomposed into data-related aleatoric uncertainty (AU) and model-related epistemic uncertainty (EU). Many methods exist for modeling AU (such as Probabilistic UNet, Diffusion) and EU (such as ensembles, MC Dropout), but it is unclear how they interact when combined. Additionally, recent work has revealed substantial entanglement between AU and EU, undermining the interpretability and practical usefulness of the decomposition. We present a comprehensive empirical study covering a broad range of AU-EU model combinations, propose a metric to quantify uncertainty entanglement, and evaluate both across downstream UQ tasks. For out-of-distribution detection, ensembles exhibit consistently lower entanglement and superior performance. For ambiguity modeling and calibration the best models are dataset-dependent, with softmax/SSN-based methods performing well and Probabilistic UNets being less entangled. A softmax ensemble fares remarkably well on all tasks. Finally, we analyze potential sources of uncertainty entanglement and outline directions for mitigating this effect.
Spectral imaging data acquired via multispectral and hyperspectral cameras can have hundreds of channels, where each channel records the reflectance at a specific wavelength and bandwidth. Time and resource constraints limit our ability to collect large spectral datasets, making it difficult to build and train predictive models from scratch. In the RGB domain, we can often alleviate some of the limitations of smaller datasets by using pretrained foundational models as a starting point. However, most existing foundation models are pretrained on large datasets of 3-channel RGB images, severely limiting their effectiveness when used with spectral imaging data. The few spectral foundation models that do exist usually have one of two limitations: (1) they are built and trained only on remote sensing data limiting their application in proximal spectral imaging, (2) they utilize the more widely available multispectral imaging datasets with less than 15 channels restricting their use with hundred-channel hyperspectral images. To alleviate these issues, we propose a large-scale foundational model and dataset built upon the masked autoencoder architecture that takes advantage of spectral channel encoding, spatial-spectral masking and ImageNet pretraining for an adaptable and robust model for downstream spectral imaging tasks.
Controlling illumination during video post-production is a crucial yet elusive goal in computational photography. Existing methods often lack flexibility, restricting users to certain relighting models. This paper introduces ReLumix, a novel framework that decouples the relighting algorithm from temporal synthesis, thereby enabling any image relighting technique to be seamlessly applied to video. Our approach reformulates video relighting into a simple yet effective two-stage process: (1) an artist relights a single reference frame using any preferred image-based technique (e.g., Diffusion Models, physics-based renderers); and (2) a fine-tuned stable video diffusion (SVD) model seamlessly propagates this target illumination throughout the sequence. To ensure temporal coherence and prevent artifacts, we introduce a gated cross-attention mechanism for smooth feature blending and a temporal bootstrapping strategy that harnesses SVD's powerful motion priors. Although trained on synthetic data, ReLumix shows competitive generalization to real-world videos. The method demonstrates significant improvements in visual fidelity, offering a scalable and versatile solution for dynamic lighting control.
Super-resolution ultrasound (SRUS) techniques, ultrasound localization microscopy (ULM) and super-resolution using erythrocytes (SURE), resolve vasculature beyond the diffraction limit, yet detection thresholds and filtering can bias width estimates. This study tests whether modest spatial smoothing, matched to the system's effective resolution (Gaussian sigma = 15 mu m), reduces bias relative to micro-CT. A Sprague-Dawley rat kidney was imaged with SURE and ULM and co-registered to ex vivo micro-CT (5.45 mu m voxels). Vessel widths were measured as the full width half maximum of SRUS line profiles and local thickness on micro-CT along identical lines, stratified by vessel type (interlobar, arcuate, and cortial radial vessels). Without smoothing, SRUS underestimated micro-CT across classes: interlobar mean normalized error was -89.9 +/- 8.9 % (SURE) and -80.6 +/- 12.4 % (ULM); arcuate -59.9 +/- 13.1 % (SURE) and -87.2 +/- 10.9 % (ULM); radial -39.3 +/- 34.3 % (SURE) and -91.1 +/- 5.3 % (ULM). With smoothing, errors moved toward zero: interlobar 9.7 +/- 35.7 % (SURE) and -1.0 +/- 25.7 % (ULM); arcuate -4.3 +/- 40.5 % (SURE) and 28.8 +/- 41.3 % (ULM); radial 23.4 +/- 31.7 %(SURE) and 14.6 +/- 23.8 % (ULM). Paired, line-wise normalized improvements were predominantly positive, e.g., interlobar gains of 60 +/- 22 % (SURE) and 59 +/- 22 %, with the largest class-wise gain up to 73 +/- 21 % for radial vessels in ULM. These results show that resolution-matched smoothing mitigates systematic underestimation in SRUS width estimates, yielding near-unbiased ULM interlobar measurements and attenuated bias across vessel classes, and support modality-aware preprocessing for structural quantification in SRUS.
While neural networks achieve strong performance in medical image analysis, effectively combining their predictions with human expertise remains a critical challenge for clinical deployment. We examine how different choices of stochastic parameter subsets used in approximate Bayesian inference impact the posterior predictive distributions and, consequently, the performance of a combined human-AI decision model. Using two medical classification tasks, we analyze the relationship between the resulting model and human uncertainty. We demonstrate that uncertainty estimates correlate differently with human uncertainty depending on the stochastic subsets. Building on these findings, we propose a framework that optimizes the choice of stochastic subsets to improve a final decision model that considers human uncertainty, enabling more reliable and interpretable integration of human and AI predictions in clinical settings. Our implementation is publicly available at https://github.com/mkreimann/uncertainty-guided-classification.
A rapid, contrast-free method for ultrasound superresolution vector flow imaging is introduced. The approach enhances the SURE-Hankel pipeline by integrating frequency-based directional separation, Hankel-matrix Singular Value Decomposition (HSVD), and localized optical flow estimation. A multi-stage process isolates overlapping erythrocyte signals based on their directional and spatiotemporal characteristics. Subsequently, optical flow is applied to the envelope of these separated modes for 2D flow vector estimation. The method's validation was performed on an in vivo rat kidney (8.7 MHz center frequency, wavelength lambda approximate to 178 mu m). For 1-second acquisitions, spatial resolution was improved to 40.8 mu m from 57.5 mu m using standard SURE. Flow direction accuracy, quantified against ULM velocimetry, resulted in a mean absolute Flow Angle Difference of 28.7 degrees and a mean bias of 1.3 degrees, successfully determining flow direction both axially and laterally. This technique enables accurate, non-invasive vector flow mapping of the microvasculature within acquisitions compatible with a single breath-hold.
Medical image segmentation often involves inherent uncertainty due to variations in expert annotations. Capturing this uncertainty is an important goal and previous works have used various generative image models for the purpose of representing the full distribution of plausible expert ground truths. In this work, we explore the design space of diffusion models for generative segmentation, investigating the impact of noise schedules, prediction types, and loss weightings. Notably, we find that making the noise schedule harder with input scaling significantly improves performance. We conclude that x- and v-prediction outperform epsilon-prediction, likely because the diffusion process is in the discrete segmentation domain. Many loss weightings achieve similar performance as long as they give enough weight to the end of the diffusion process. We base our experiments on the LIDC-IDRI lung lesion dataset and obtain state-of-the-art (SOTA) performance. Additionally, we introduce a randomly cropped variant of the LIDC-IDRI dataset that is better suited for uncertainty in image segmentation. Our model also achieves SOTA in this harder setting.
The human placenta exhibits a complex three-dimensional (3D) structure with a interpenetrating vascular tree and large internal interfacial area. In a unique and yet insufficiently explored way, this parenchymal structure enables its multiple functions as a respiratory, renal, and gastrointestinal multiorgan. The histopathological states are highly correlated with complications and health issues of mother, and fetus or newborn. Macroscopic and microscopic examination has so far been challenging to reconcile on the entire organ. Here we show that anatomical and histological scales can be bridged with the advent of hierarchical phase-contrast tomography and highly brilliant synchrotron radiation. To this end, we are exploiting the new capabilities offered by the BM18 beamline at ESRF, Grenoble for whole organ as well as the coherence beamline P10 at DESY, Hamburg for high-resolution, creating unique multiscale datasets. We also show that within certain limits, translation to μCT instrumentation for 3D placenta examination becomes possible based on advanced preparation and CT protocols, while segmentation of the datasets by machine learning now remains the biggest challenge.
We used diffusion MRI and x-ray synchrotron imaging on monkey and mice brains to examine the organisation of fibre pathways in white matter across anatomical scales. We compared the structure in the corpus callosum and crossing fibre regions and investigated the differences in cuprizone-induced demyelination in mouse brains versus healthy controls. Our findings revealed common principles of fibre organisation that apply despite the varying patterns observed across species; small axonal fasciculi and major bundles formed laminar structures with varying angles, according to the characteristics of major pathways. Fasciculi exhibited non-straight paths around obstacles like blood vessels, comparable across the samples of varying fibre complexity and demyelination. Quantifications of fibre orientation distributions were consistent across anatomical length scales and modalities, whereas tissue anisotropy had a more complex relationship, both dependent on the field-of-view. Our study emphasises the need to balance field-of-view and voxel size when characterising white matter features across length scales.