
Most of color image restoration models overlook the valid information across color spaces and the coupling between color channels, which are essential properties. To address such limitations, we present a new two-stage image restoration model: neural cross-space total variation regularization model with quaternion-based group sparse representation (QGS-NCSTV). In the first stage, a novel neural cross-space total variation regularization functional effectively captures the cross-space information to restore color images by incorporating the perceptual characteristics of the HSV space and the structural correlations of the RGB space. In the second stage, the quaternion-based group sparse representation preserves the coupling between color channels with leveraging nonlocal self-similarity. Experimental results show QGS-NCSTV outperforms the state-of-the-art methods in terms of PSNR, SSIM and UQI, where the reflection in colonoscopy images is successfully solved.
In the framework of edge-weighted graphs, watersheds have proven to be linked to well-known optimization problems, as Minimum Spanning Tree, which allowed the design of efficient algorithms for computing (hierarchical) watershed segmentations. In the present article, after reviewing the literature related to watershed segmentation, we present a detailed end-to-end pipeline of algorithms to compute (hierarchical) watershed segmentations, starting from the computation of graph-based image representations, up to the computation of connected components of the final (hierarchical) segmentation. We consider the several variations of watersheds, including their supervised and unsupervised versions, and the various ways of computing seeds, to name a few. For the first time, we bring together all these watershed notions and algorithms in a compact and understandable way. We aim at providing a reference for those interested in employing and reimplementing the watershed segmentation framework for their task at hand.
Initially explored by the mathematical morphology community in the early 1990 s, the integration of morphological operations into neural network models has been a fruitful line of research, which has also regained momentum over the past decade with the rise of deep learning. Two broad families of morphological neural networks have emerged in this context: morphological perceptrons and morphological convolutional neural networks, reflecting the evolution of classical neural architectures. This state-of-the-art report provides a comprehensive review of these developments. For each family, we present the mathematical formulation underlying the corresponding morphological network architectures, review the various architectures and training strategies introduced in the literature, and highlight practical applications. We conclude by discussing open challenges and promising directions for future research in this field.
In a recent article, we introduced the new notion of a complete tree of shapes as the keystone of a unifying framework for the partial partition trees. Beyond its ability to provide a continuum between usual morphological trees, namely the component trees and the tree of shapes, we also established that the complete tree of shapes could be contracted to build a topological tree of shapes, which is both an image model and a topological invariant. In this article, we extend this study by providing new theoretical, algorithmic and applicative results. We present a general scheme for building the so-called compact trees of shapes, which are trees obtained by applying decreasing homeomorphisms on the complete tree of shapes, leading to a wide family of known and new morphological trees, including many topological trees of shapes, built from topological invariants natively designed for binary images. We also propose the first optimal algorithm for building the complete tree of shapes, which allows to effectively use this tree, but also the derived compact and topological trees. Finally, we provide an example of the relevance of the notion of topological tree of shapes, by showing how we can design connected operators that preserve the topological tree of shapes, seen as a topological invariant, opening the way to the construction of morpho-topological operators.
Algorithm unfolding networks provide a bridge between classical variational models and deep learning for interpretable image segmentation. However, existing methods based on the Potts model typically implement region force and regularization terms using UNet-style modules and fixed regularizers, respectively, which limits their ability to capture global contextual information and complex boundaries. To overcome these limitations, we propose PottsSAGNet, which integrates a lightweight convolution-based self-attention module to approximate the region force term and trainable Gaussian convolutions within a fields of experts (FoE) framework for adaptive regularization. Additionally, a composite loss function combining Chan-Vese fidelity, mean squared error, and dice score stabilizes training and enhances segmentation accuracy. Experiments on multiple benchmarks demonstrate that PottsSAGNet achieves superior performance with significantly fewer parameters–only 3.52M, substantially lower than competing models while maintaining higher accuracy and dice scores. This work validates the effectiveness of attention-driven and learnable regularization designs in interpretable deep segmentation networks.
Video data from coherent systems are often corrupted by multiplicative noise, which poses a significant restoration challenge due to its inherent nonlinearity. To address this problem, we propose NL-LMURE, an unbiased and nonlocal linear regression framework for denoising video data. We derive an unbiased risk estimator for arbitrary linear denoising functions that accurately estimates the true squared Euclidean norm under multiplicative noise. By integrating this estimator with a nonlocal linear functional form, we obtain a closed-form solution that frames the denoising process as a principled ridge-regression problem. Furthermore, we introduce a robust matching criterion and a variance estimation method tailored to multiplicative noise. Theoretically, we prove that our framework is applicable to a broad class of noise distributions. Experimentally, we demonstrate the effectiveness and efficiency of our approach across diverse multiplicative noise distributions. Code is available at https://github.com/TISGroup/NL-LMURE .
A new mathematical analysis is given for the problem of recovering the position and orientation of a camera equipped with an accelerometer from sensor images of two labeled landmarks whose positions in a coordinate system aligned in a known way with gravity are known (P2PA). This is a variant on the much studied PnP problem of recovering camera position and orientation from n points without any gravitational data. It is proved that in three types of degenerate cases there are infinitely many solutions, in another type of case there is one, and in a final type of case there are two. A precise geometric characterization of each type of case is given. This completes the incomplete and at times incorrect characterizations in previous research. In particular, there is always a unique solution in the practically interesting case where the two landmarks are at the same altitude and the camera is at a different altitude. It is also proved that if the two landmarks are unlabeled, then apart from the same degenerate cases, there are still always one or two solutions.
We propose an image registration model for computing a dense, piecewise diffeomorphic deformation map between two images that facilitates sliding motion. Our approach is based on stationary velocity fields and implicit neural representations, with a particular emphasis on accurate motion interpolation. We investigate the effect of composing deformations in motion interpolation. To enforce domain alignment, we employ a hybrid technique that combines velocity field restriction with a soft penalty. We use 3D MR images of the thorax, acquired at end-inspiration and end-expiration, to create a time-continuous model of the respiratory cycle that captures pleural sliding motion. We experimentally validate the registration accuracy and the anatomical plausibility of the motion interpolation results, and discuss the inclusion of 2D intermediate data to improve interpolation accuracy.
Networks containing learnable morphological operators are challenging to train. In order to investigate their optimization difficulties, we focus on neural architectures inspired by the morphological representation theory, that guarantees expressivity. The sparsity of the gradient (or subgradient) of morphological layers is identified as the main optimization limitation, rendering the training sensitive to initialization. In order to make the training robust, we propose tailored to the data initialization and optimization methods that update multiple parameters simultaneously. Along the way we introduce two new theoretical representation results, one inspiring our adaptive initialization strategy, and the other one helping interpret results which are unexpected in the existing theory.
A key challenge in image restoration is to define a realistic prior on clean images to complete the missing information in the observation. Using a pretrained neural network to encode this prior has shown to be state of the art for many problems. Typical image distributions are invariant to some set of transformations, such as rotations or flips. However, most deep architectures are not designed to represent an image distribution with some invariances. Recent works have proposed to overcome this difficulty by including equivariance properties within a Plug-and-Play paradigm. In this work, we propose two frameworks named Equivariant Regularization by Denoising (ERED) and Equivariant Plug-and-Play (EPnP) based on equivariant denoisers and stochastic optimization. ERED unifies in a single formulation several existing algorithms, and both ERED and EPnP generalize the well-known RED and PnP algorithms. We analyse the convergence of the proposed algorithms and discuss their practical benefit.
The full-rank decomposition of reduced biquaternion matrices is fundamental in quaternion algebra and critical for its engineering applications. However, the high computational complexity of biquaternion operations and the lack of unified explicit expressions remain key challenges. We investigate the full-rank decomposition of row or column full-rank reduced biquaternion matrices, derive its explicit expression via their complex representation, and propose a novel numerical algorithm that only involves complex field operations, with theoretical analysis verifying its correctness. We also theoretically demonstrate the decomposition algorithm’s applicability to idempotent matrix factorization and least squares matrix equation solving. Numerical experiments confirm the decomposition algorithm’s efficiency and simplicity, especially in handling the aforementioned matrix tasks. We further apply this decomposition to color image encryption and decryption in image processing and experimental results prove the corresponding encryption–decryption algorithm’s effectiveness and robustness. This research enriches the theoretical system of reduced biquaternion matrix decomposition and provides a practical solution for related engineering scenarios.
We contribute to an uncertainty quantification problem in imaging that evaluates a hypothesis test questioning the existence of local “artifacts” appearing in the maximum a posteriori (MAP) estimate (obtained from standard numerical tools). Such a method, called Bayesian uncertainty quantification by optimization (BUQO), was introduced a few years ago as an efficient and scalable alternative to sampling methods when per-pixel error bars are not needed. BUQO formulates a hypothesis test for probing the existence of local structures in the MAP estimate as a minimization problem that can be solved efficiently with standard optimization algorithms. In this context, BUQO requires a “mathematical” definition of the “local artifact.” This definition can be interpreted as an inpainting of the structure. However, only simple handcrafted techniques have been proposed so far due to the complexity of the problem. In this work, we propose a data-driven alternative to BUQO where the inpainting procedure in the algorithm is performed using a convolutional inpainting neural network (NN). This results in a plug-and-play algorithm, based on the primal-dual Condat–Vũ iterations, where the inpainting procedure is performed with a NN. The proposed approach is assessed on two image reconstruction problems inspired by medicine. We specifically perform simulations on two Fourier undersampling problems (discrete and non-uniform) encountered in magnetic resonance imaging, as well as a computed tomography problem using the Radon measurement operator.
We study a particular class of greedy algorithms for combinatorial optimization problems and present a generalized version of this algorithmic pattern encompassing several previously published algorithms. We analyze the properties of the solutions produced by such algorithms and provide proofs of their optimality through the concept of lexicographic max-ordering. By presenting a unified formulation of this class of greedy algorithms, we hope to facilitate the development of new such algorithms and to provide a deeper understanding on the properties of existing ones. To illustrate the utility of our results, we present two case studies, where we study two previously published optimization algorithms and show that they can be seen as instances of the proposed general algorithmic pattern. In doing so, we provide alternative proofs of correctness for these algorithms and draw new conclusions about their properties.
Parallel magnetic resonance imaging (pMRI) accelerates data acquisition by undersampling multi-coil k-space. Its reconstruction quality, however, deteriorates when coil sensitivity maps (CSMs) or calibration kernels are inaccurate, especially when only limited auto-calibration signal (ACS) data are available. We propose SRSC+, a model-driven bilevel optimization framework that couples SENSE-based image reconstruction with SPIRiT-based k-space calibration through shared CSMs. The bilevel formulation explicitly decouples sensitivity estimation from kernel calibration, thereby enabling iterative correction of both components and reducing error accumulation that often arises in dual-domain methods. In addition, SRSC+ introduces a deep-prior-guided regularization strategy that preserves the structure of classical linear regularizers while adaptively learning spatially varying regularization weights from denoised intermediate reconstructions. Experiments on out-of-distribution datasets under diverse sampling patterns show that SRSC+ achieves state-of-the-art performance across multiple fidelity and perceptual metrics, while remaining robust to scarce ACS data and imperfect CSM initialization. Visual comparisons further demonstrate effective artifact suppression without pseudo-structural distortions, together with strong generalization across scanners and acquisition protocols. The implementation code is available at https://github.com/Chenvp/SRSC .
We propose a unified variational framework for image segmentation under sparse pixel-level supervision. Our method is based on a simplex-constrained Potts model with a smooth perimeter regularizer, yielding a convex, smooth energy functional that can be used as a training loss in weakly supervised deep learning paradigms or optimized efficiently using iterative methods. Sparse labels are incorporated by constructing a fuzzy membership function via a function extension problem in a Reproducing Kernel Hilbert Space (RKHS), which effectively captures inhomogeneous intensity statistics. The derived discrete loss for training standard networks demonstrates robustness and consistent improvements over non-training and partial cross-entropy (PCE) baselines in experiments, achieving comparable performance without requiring ground-truth segmentation images.
We present a continuous-domain variational model for the morphological decomposition of 1D noisy signals into a sum of meaningful constituent components. The model relies on sparsifying fractional-order derivatives of the sought components to capture intricate signal structures. An in-depth analysis of the model leads to a representer theorem, establishing the equivalence between the infinite-dimensional problem and a finite-dimensional counterpart, which serves as its exact discretization. To efficiently solve the resulting discrete and convex optimization problem, an alternating direction method of multipliers-based algorithm is presented. Furthermore, we introduce a bilevel optimization framework for the automatic selection of all free model parameters, including the fractional derivative orders, based on the generalized whiteness principle. Numerical results validate the effectiveness of our approach, which can provide accurate decompositions even in demanding scenarios characterized by high noise levels and abrupt signal discontinuities.
Plug-and-Play (PnP) image restoration provides a flexible framework that integrates model-based optimization with powerful deep denoising priors. However, PnP methods based on deep neural networks are highly sensitive to perturbations. In this work, we propose a robust PnP framework embedding a frequency-domain robust correction block and an adaptive noise-level scheduling block for adaptive image restoration. The frequency-domain block selectively suppresses corrupted high-frequency components to align adversarial noise with AWGN assumptions, serving as a general-purpose preprocessing strategy that can be applied beyond PnP. The adaptive noise-level scheduling block based on patch-wise PCA estimation dynamically adjusts the denoiser noise level throughout the iterative process. Experiments on deblurring and superresolution demonstrate that our method achieves stronger robustness and restoration performance than existing PnP methods under adversarial attacks. Ablation studies validate the complementary contributions of the two proposed blocks.
We propose a novel extension of the empirical Bayes framework (EBF) for inverse problems by introducing a class of nonconvex and nonconcave L -smooth hyperpriors. These hyperpriors relax traditional convexity or concavity assumptions while preserving sufficient regularity for theoretical and algorithmic analysis, thereby allowing heavy-tailed hyperpriors with changing curvature to be incorporated into EBF. To solve the resulting optimization problem, we develop an alternating majorization–minimization (AMM) algorithm and establish its convergence to a stationary point. Through extensive numerical experiments on one-dimensional signal deconvolution and image deblurring, we investigate how the heavy-tailed behavior in the hyperprior affects sparsity, reconstruction accuracy, and structural preservation. In addition, the proposed method naturally enables uncertainty quantification via marginal posterior variance estimation. This work broadens the scope of EBF to a wider class of hyperpriors and provides practical guidance for hyperprior design and selection in ill-posed inverse problems.
Data-driven image segmentation methods frequently encounter challenges related to object size imbalance, particularly in cases where small targets occupy only a tiny proportion of the image. Furthermore, the theoretical foundation underlying how these methods learn object size remains poorly understood. To address this problem, we study data-driven image segmentation from the perspective of semi-dual optimal transport and introduce image-specific volume prior into the activation layer of the segmentation neural network. In this framework, the neural network logits are interpreted to the transport cost, and the bias term in the activation operator is associated with the dual variable that controls the distribution of target size. Based on this theoretical discovery, we develop a nonlinear activation mechanism, termed VP-Sparsemax, in image segmentation neural networks with volume priors from smooth approximations of the max operator in the semi-dual optimal transport problem. Different from existing loss-based modifications, the proposed nonlinear activation module incorporates both volume prior and spatial information, and can be embedded into the segmentation networks for both training and inference. Theoretical analysis and numerical results show that the proposed method can improve the ability of the networks to learn object size. Experiments on several real-world datasets and some representative segmentation network backbones such as U-Net, DeepLabV3+, and SAM, show that the proposed method consistently outperforms the existing related methods.