
Texture analysis consists of a classical and everlasting task in image processing, often involved in a broad range of applications, possibly very different in nature. For large classes of textures, scale-free (or fractal) spatial dynamics as well as anisotropy constitute key properties. However, the local (pixelwise) joint estimation of anisotropy and scale-free attributes constitutes a difficult challenge, as both consist of nonlocal properties. Yet accurately detecting variations of these attributes across the image is often crucial. The overarching goal of the present work is thus to propose an inverse problem formulation for the analysis of piecewise homogeneous textures, grounded jointly on scale-free dynamics and anisotropy, and to study its performance for local assessment of textures. The formulation combines several major contributions. First, piecewise homogeneous Gaussian fields are defined as texture mixture models, with prescribed scale-free and anisotropy properties. Second, multiband complex wavelet coefficients, implemented via (nondecimated) dual-tree fast algorithms, are theoretically shown to be sensitive to both scale-free dynamics and anisotropy. Notably, it is shown that the squared-modulus of the wavelet coefficients behaves locally (or pixelwise) approximately as power-laws, with both scaling exponent and intercept jointly sensitive to anisotropy and scale-free dynamics. Thus, a key originality of the present work is to propose to perform local analysis of intrinsically nonlocal properties without having recourse to local averages, which would significantly impair accurate segmentation or accurate local characterization. Third, these local power-law-like behaviors, combined with regularization terms enforcing piecewise homogeneity, are embedded into a minimization procedure, solved by an optimization algorithm, whose convergence conditions are theoretically well-studied, and practically implemented through an efficient proximal algorithm due to strong convexity of the resulting minimization problem. Segmentation performance achieved by the proposed procedure is quantified and compared with respect to the difficulty of the segmentation task, for synthetic piecewise homogeneous Gaussian fields. The potential of the proposed texture analysis tool is also illustrated at work on real-world textures.
Abstract. Minkowski tensors, also known as tensor valuations, provide robust [Formula: see text]-point information for a wide range of random spatial structures. Local estimators for point clouds, e.g., representing voxelized data, however, are unavoidably biased even in the limit of infinitely high resolution. Here, we substantially improve a recently proposed, asymptotically unbiased algorithm to estimate Minkowski tensors from point clouds. Our improved algorithm is more robust and efficient. Moreover we generalize the theoretical foundations for an asymptotically bias-free estimation of the interfacial tensors, among others, to the case of finite unions of compact sets with positive reach, which is relevant for many applications like rough surfaces or composite materials. As a realistic test case of random spatial structures, we consider random (beta) polytopes. We first derive explicit expressions of the expected Minkowski tensors, which we then compare to our simulation results. We obtain precise estimates with relative errors of a few percent for practically relevant resolutions. Finally, we apply our methods to real data of metallic grains and nanorough surfaces, and we provide an open-source Python package, which works in any dimension.
Abstract. One-bit quantized measurements are commonly encountered in resource-constrained imaging applications such as medical imaging and satellite remote sensing. Although stable signal reconstruction from noise-free one-bit measurements has been widely demonstrated, noise in practical data acquisition systems often leads to unknown sign flips, significantly increasing the difficulty of accurate recovery. While numerous robust methods have been developed for signal reconstruction from noisy one-bit measurements, most of them necessitate prior knowledge of the number of sign flips, which is typically unavailable in practice. In this paper, we propose two novel optimization models based on continuous approximation functions (CAFs) of the sign function. The first model employs the quadratic loss function, specifically designed for scenarios with low sign flipping ratios, while the second incorporates a Hybrid Ordinary-Welsch (HOW) loss function, yielding enhanced robustness under higher flipping ratios without requiring prior knowledge. Then, we propose two algorithmic frameworks, OBCSA_CAF and OBCSA_CAF_HOW, to solve the respective models, and prove that both algorithms generate convergent sequences of objective function values and subsequences that, with high probability, converge to specific accumulation points. Extensive simulations demonstrate the superior performance of our methods over the state-of-the-art algorithms in terms of signal-to-noise ratio (SNR), rate of support recovery, Hamming error, and Hamming distance. In particular, OBCSA_CAF achieves 0.5–4 dB gains in SNR and [Formula: see text]–[Formula: see text] improvements in support recovery rate at low flipping ratios, whereas OBCSA_CAF_HOW yields 2–6.7 dB gains and [Formula: see text]–[Formula: see text] improvements, respectively, under high flipping ratios. Experiments on the MNIST and CIFAR-10 datasets further illustrate the superior performance of our methods in image reconstruction.
Abstract. Multislice ptychography (msPtycho) extends conventional two-dimensional ptychographic phase retrieval to thick or strongly scattering specimens by modeling the object as a stack of slices and accounting for multiple scattering. While this improves physical fidelity, existing reconstruction algorithms lack rigorous theoretical guarantees and often suffer from either slow convergence or high computational complexity. In this work, we propose a fast and theoretically grounded framework for msPtycho. We formulate a constrained optimization model that exploits the successive dependence of adjacent exit waves, reducing the multilinear coupling into a chain of bilinear relations. Building on this structure, we develop an alternating minimization of msPtycho (AM[Formula: see text]SP), which admits closed-form updates for all subproblems and is proven to converge globally to stationary points. To further enhance the performance for reconstructing deeply layered specimens, we introduce a forward-propagating correction to suppress error accumulation across slices and an extrapolation strategy to accelerate convergence. Together, these yield the accelerated AM[Formula: see text]SP-FX algorithm. Extensive numerical experiments demonstrate that the proposed methods achieve faster convergence and better reconstructions. Notably, AM[Formula: see text]SP-FX achieves superior reconstruction quality and consistently faster per-iteration runtimes than both layer-wise optimization and 3D Ptychographical Iterative Engine (3PIE); in particular, it requires an order of magnitude less time per iteration than 3PIE. Overall, the proposed framework offers a practical and theoretically sound solution for high-fidelity multislice ptychographic imaging.
Abstract. Diffuse optical tomography (DOT) is an imaging modality in which images of the optical properties of a biological tissue, specifically diffusion [Formula: see text] and absorption [Formula: see text], are estimated based on measurements of near-infrared light on the surface of the body. In practical applications, one often lacks exact knowledge of the measurement domain boundary. This poses a significant challenge, as inaccuracies in the boundary shape of the computational domain may result in substantial artifacts in the reconstructed images. In this study, the following two results are achieved in a two-dimensional setting: (i) when the measurement domain [Formula: see text] is known, it is shown that knowledge of the Robin-to-Neumann map for two modulation frequencies uniquely determines [Formula: see text] and [Formula: see text], both assumed to be isotropic. (ii) When the measurement domain [Formula: see text] is not exactly known, a method is proposed for simultaneous reconstruction of [Formula: see text] and [Formula: see text] as well as the boundary [Formula: see text]. For this, DOT measurements are needed for two modulation frequencies. This approach yields reconstructed coefficients that are a conformal deformation of the true coefficients in the exact domain. The new method is demonstrated using simulated noisy data.
This paper proposes a two-stage framework to address the inverse diffraction grating problem with limited-aperture data. The first stage introduces a deep learning network for data retrieval, featuring a dual-branch, cross-attention architecture. Motivated by an information-theoretic analysis, this design is tailored to separate and adaptively fuse the diffracted field's low-and high-frequency components, effectively handling their distinct noise sensitivities. The second stage employs a computationally efficient Newton-type algorithm for shape reconstruction, which avoids the need for a forward solver at each iteration. Numerical experiments show that our framework provides accurate and robust reconstructions.
Computational hemo dynamics can enhance image-based diagnosis and provide complementary insights to predict, understand, and monitor treatments. The high computational costs and the complexity associated with handling patient-specific settings remain a major challenge toward clinical applications. In this work, we propose a novel robust shape registration method for nonparametric aortic geometries, describing different applications for projection-based reduced-order modeling for the training of graph neural networks and for data assimilation. The registration approach is based on ResNet-LDDMM, trained with a dataset of synthetic shapes, generated from real ones with statistical shape modeling. The optimization is tailored to surface meshes and does not rely on a priori assumptions on domain parameterization. We employ a multigrid strategy during the training phase that allows handling realistic mesh sizes. The registration enables the definition of geometric encoding of different blood flow solutions on a single reference shape, as well as the design of projection-based reduced-order models. We use this geometrical encoding to improve the training of graph neural networks and to present potential applications in data assimilation problems, combined with a generalized parameterized-Background Data-Weak formulation. As a particular example of data assimilation problem, we address the reconstruction of velocity fields and wall shear stresses, as well as the estimation of pressure fields and pressure-related biomarkers, such as the pressure drop, from low-resolution velocity observations. We show various numerical tests based on synthetic data, comparing the proposed strategies with state-of-the-art estimators.
Motivated by the use of unmanned aerial vehicles (UAVs) equipped with ground penetrating synthetic aperture radar (GP-SAR) for imaging buried landmines, we study imaging of subsurface targets. A key challenge in this problem is the removal of so-called ground bounce signals which are reflections from the air-soil interface. When ground bounce signals can be removed exactly from GP-SAR measurements, one can readily image subsurface targets using traditional synthetic aperture imaging methods such as Kirchhoff migration (KM). However, ground bounce signals dominate measurements, carry no useful information about subsurface targets, and disrupt their imaging. In this paper we present an asymptotic theory for this problem assuming that the elevation of the UAV flight path is the largest length scale in the problem. Using this asymptotic theory, we introduce data-driven methods for (i) estimating the ground permittivity, and (ii) approximately removing ground bounce signals from GP-SAR. We first establish these methods for the simple case of a flat air-soil interface and then apply them to an unknown rough air-soil interface. We validate this asymptotic theory using numerical simulations.
Fusing hyperspectral images (HSIs) with multispectral images (MSIs) has become a mainstream approach to enhance the spatial resolution of HSIs. Recently, many HSI-MSI fusion techniques have been developed and have achieved notable performance. Nevertheless, certain challenges still persist, including (a) Most existing fusion methods assume precise registration between HSIs and MSIs, which can be challenging to achieve in practice. (b) The obtained HSI-MSI pairs may not be fully utilized. To address these issues, we propose a coarse-to-fine hybrid registration and fusion approach. In the coarse stage, a joint registration and fusion (JRF) model is developed by integrating batch image alignment into the fusion process, which can achieve HSI-MSI registration and fusion simultaneously. To further enhance fusion performance, we propose a nonconvex low-rank and group-sparse (NLG) model in the fine stage. This model exploits the low-rank property to reconstruct the global image structure (i.e., the clean image), while simultaneously leveraging group sparsity to retrieve the textural details. This combination yields the proposed JRF-NLG method. Then, the JRF-NLG model is solved by the generalized Gauss-Newton (GGN) algorithm and the proximal alternating optimization (PAO) algorithm, respectively. Theoretically, we establish an error bound for the NLG model and prove the quadratic convergence rate of the GGN algorithm. Finally, extensive numerical experiments on simulated and real datasets are performed to verify the effectiveness of our method in registration and fusion. We also validate its effectiveness in enhancing classification performance.
We address the optimization problem in a data-driven variational reconstruction framework, where the regularizer is parametrized by an input-convex neural network. While gradient-based methods are commonly used to solve such problems, they struggle to effectively handle nonsmooth problems, which often leads to slow convergence. Moreover, the nested structure of the neural network complicates the application of standard nonsmooth optimization techniques, such as proximal algorithms. To overcome these challenges, we reformulate the problem and eliminate the network's nested structure. By relating this reformulation to epigraphical projections of the activation functions, we transform the problem into a convex optimization problem that can be efficiently solved using a primal-dual algorithm. We also prove that this reformulation is equivalent to the original variational problem. Through experiments on several imaging tasks, we show that the proposed approach not only outperforms subgradient methods and even accelerated methods in the smooth setting but also facilitates the training of the regularizer itself.
This work considers a time domain inverse acoustic obstacle scattering problem due to passive data. Motivated by the Helmholtz-Kirchhoff identity in the frequency domain, we propose to relate the time domain measurement data in passive imaging to an approximate data set given by the subtraction of two scattered wave fields. We propose a time domain linear sampling method for the approximate data set and show how to tackle the measurement data in passive imaging. An imaging functional is built based on the linear sampling method, which reconstructs the support of the unknown scattering object using directly the time domain measurements. The functional framework is based on the Laplace transform, which relates the mapping properties of Laplace domain factorized operators to their counterparts in the time domain. Numerical examples are provided to illustrate the capability of the proposed method.
Stochastic gradient descent (SGD) and its variants are widely used and highly effective optimization methods in machine learning, especially for neural network training. By using a single datum or a small subset of the data, selected randomly at each iteration, SGD scales well to problem size and has been shown to be effective for solving large-scale inverse problems. In this work, we investigate SGD for solving nonlinear inverse problems in Banach spaces through the lens of iterative regularization. Under general assumptions, we prove almost sure convergence of the iterates to the minimumdistance solution and show the regularizing property in expectation under an a priori stopping rule. Further, we establish convergence rates under the conditional stability assumptions for both exact and noisy data. Numerical experiments on Schlieren tomography and electrical impedance tomography are presented to show distinct features of the method.
One-bit quantized measurements are commonly encountered in resource-constrained imaging applications such as medical imaging and satellite remote sensing. Although stable signal reconstruction from noise-free one-bit measurements has been widely demonstrated, noise in practical data acquisition systems often leads to unknown sign flips, significantly increasing the difficulty of accurate recovery. While numerous robust methods have been developed for signal reconstruction from noisy one-bit measurements, most of them necessitate prior knowledge of the number of sign flips, which is typically unavailable in practice. In this paper, we propose two novel optimization models based on continuous approximation functions (CAFs) of the sign function. The first model employs the quadratic loss function, specifically designed for scenarios with low sign flipping ratios, while the second incorporates a Hybrid Ordinary-Welsch (HOW) loss function, yielding enhanced robustness under higher flipping ratios without requiring prior knowledge. Then, we propose two algorithmic frameworks, OBCSA CAF and OBCSA CAF HOW, to solve the respective models, and prove that both algorithms generate convergent sequences of objective function values and subsequences that, with high probability, converge to specific accumulation points. Extensive simulations demonstrate the superior performance of our methods over the state-of-the-art algorithms in terms of signal-tonoise ratio (SNR), rate of support recovery, Hamming error, and Hamming distance. In particular, OBCSA CAF achieves 0.5--4 dB gains in SNR and 1\%--3.6\% improvements in support recovery rate at low flipping ratios, whereas OBCSA CAF HOW yields 2--6.7 dB gains and 2\%--8\% improvements, respectively, under high flipping ratios. Experiments on the MNIST and CIFAR-10 datasets further illustrate the superior performance of our methods in image reconstruction.
Deep learning on non-Euclidean domains is important for analyzing complex geometric data that lacks common coordinate systems and familiar Euclidean properties. A central challenge in this field is to define convolution on domains, which inherently possess irregular and non-Euclidean structures. In this work, we introduce quasi-conformal convolution (QCC), a novel framework for defining convolution on simply-connected open surfaces using quasi-conformal theories. Each QCC operator is linked to a specific quasi-conformal mapping, enabling the adjustment of the convolution operation through manipulation of this mapping. By utilizing trainable estimator modules that produce quasi-conformal mappings, QCC facilitates adaptive and learnable convolution operators that can be dynamically adjusted according to the underlying data structured on the surfaces. QCC unifies a broad range of spatially defined convolutions, facilitating the learning of tailored convolution operators on each underlying surface optimized for specific tasks. Building on this foundation, we develop the quasi-conformal convolutional neural network (QCCNN) to address a variety of tasks related to geometric data. We validate the efficacy of QCCNN through the classification of images defined on curvilinear simply-connected open Riemann surfaces, demonstrating superior performance in this context. Additionally, we explore its potential in medical applications, including craniofacial analysis using 3D facial data and lesion segmentation on 3D human faces, achieving enhanced accuracy and reliability.
We derive an algorithm for compression of the currents and varifolds representations of shapes using Ridge Leverage Score sampling and the theory of Nystrom approximation in Reproducing Kernel Hilbert Spaces. Our method is faster than existing compression techniques and comes with theoretical guarantees on the rate of convergence of the compressed approximation as a function of the smoothness of the associated shape representation. The obtained compressions are shown to be useful for down-line tasks such as nonlinear shape registration in the Large Deformation Diffeomorphic Metric Mapping (LDDMM) framework, even for very high compression ratios. The performance of our algorithm is demonstrated on large-scale shape data from modern geometry processing datasets and is shown to be fast and scalable with rapid error decay.
In photoacoustic tomography (PAT), a hybrid imaging modality that is based on the acoustic detection of optical absorption from biological tissue exposed to a pulsed laser, a short pulse laser generates an initial pressure proportional to the absorbed optical energy, which then propagates acoustically and is measured on the boundary. To account for the significant signal distortion caused by acoustic attenuation in biological tissue, we model PAT in heterogeneous media using a damped wave equation featuring spatially varying sound speed and a time-dependent damping term. Under natural assumptions, we show that the initial pressure is uniquely determined by the boundary measurements using a harmonic extension of the boundary data with energy decay. For constant damping, an expansion in Dirichlet eigenfunctions of -c^2()Δ leads to an explicit series reconstruction formula for the initial pressure. Finally, we develop a gradient free numerical method based on the Pontryagin's maximum principle to provide a robust and computationally viable approach to image reconstruction in attenuating PAT.
This paper introduces a novel intrinsic semiparametric mixed-effects model (ISMEM) to characterize the complex relationships between manifold-valued responses and Euclidean-valued covariates in the longitudinal setting. Manifold-valued data, characterized by their inherent nonlinearity, high dimensionality, and specific geometric structures, present significant challenges for longitudinal analysis. Most existing intrinsic models are developed for cross-sectional data and unsuited for longitudinal studies. To address this limitation, we propose ISMEM, an efficient framework designed for longitudinal manifold-valued datasets that capture such relationships at both group and individual levels. Key features of ISMEM include (i) the integration of fixed and random effects on Riemannian manifolds, enabling analysis at multiple levels; (ii) a semiparametric structure combining parametric and nonparametric components, enhancing both flexibility and interpretability; and (iii) the preservation of the geometric properties of manifold-valued observations, offering improved robustness and interpretability compared to Euclidean-based models. In addition, we develop a two-stage iterative estimation procedure and validate our approach through simulations on the symmetric positive definite (SPD) manifold. Finally, we apply ISMEM to longitudinal data from the Alzheimer's Disease Neuroimaging Initiative (ADNI), reconstructing and comparing continuous three-dimensional (3D) shape trajectories of the lateral ventricle in Kendall shape space across distinct groups.
This paper addresses the challenging problem of modeling and computing multivalued mappings with varying cardinality, where a single input can correspond to multiple valid outputs, and the number of possible outputs may vary for different inputs. Such scenarios are prevalent in various imaging applications where multiple plausible solutions exist for a given input. We introduce a deep neural network framework to model multivalued mappings with varying cardinality. The framework integrates a discrete codebook with a generative network to produce valid outputs for each input. The discrete codebook variables are combined with the input to guide the generator in producing different valid solutions. The discrete nature of the codebook enables the framework to efficiently estimate the conditional probability distribution of possible outputs through a fixed equiangular tight frame classifier. By jointly optimizing the discrete codebook and its uncertainty estimation during training using a specially designed covariance loss function, an accurate computation of multiple solution candidates with reliable confidence measures can be achieved. We demonstrate the effectiveness of the proposed framework on various imaging applications, using both synthetic and real datasets. Experimental results show the efficacy of our proposed model to generate multiple high-quality outputs while providing meaningful uncertainty estimates for each solution.
In this paper, we address the problem of robust orthogonal nonnegative matrix factorization (RONMF), a crucial challenge in data analysis. We first propose a RONMF model that explicitly handles both dense and sparse noise, making it suitable for a wide range of real-world applications. To circumvent the computational complexity associated with the Stiefel manifold and effectively solve the proposed model, we introduce an exact penalty method that transforms the optimization problem from the Stiefel manifold to the Oblique manifold. To achieve this, we develop the EP-RONMF algorithm, which seeks a point satisfying the weak second-order optimality conditions through an alternating proximal method and iterative updates of the penalty parameter. This approach successfully addresses the nonconvex nature of the ONMF problem and ensures convergence to a stable solution. To validate the efficacy of our method, we conducted extensive experiments on diverse datasets, including image, text, and hyperspectral data. The results clearly demonstrate the superiority of our approach compared to existing techniques.
Stochastic gradient descent (SGD) and its variants are widely used and highly effective optimization methods in machine learning, especially for neural network training. By using a single datum or a small subset of the data, selected randomly at each iteration, SGD scales well to problem size and has been shown to be effective for solving large-scale inverse problems. In this work, we investigate SGD for solving nonlinear inverse problems in Banach spaces through the lens of iterative regularization. Under general assumptions, we prove almost sure convergence of the iterates to the minimum distance solution and show the regularizing property in expectation under an a priori stopping rule. Further, we establish convergence rates under the conditional stability assumptions for both exact and noisy data. Numerical experiments on Schlieren tomography and electrical impedance tomography are presented to show distinct features of the method.