We study the recovery of an unknown three-dimensional band-limited signal from multiple noisy observations that are randomly rotated by latent elements of SO(3), where the rotations are drawn from an unknown, non-uniform distribution. Because the rotations are unobserved, only the signal orbit under the rotation group can be recovered. We show that the signal orbit and the rotation distribution are jointly identifiable from the first and second moments. This yields an improved high-noise sample complexity that scales quadratically with the noise variance, rather than cubically as in the uniform-rotation case. We further develop a provable, computationally efficient reconstruction algorithm that recovers the 3-D signal by successively solving a sequence of well-conditioned linear systems. The algorithm is validated through extensive numerical experiments. Our results provide a principled and tractable framework for high-noise 3-D orbit recovery, with potential relevance to cryo-electron microscopy and cryo-electron tomography modeling, where molecules are observed in unknown orientations.
In many applications of matching, the point clouds to be matched are not merely unstructured sets of points but rather samples from distributions with an intrinsic cluster structure. In such cases, as individual points are often interchangeable within a coherent region, finding a robust region-to-region alignment is more desirable than establishing a precise point-to-point correspondence. To this end, we propose a novel approach for cluster-aware matching based on Laplacian Optimal Transport (LapOT). The key idea is to regularize the optimal transport problem with quadratic Laplacian terms constructed from similarity graphs of the point clouds, which encourages the optimal coupling to respect the cluster structure of both point sets. We also introduce Refined Simultaneous Clustering (RSC), a method that leverages the cluster-aware coupling obtained from LapOT to produce consistent partitions across the point sets, which can overcome the limitations of independent clustering and yield more stable and interpretable results. We demonstrate the effectiveness of our approach through theoretical analysis and empirical experiments, showing that LapOT indeed produces cluster-aware matching that leads to more consistent and meaningful alignments between point clouds.
The approximation of unknown functions from scattered, possibly high-dimensional data is central to many scientific applications. Advances in data acquisition have driven the need for flexible nonlinear models, including manifold-valued functions. Approximating and learning such functions differs fundamentally from classical linear methods and requires tools from numerical analysis, linear algebra, and differential geometry. This interdisciplinary framework has applications ranging from data science and machine learning to numerical PDEs and quantum chemistry. This mini-workshop brings together researchers developing constructive approximation methods for manifold-valued functions, their theory, and applications.
Modeling data using manifold values is a powerful concept with numerous advantages, particularly in addressing nonlinear phenomena. This approach captures the intrinsic geometric structure of the data, leading to more accurate descriptors and more efficient computational processes. However, even fundamental tasks like compression and data enhancement present meaningful challenges in the manifold setting. This paper introduces a multiscale transform that aims to represent manifold-valued sequences at different scales, enabling novel data processing tools for various applications. Similar to traditional methods, our construction is based on a refinement operator that acts as an upsampling operator and a corresponding downsampling operator. Inspired by Wiener's lemma, we term the latter as the reverse of the former. It turns out that some upsampling operators, for example, least-squares-based refinement, do not have a practical reverse. Therefore, we introduce the notion of pseudo-reversing and explore its analytical properties and asymptotic behavior. We derive analytical properties of the induced multiscale transform and conclude the paper with numerical illustrations showcasing different aspects of the pseudo-reversing and two data processing applications involving manifolds.
Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations – common in structural biology – presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.
Capturing data from dynamic processes through cross-sectional measurements is seen in many fields, such as computational biology. Trajectory inference deals with the challenge of reconstructing continuous processes from such observations. In this work, we propose methods for B-spline approximation and interpolation of point clouds through consecutive averaging that is intrinsic to the Wasserstein space. Combining subdivision schemes with optimal transport-based geodesic, our methods carry out trajectory inference at a chosen level of precision and smoothness, and can automatically handle scenarios where particles undergo division over time. We prove linear convergence rates and rigorously evaluate our method on cell data characterized by bifurcations, merges, and trajectory splitting scenarios like supercells, comparing its performance against state-of-the-art trajectory inference and interpolation methods. The results not only underscore the effectiveness of our method in inferring trajectories but also highlight the benefit of performing interpolation and approximation that respect the inherent geometric properties of the data.
Pyramid transforms are constructive methods for analyzing sequences in a multiscale fashion. Traditionally, these transforms rely on stationary upsampling and downsampling operations. In this paper, we propose employing nonstationary subdivision schemes as upsampling operators that vary with the refinement level. These schemes offer greater flexibility, enabling the development of more expressive multiscale transforms, including geometric multiscale analysis. We establish the fundamental properties of the resulting nonstationary transforms, namely the decay of their detail coefficients and the stability of both the decomposition and the reconstruction, and we demonstrate their effectiveness in capturing and analyzing geometric features. In particular, we apply the framework to assess the consistency of planar samples with a circular shape and to detect localized geometric anomalies, including a study of circles generated by a neural network. The accompanying code is publicly available.
We present a novel multiscale framework for analyzing sequences of probability measures in Wasserstein spaces over Euclidean domains. Exploiting the intrinsic geometry of optimal transport, we construct a multiscale transform applicable to both absolutely continuous and discrete measures. Central to our approach is a refinement operator based on McCann's interpolants, which preserves the geodesic structure of measure flows and serves as an upsampling mechanism. Building on this, we introduce the optimality number, a scalar that quantifies deviations of a sequence from Wasserstein geodesicity across scales, enabling the detection of irregular dynamics and anomalies. We establish key theoretical guarantees, including stability of the transform and geometric decay of coefficients, ensuring robustness and interpretability of the multiscale representation. Finally, we demonstrate the versatility of our methodology through numerical experiments: denoising and anomaly detection in Gaussian flows, analysis of point cloud dynamics under vector fields, and the multiscale characterization of neural network learning trajectories.
We develop a unified framework for nonlinear subdivision schemes on complete metric spaces (CMS). We begin with CMS preliminaries and formalize refinement in CMS, retaining key structural properties, such as locality. We prove a convergence theorem under contractivity and demonstrate its applicability. To address schemes where contractivity is unknown, we introduce two notions of proximity. Our proximity methods relate a nonlinear scheme to another nonlinear scheme with known contractivity, rather than to a linear scheme, as in much of the literature. Specifically, the first type proximity compares the two schemes after a single refinement step and, as in the classical theory, yields convergence from sufficiently dense initial data. The proximity of the second type monitors alignment across all refinement levels and provides strong convergence without density assumptions. We formulate and prove the corresponding theorems, and illustrate them with various examples, such as schemes over metric spaces of compact sets in ^n and schemes over the Wasserstein space, as well as a geometric Hermite metric space. These results extend subdivision theory beyond Euclidean and manifold-valued data for data in metric spaces.
The process of particle picking, a crucial step in cryo-electron microscopy (cryo-EM) image analysis, often encounters challenges due to outliers, leading to inaccuracies in downstream processing. In response to this challenge, this research introduces an additional automated step to reduce the number of outliers identified by the particle picker. The proposed method enhances both the accuracy and efficiency of particle picking, thereby reducing the overall running time and the necessity for expert intervention in the process. Experimental results demonstrate the effectiveness of the proposed approach in mitigating outlier inclusion and its potential to enhance cryo-EM data analysis pipelines significantly. This work contributes to the ongoing advancement of automated cryo-EM image processing methods, offering novel insights and solutions to challenges in structural biology research.
The multi-reference alignment (MRA) problem involves reconstructing a signal from multiple noisy observations, each transformed by a random group element. In this paper, we focus on the group SO(2) of in-plane rotations and propose two computationally efficient algorithms with theoretical guarantees for accurate signal recovery under a non-uniform distribution over the group. The first algorithm exploits the spectral properties of the second moment of the data, while the second utilizes the frequency matching principle. Both algorithms achieve the optimal estimation rate in high-noise regimes, marking a significant advancement in the development of computationally efficient and statistically optimal methods for estimation problems over groups.
We study the bias-variance tradeoff within a multiscale approximation framework. Our approach uses a given quasi-interpolation operator, which is repeatedly applied within an error-correction scheme over a hierarchical data structure. We introduce a new bias measure, the bias ratio, to quantitatively assess the improvements afforded by multiscale approximations and demonstrate that this strategy effectively reduces the bias component of the approximation error, thereby providing an operator-level bias reduction framework for addressing scattered-data approximation problems. Our findings establish multiscale approximation as a bias-reduction methodology applicable to general quasi-interpolation operators, including applications to manifold-valued functions.
This work presents several new results concerning the analysis of the convergence of binary, univariate, and linear subdivision schemes, all related to the contractivity factor of a convergent scheme. First, we prove that a convergent scheme cannot have a contractivity factor lower than half. Since the lower this factor is, the faster the scheme's convergence, and schemes with contractivity factor 1/2, such as those generating spline functions, have optimal convergence rates. Additionally, we provide further insights and conditions for the convergence of linear schemes and demonstrate their applicability in an improved algorithm for determining the convergence of such subdivision schemes.
In synchronization problems, the goal is to estimate elements of a group from noisy measurements of their ratios. A popular estimation method for synchronization is the spectral method. It extracts the group elements from eigenvectors of a block matrix formed from the measurements. The eigenvectors must be projected, or 'rounded', onto the group. The rounding procedures are constructed ad hoc and increasingly so when applied to synchronization problems over non-compact groups. In this paper, we develop a spectral approach to synchronization over the non-compact group $\mathrm{SE}(3)$, the group of rigid motions of $\mathbb{R}<^>{3}$. We based our method on embedding $\mathrm{SE}(3)$ into the algebra of dual quaternions, which has deep algebraic connections with the group $\mathrm{SE}(3)$. These connections suggest a natural rounding procedure considerably more straightforward than the current state of the art for spectral $\mathrm{SE}(3)$ synchronization, which uses a matrix embedding of $\mathrm{SE}(3)$. We show by numerical experiments that our approach yields comparable results with the current state of the art in $\mathrm{SE}(3)$ synchronization via the spectral method. Thus, our approach reaps the benefits of the dual quaternion embedding of $\mathrm{SE}(3)$ while yielding estimators of similar quality.
Orbit recovery problems are a class of problems that often arise in practice and various forms. In these problems, we aim to estimate an unknown function after being distorted by a group action and observed via a known operator. Typically, the observations are contaminated with a non-trivial level of noise. Two particular orbit recovery problems of interest in this paper are multireference alignment and single-particle cryo-EM modeling. In order to suppress the noise, we suggest using the method of moments approach for both problems while introducing deep neural network priors. In particular, our neural networks should output the signals and the distribution of group elements, with moments being the input. In the multireference alignment case, we demonstrate the advantage of using the NN to accelerate the convergence for the reconstruction of signals from the moments. Finally, we use our method to reconstruct simulated and biological volumes in the cryo-EM setting.
In this paper, we take a new approach to autoregressive image generation that is based on two main ingredients. The first is wavelet image coding, which allows to tokenize the visual details of an image from coarse to fine details by ordering the information starting with the most significant bits of the most significant wavelet coefficients. The second is a variant of a language transformer whose architecture is re-designed and optimized for token sequences in this 'wavelet language'. The transformer learns the significant statistical correlations within a token sequence, which are the manifestations of well-known correlations between the wavelet subbands at various resolutions. We show experimental results with conditioning on the generation process.
This paper introduces a family of subdivision schemes that generate curves over manifolds from manifold-Hermite data. This data consists of points and tangent directions sampled from a curve over a manifold. Using a manifold-Hermite average based on the De Casteljau algorithm as our main building block, we show how to adapt a geometric approach for curve approximation over manifold-Hermite data. The paper presents the various definitions and provides several analysis methods for characterizing properties of both the average and the resulting subdivision schemes based on it. Demonstrative figures accompany the paper's presentation and analysis.
This paper addresses the problem of accurately estimating a function on one domain when only its discrete samples are available on another domain. To answer this challenge, we utilize a neural network, which we train to incorporate prior knowledge of the function. In addition, by carefully analyzing the problem, we obtain a bound on the error over the extrapolation domain and define a condition number for this problem that quantifies the level of difficulty of the setup. Compared to other machine learning methods that provide time series prediction, such as transformers, our approach is suitable for setups where the interpolation and extrapolation regions are general subdomains and, in particular, manifolds. In addition, our construction leads to an improved loss function that helps us boost the accuracy and robustness of our neural network. We conduct comprehensive numerical tests and comparisons of our extrapolation versus standard methods. The results illustrate the effectiveness of our approach in various scenarios.
Multiscale transforms have become a key ingredient in many data processing tasks. With technological development, we observe a growing demand for methods to cope with non-linear data structures such as manifold values. In this paper, we propose a multiscale approach for analyzing manifold-valued data using a pyramid transform. The transform uses a unique class of downsampling operators that enable a non-interpolating subdivision schemes as upsampling operators. We describe this construction in detail and present its analytical properties, including stability and coefficient decay. Next, we numerically demonstrate the results and show the application of our method to denoising and anomaly detection.
We address the problem of approximating an unknown function from its discrete samples given at arbitrarily scattered sites. This problem is essential in numerical sciences, where modern applications also highlight the need for a solution to the case of functions with manifold values. In this paper, we introduce and analyze a combination of kernel-based quasi-interpolation and multiscale approximations for both scalar- and manifold-valued functions. While quasi-interpolation provides a powerful tool for approximation problems if the data is defined on infinite grids, the situation is more complicated when it comes to scattered data. Here, higher-order quasi-interpolation schemes either require derivative information or become numerically unstable. Hence, this paper principally studies the improvement achieved by combining quasi-interpolation with a multiscale technique. The main contributions of this paper are as follows. First, we introduce the multiscale quasi-interpolation technique for scalar-valued functions. Second, we show how this technique can be carried over using moving least-squares operators to the manifold-valued setting. Third, we give a mathematical proof that converging quasi-interpolation will also lead to converging multiscale quasi-interpolation. Fourth, we provide ample numerical evidence that multiscale quasi-interpolation has superior convergence to quasi-interpolation. In addition, we will provide examples showing that the multiscale quasi-interpolation approach offers a powerful tool for many data analysis tasks, such as denoising and anomaly detection. It is especially attractive for cases of massive data points and high dimensionality.