We propose a method for sampling from Gibbs distributions of the form i(x) o* exp(*U(x)) that leverages a family (it)t of approximations of the target density which is deliberately constructed such that it exhibits favorable properties for sampling when t is large and such that it approaches i as t approaches 0. This sequence is obtained by replacing (parts of) the potential U with its Moreau envelope. Through the sequential sampling from it for decreasing values oft by a Langevin algorithm with appropriate step size, the samples are guided from a simple starting density to the more complex target quickly. We prove that t \mapsto-* it is Lipschitz continuous in the total variation distance and Ho"\lder continuous in the Wasserstein-p distance, that the sampling algorithm is ergodic, and that it converges to the target density without assuming convexity or differentiability of the potential U. In addition to the theoretical analysis, we show experimental results that support the superiority of the method in terms of convergence speed and mode-coverage of multimo dal densities over current algorithms. The experiments range from one-dimensional toy-problems to high-dimensional inverse imaging problems with learned potentials.
Diffusion models play an important role in many state-of-the-art algorithms to solve inverse problems in imaging. In particular, reconstructing MRI and CT images from undersampled data can be reformulated as a canonical inverse problem that benefits from strong priors. Researchers typically re-purpose models originally designed for unconditional sampling without modifications and achieve remarkable accuracy on in-distribution (ID) reconstruction. However, due to the scarce availability of training data and the large number of different imaging setups, increasing the generalization capabilities of diffusion-based priors is key to clinical adoption. To do so, we propose two solutions: (i) using smaller models and (ii) training on natural images. Using three different posterior sampling algorithms, we evaluate the influence of network size and training data. Our smallest model, effectively a ResNet, performs almost as good as an attention U-Net on ID reconstruction, while being significantly more robust towards distribution shifts. Furthermore, we introduce models trained on natural images and demonstrate that they can be used in both MRI and CT reconstruction, outperforming models trained on medical images in OOD cases. As a result of our findings, we strongly caution against simply re-using very large networks and encourage researchers to adapt the model complexity to the respective task.
The semi-arid regions of the world are populated by highly specialized groups of organisms that can cope with the harsh climatic conditions. However, the climate and, as a result, the floristic composition have changed in these regions in recent decades. Dryland regions throughout the world host biological soil crusts, which colonize the uppermost soil layer. This superficial growth facilitates the utilization of imaging methods for monitoring purposes. In this study a deep-learning model called Joint Energy-Based Semantic Segmentation that enables robust analysis of images captured over long periods of time is proposed. It is shown how biological soil crusts have changed over time in two very different areas (Succulent Karoo, South Africa and Colorado Plateau, USA) with an accuracy of 91 % and 77 %, respectively. This provides a detailed analysis of the complex interactions between the individual taxa and external climatic influences. The results show that conditions of extreme drought led to degradation and that the soil crust organisms were unable to fully recover from this during wetter periods. On both sites, biocrusts degraded within the last 1-2 decades. A detailed time series analysis of the interactions between the occurring taxa, using time lagged cross correlation and transfer entropy metrics, identified Psora sp. and Fulgensia sp. as key indicator species, as they were highly reactive to climate alterations, and therefore inform about biocrust degradation already at an early state. This study demonstrates how modern image-based deep learning methods enable a very detailed analysis of the development of the world's dryland flora.
Langevin sampling from distributions of the form $p(x) \propto \exp(-Ψ(x))$ faces two major challenges: (global) mode coverage and (local) mode exploration. The first challenge is particularly relevant for multi-modal distributions with disjoint modes, whereas the second arises when the potential $Ψ$ exhibits diverse and ill-conditioned local mode geometry. To address these challenges, a common approach is to precondition Langevin dynamics with problem-specific information, such as the sample covariance or the local curvature of $Ψ$. However, existing preconditioner choices inherently involve a trade-off between global mode coverage and local mode exploration, and no prior method resolves both simultaneously. To overcome this limitation, we propose the TIPreL, which introduces a time- and position-dependent preconditioner. This design effectively addresses both challenges mentioned above within a single framework. We establish convergence of the resulting dynamics in the Wasserstein-2 distance both in continuous time and for a tamed Euler discretization. In particular, our analysis extends the existing state of the art by proving convergence under time- and space-dependent diffusion coefficients, and only locally Lipschitz drifts, which has not been covered by prior work. Finally, we experimentally compare TIPreL with competing preconditioning schemes on a two-dimensional, severely ill-posed example and on a Bayesian logistic regression task in higher dimensions, confirming the efficiency of the proposed method.
Experimentally acquired microscopy images are unavoidably affected by the presence of noise and other unwanted signals, which degrade their quality and might hide relevant features. With the recent increase in image acquisition rate, modern denoising and restoration solutions become necessary. This study focuses on image decomposition and denoising of microscopy images through a workflow based on total variation (TV), addressing images obtained from various microscopy techniques, including atomic force microscopy (AFM), scanning tunneling microscopy (STM), and scanning electron microscopy (SEM). Our approach consists in restoring an image by extracting its unwanted signal components and subtracting them from the raw one, or by denoising it. We evaluate the performance of TV-L^1, Huber-ROF, and TGV-L^1 in achieving this goal in distinct study cases. Huber-ROF proved to be the most flexible one, while TGV-L^1 is the most suitable for denoising. Our results suggest a wider applicability of this method in microscopy, restricted not only to STM, AFM, and SEM images. The Python code used for this study is publicly available as part of AiSurf. It is designed to be integrated into experimental workflows for image acquisition or can be used to denoise previously acquired images.
Molecules in equilibrium follow a Boltzmann distribution, making the underlying energy landscape a physically grounded modeling objective. However, such landscapes are difficult to learn from data and, once learned, hard to sample from. Diffusion and flow-matching models sidestep these difficulties by learning a time-conditional score or transport field between noise and data, losing the energy inductive bias in exchange for a more tractable training objective. We introduce EBMol, an energy-based model (EBM) that restores this inductive bias by learning an atom-additive scalar potential without explicit simulation during training. Our method employs a flow-inspired Restoring Field Matching objective to approximate the energy landscape. We adopt the Mirror-Langevin algorithm for sampling, enabling unified updates of atomic positions and types, and incorporate parallel tempering for inference-time compute scaling. EBMol is the first EBM for 3D molecular generation to achieve state-of-the-art performance on QM9 and GEOM-Drugs. Moreover, we show that the learned energy landscape serves as a principled quality metric for ranking and filtering configurations, and demonstrate controllable generation without retraining through shape-steered sampling via potential composition and zero-shot linker design.
Recently, diffusion models have attracted considerable attention for magnetic resonance image reconstruction due to their high sample quality. However, most existing methods rely on large networks with opaque time-conditioning mechanisms, and require offline coil sensitivity estimation. This results in limited interpretability of the reconstruction process and reduced flexibility in the acquisition setup. To address these limitations, we jointly reconstruct the image and the coil sensitivities by combining the parameter-efficient product-of-Gaussian-mixture diffusion model as an image prior with a classical smoothness prior on the coil sensitivities. The proposed method is fast and robust to both contrast and anatomical distribution shifts as well as changing k-space trajectories. Finally, we propose a more expressive parameterization of the image prior which improves results in denoising and magnetic resonance image reconstruction.
BACKGROUND:It is unclear, whether photoplethysmography (PPG) waveforms from wearable devices can differentiate between supraventricular and ventricular arrhythmias. We assessed, whether a neural network-based classifier can distinguish the origin of PPG pulse waveforms. METHODS:In thirty patients undergoing invasive electrophysiological (EP) studies for narrow complex tachycardia, PPG waveforms were recorded using a PPG wristband (Empatica E4) in parallel to 12-lead surface electrocardiograms (ECGs) and intracardiac bipolar electrograms. PPG waveforms were annotated to either atrial (AP, supraventricular) or ventricular pacing (VP) based on bipolar electrograms, ECGs and stimulation protocols. 25 221 samples were split into training, testing, and validation data sets and used to develop, optimize and validate a residual network based on convolutional layers for classifying PPG waveforms according to their origin into AP or VP. RESULTS:Datasets were complete for 27 patients. 74 % were female, median age was 53 (range 18, 78) years and median BMI was 27±5 kg/m². The electrophysiological study revealed typical atrioventricular nodal re-entrant tachycardias in 63 %, atrial tachycardias in 15 % and no inducible tachyarrhythmias in 12 % of patients. On an independent patient level, correct prediction was possible in ∼73 % for AP and ∼59 % for VP. With adaptive performance built on previous patient-specific annotations, the classifier correctly predicted the origins of PPG-derived pulse waves in ∼97 % for AP and ∼95 % for VP. CONCLUSIONS:A neural network trained on ground truth PPG data collected during EP studies could distinguish between supraventricular or ventricular origin from PPG waveforms alone.
In this chapter we provide a thorough overview of the use of energy-based models (EBMs) in the context of inverse imaging problems. EBMs are probability distributions modeled via Gibbs densities p(x) ∝exp-E(x) with an appropriate energy functional E. Within this chapter we present a rigorous theoretical introduction to Bayesian inverse problems that includes results on well-posedness and stability in the finite-dimensional and infinite-dimensional setting. Afterwards we discuss the use of EBMs for Bayesian inverse problems and explain the most relevant techniques for learning EBMs from data. As a crucial part of Bayesian inverse problems, we cover several popular algorithms for sampling from EBMs, namely the Metropolis-Hastings algorithm, Gibbs sampling, Langevin Monte Carlo, and Hamiltonian Monte Carlo. Moreover, we present numerical results for the resolution of several inverse imaging problems obtained by leveraging an EBM that allows for the explicit verification of those properties that are needed for valid energy-based modeling.
Medical image segmentation plays an important role in accurately identifying and isolating regions of interest within medical images. Generative approaches are particularly effective in modeling the statistical properties of segmentation masks that are closely related to the respective structures. In this work we introduce FlowSDF, an image-guided conditional flow matching framework, designed to represent the signed distance function (SDF), and, in turn, to represent an implicit distribution of segmentation masks. The advantage of leveraging the SDF is a more natural distortion when compared to that of binary masks. Through the learning of a vector field associated with the probability path of conditional SDF distributions, our framework enables accurate sampling of segmentation masks and the computation of relevant statistical measures. This probabilistic approach also facilitates the generation of uncertainty maps represented by the variance, thereby supporting enhanced robustness in prediction and further analysis. We qualitatively and quantitatively illustrate competitive performance of the proposed method on a public nuclei and gland segmentation data set, highlighting its utility in medical image segmentation applications.
We present a novel method for drawing samples from Gibbs distributions with densities of the form π(x) ∝exp(-U(x)). The method accelerates the unadjusted Langevin algorithm by introducing an inertia term similar to Polyak's heavy ball method, together with a corresponding noise rescaling. Interpreting the scheme as a discretization of kinetic Langevin dynamics, we prove ergodicity (in continuous and discrete time) for twice continuously differentiable, strongly convex, and L-smooth potentials and bound the bias of the discretization to the target in Wasserstein-2 distance. In particular, the presented proofs allow for smaller friction parameters in the kinetic Langevin diffusion compared to existing literature. Moreover, we show the close ties of the proposed method to the over-relaxed Gibbs sampler. The scheme is tested in an extensive set of numerical experiments covering simple toy examples, total variation image denoising, and the complex task of maximum likelihood learning of an energy-based model for molecular structure generation. The experimental results confirm the acceleration provided by the proposed scheme even beyond the strongly convex and L-smooth setting.
Typically, sequences designed de novo are assessed in silico using deep learning-based protein structure prediction methods prior to wetlab testing. While these deep learning (DL) models excel at predicting well-ordered regions, accurate prediction of loop regions, which often are flexible and crucial for protein function, remains a significant challenge. To address this, we introduce the Equivariant Loop Evaluation Network (ELEN), a local model quality assessment (MQA) method that is tailored towards evaluating the accuracy of protein loops at the per-residue level. ELEN jointly predicts three quality metrics, local Distance Difference Test (lDDT), Contact Area Difference Score (CAD-score), and Root Mean Squared Deviation (RMSD), by comparing predicted to experimental reference structures. Learning these metrics simultaneously enables ELEN to capture complementary structural insights, providing a richer assessment of model accuracy. The network operates at all-atom resolution and employs 3D equivariant group convolutions to learn the local geometric environment of each atom. By incorporating sequence embeddings from large language models (LLMs), such as SaProt, we enhance the sequence and evolutionary awareness of the model. Furthermore, by informing ELEN with per-residue physicochemical features, the model achieves competitive accuracy relative to state-of-the-art MQA methods on the Continuous Automated Model EvaluatiOn (CAMEO) benchmark. Although ELEN was primarily developed for assessing loop quality, its architecture also demonstrates strong potential for general MQA tasks. We used ELEN to perform detailed analysis, including identification of flexible or disordered regions and assessment of structural effects from single-residue mutations on three sets of redesigned enzymes. We show that for all sets ELEN successfully identifies poor design positions and thus serves as a powerful tool for advancing both the study and modeling of loops in protein structures. ### Competing Interest Statement The authors have declared no competing interest. European Research Council, https://ror.org/0472cxd90, 802217
We consider the problem of sampling from a product-of-experts-type model that encompasses many standard prior and posterior distributions commonly found in Bayesian imaging. We show that this model can be easily lifted into a novel latent variable model, which we refer to as a Gaussian latent machine. This leads to a general sampling approach that unifies and generalizes many existing sampling algorithms in the literature. Most notably, it yields a highly efficient and effective two-block Gibbs sampling approach in the general case, while also specializing to direct sampling algorithms in particular cases. Finally, we present detailed numerical experiments that demonstrate the efficiency and effectiveness of our proposed sampling approach across a wide range of prior and posterior sampling problems from Bayesian imaging.
Diffusion models have recently shown remarkable results in magnetic resonance imaging reconstruction. However, the employed networks typically are black-box estimators of the (smoothed) prior score with tens of millions of parameters, restricting interpretability and increasing reconstruction time. Furthermore, parallel imaging reconstruction algorithms either rely on off-line coil sensitivity estimation, which is prone to misalignment and restricting sampling trajectories, or perform per-coil reconstruction, making the computational cost proportional to the number of coils. To overcome this, we jointly reconstruct the image and the coil sensitivities using the lightweight, parameter-efficient, and interpretable product of Gaussian mixture diffusion model as an image prior and a classical smoothness priors on the coil sensitivities. The proposed method delivers promising results while allowing for fast inference and demonstrating robustness to contrast out-of-distribution data and sampling trajectories, comparable to classical variational penalties such as total variation. Finally, the probabilistic formulation allows the calculation of the posterior expectation and pixel-wise variance.
We consider a bilevel learning framework for learning linear operators. In this framework, the learnable parameters are optimized via a loss function that also depends on the minimizer of a convex optimization problem (denoted lower-level problem). We utilize an iterative algorithm called `piggyback' to compute the gradient of the loss and minimizer of the lower-level problem. Given that the lower-level problem is solved numerically, the loss function and thus its gradient can only be computed inexactly. To estimate the accuracy of the computed hypergradient, we derive an a-posteriori error bound, which provides guides for setting the tolerance for the lower-level problem, as well as the piggyback algorithm. To efficiently solve the upper-level optimization, we also propose an adaptive method for choosing a suitable step-size. To illustrate the proposed method, we consider a few learned regularizer problems, such as training an input-convex neural network.
Machine learning methods are increasingly used in cardiovascular research. In order to highlight opportunities and challenges of the evaluation of studies applying machine learning, we use examples from cardiac electrophysiology, a field characterized by large and often imbalanced amounts of data. We provide recommendations and guidance on evaluating and presenting supervised machine learning studies. We recommend proper cohort selection, keeping training and testing data strictly separate, and comparing results to a reference model without machine learning as basic principles to ensure the quality of studies using machine learning methods. We furthermore recommend specific metrics and plots when reporting on machine learning including on models for multi-channel time series or images. This Best Practice paper represents a possible blueprint to help evaluate machine learning-based medical tests in cardiac electrophysiology and beyond.
In this work, we tackle the problem of estimating the density f_X of a random variable X by successive smoothing, such that the smoothed random variable Y fulfills the diffusion partial differential equation (∂ _t - Δ _1)f_Y( · , t) = 0 with initial condition f_Y( · , 0) = f_X . We propose a product-of-experts-type model utilizing Gaussian mixture experts and study configurations that admit an analytic expression for f_Y ( · , t) . In particular, with a focus on image processing, we derive conditions for models acting on filter, wavelet, and shearlet responses. Our construction naturally allows the model to be trained simultaneously over the entire diffusion horizon using empirical Bayes. We show numerical results for image denoising where our models are competitive while being tractable, interpretable, and having only a small number of learnable parameters. As a by-product, our models can be used for reliable noise level estimation, allowing blind denoising of images corrupted by heteroscedastic noise.
We examine the assumption that the hidden-state vectors of recurrent neural networks (RNNs) tend to form clusters of semantically similar vectors, which we dub the clustering hypothesis. While this hypothesis has been assumed in the analysis of RNNs in recent years, its validity has not been studied thoroughly on modern neural network architectures. We examine the clustering hypothesis in the context of RNNs that were trained to recognize regular languages. This enables us to draw on perfect ground-truth automata in our evaluation, against which we can compare the RNN's accuracy and the distribution of the hidden-state vectors. We start with examining the (piecewise linear) separability of an RNN's hidden-state vectors into semantically different classes. We continue the analysis by computing clusters over the hidden-state vector space with multiple state-of-the-art unsupervised clustering approaches. We formally analyze the accuracy of computed clustering functions and the validity of the clustering hypothesis by determining whether clusters group semantically similar vectors to the same state in the ground-truth model. Our evaluation supports the validity of the clustering hypothesis in the majority of examined cases. We observed that the hidden-state vectors of well-trained RNNs are separable, and that the unsupervised clustering techniques succeed in finding clusters of similar state vectors.
In this work, a method for obtaining pixel-wise error bounds in Bayesian regularization of inverse imaging problems is introduced. The proposed method employs estimates of the posterior variance together with techniques from conformal prediction in order to obtain coverage guarantees for the error bounds, without making any assumption on the underlying data distribution. It is generally applicable to Bayesian regularization approaches, independent, e.g., of the concrete choice of the prior. Furthermore, the coverage guarantees can also be obtained in case only approximate sampling from the posterior is possible. With this in particular, the proposed framework is able to incorporate any learned prior in a black-box manner. Guaranteed coverage without assumptions on the underlying distributions is only achievable since the magnitude of the error bounds is, in general, unknown in advance. Nevertheless, experiments with multiple regularization approaches presented in the paper confirm that in practice, the obtained error bounds are rather tight. For realizing the numerical experiments, also a novel primal-dual Langevin algorithm for sampling from non-smooth distributions is introduced in this work.
Biological soil crusts (biocrusts) form a layer of only one to few centimeters depth on the soil surface and occur mostly in hot and cold deserts. Biocrusts have a major impact on different processes in these ecosystems, like carbon and nitrogen cycling, biodiversity preservation, erosion protection and soil dust emission reduction, but also react highly sensitive upon climate alterations and land use intensification. Therefore, monitoring tools are required to keep track of the changes of these specialized communities in an altering environment. In the current study, we applied a semantic image segmentation approach, using neural networks. One main problem to be solved was, that the training data and target data, on which the model is applied, are often recorded with different camera devices. This leads to different statistical properties of the image data, like different scale, resolution, brightness etc., which could significantly affect the model's performance. To solve this problem, we propose a new domain adaption method using a joint energy-based approach. To test a semantic segmentation approach in general, we utilized biocrust imagery taken in Utah (United States of America) and two sub datasets from the National Park Gesäuse (Austria). Here, we achieved highly reliable results with an overall classification accuracy of 85.9% for the USA data and 88.6% and 91.4%, respectively, for the two sub datasets of the National Park Gesäuse. To test our joint energy-based domain adaption approach, we used the two sub datasets from the National Park Gesäuse, which were recorded with different camera devices. With this newly established approach, we improved the accuracy of our segmentation on the unlabeled sub dataset from 70.4% to 75.3%. The results suggest that joint energy-based modelling is a well-suited domain adaption method for semantic segmentation that could be applied to face various deep learning and image-based biomonitoring challenges.