Uncertainty quantification plays an important role in achieving trustworthy and reliable learning-based computational imaging. Recent advances in generative modeling and Bayesian neural networks have enabled the development of uncertainty-aware image reconstruction methods. Current generative model-based methods seek to quantify the inherent (aleatoric) uncertainty on the underlying image for given measurements by learning to sample from the posterior distribution of the underlying image. On the other hand, Bayesian neural network-based approaches aim to quantify the model (epistemic) uncertainty on the parameters of a deep neural network-based reconstruction method by approximating the posterior distribution of those parameters. Unfortunately, an ongoing need for an inversion method that can jointly quantify complex aleatoric uncertainty and epistemic uncertainty patterns still persists. In this paper, we present a scalable framework that can quantify both aleatoric and epistemic uncertainties. The proposed framework accepts an existing generative model-based posterior sampling method as an input and introduces an epistemic uncertainty quantification capability through Bayesian neural networks with latent variables and deep ensembling. Furthermore, by leveraging the conformal prediction methodology, the proposed framework can be easily calibrated to ensure rigorous uncertainty quantification. We evaluated the proposed framework on magnetic resonance imaging, computed tomography, and image inpainting problems and showed that the epistemic and aleatoric uncertainty estimates produced by the proposed framework display the characteristic features of true epistemic and aleatoric uncertainties. Furthermore, our results demonstrated that the use of conformal prediction on top of the proposed framework enables marginal coverage guarantees consistent with frequentist principles.
Spatially varying image deblurring remains a fundamentally ill-posed problem, especially when degradations arise from complex mixtures of motion and other forms of blur under significant noise. State-of-the-art learning-based approaches generally fall into two paradigms: model-based deep unrolling methods that enforce physical constraints by modeling the degradations, but often produce over-smoothed, artifact-laden textures, and generative models that achieve superior perceptual quality yet hallucinate details due to weak physical constraints. In this paper, we propose a novel framework that uniquely reconciles these paradigms by taming a powerful generative prior with explicit, dense physical constraints. Rather than oversimplifying the degradation field, we model it as a dense continuum of high-dimensional compressed kernels, ensuring that minute variations in motion and other degradation patterns are captured. We leverage this rich descriptor field to condition a ControlNet architecture, strongly guiding the diffusion sampling process. Extensive experiments demonstrate that our method effectively bridges the gap between physical accuracy and perceptual realism, outperforming state-of-the-art model-based methods as well as generative baselines in challenging, severely blurred scenarios.
In cataract surgery, the opacified crystalline lens is replaced by an artificial intraocular lens (IOL), requiring precise preoperative selection of parameters to optimize postoperative visual quality. Three-dimensional customized eye models, which can be constructed using quantitative data from anterior segment optical coherence tomography, provide a robust platform for virtual surgery. These models enable simulations and predictions of the optical outcomes for specific patients and selected IOLs. A critical step in building these models is estimating the IOL's tilt and position preoperatively based on the available preoperative geometrical information (ocular parameters). In this study, we present a machine learning model that, for the first time, incorporates the full shape geometry of the crystalline lens as candidate input features to predict the postoperative IOL tilt. Furthermore, we identify the most relevant features for this prediction task. Our model demonstrates statistically significantly lower estimation errors compared to a simple linear correlation method, reducing the estimation error by approximately 6%. These findings highlight the potential of this approach to enhance the accuracy of postoperative predictions. Further work is needed to examine the potential for such postoperative predictions to improve visual outcomes in cataract patients.
Abstract:Our planet’s mantle is the largest rock-layer by volume. Across its old and stable Archean and Proterozoic terranes, seismological evidence suggests ubiquitous, spatially variable, and puzzling discontinuities, within, across and beneath the upper mantle lithosphere (~50- 350 km). A variety of explanations have been proposed, including phase transformations, melting and compositional anomalies, anisotropy, and elastically accommodated grain. To evaluate these, and other models, it is crucial to improve our threshold for detecting such discontinuities especially in reverberant and noisy environments. Here, we present a new method for sifting through the echoes and reverberations: CRISP-RF (Clean Receiver function Imaging with Sparse Radon Filters). With a global dataset of Ps converted waves, we use CRISP-RF to isolate hard-to-detect wave conversions buried in reverberations and noise. This refined, high-resolution, global view of upper mantle stratification will ensure robust evaluation of proposed models of upper mantle structure, evolution, and dynamics.
Recently, physiological data such as electroencephalography (EEG) signals have attracted significant attention in affective computing. In this context, the main goal is to design an automated model that can assess emotional states. Lately, deep neural networks have shown promising performance in emotion recognition tasks. However, designing a deep architecture that can extract practical information from raw data is still a challenge. Here, we introduce a deep neural network that acquires interpretable physiological representations by a hybrid structure of spatio-temporal encoding and recurrent attention network blocks. Furthermore, a preprocessing step is applied to the raw data using graph signal processing tools to perform graph smoothing in the spatial domain. We demonstrate that our proposed architecture exceeds state-of-the-art results for emotion classification on the publicly available DEAP dataset. To explore the generality of the learned model, we also evaluate the performance of our architecture towards transfer learning (TL) by transferring the model parameters from a specific source to other target domains. Using DEAP as the source dataset, we demonstrate the effectiveness of our model in performing cross-modality TL and improving emotion classification accuracy on DREAMER and the Emotional English Word (EEWD) datasets, which involve EEG-based emotion classification tasks with different stimuli.
Ultrasound elastography images, which enable quantitative visualization of tissue stiffness, can be reconstructed by solving an inverse problem. Classical model-based methods are usually formulated in terms of constrained optimization problems. To stabilize the elasticity reconstructions, regularization techniques, such as Tikhonov method, are used with the cost of promoting smoothness and blurriness in the reconstructed images. Thus, incorporating a suitable regularizer is essential for reducing the elasticity reconstruction artifacts, while finding the most suitable one is challenging. In this work, we present a new statistical representation of the physical imaging model, which incorporates effective signal-dependent colored noise modeling. Moreover, we develop a learning-based integrated statistical framework, which combines a physical model with learning-based priors. We use a dataset of simulated phantoms with various elasticity distributions and geometric patterns to train a denoising regularizer as the learning-based prior. We use fixed-point approaches and variants of gradient descent for solving the integrated optimization task following learning-based plug-and-play (PnP) prior and regularization by denoising (RED) paradigms. Finally, we evaluate the performance of the proposed approaches in terms of relative mean square error (RMSE) with nearly 20% improvement for both piecewise smooth simulated phantoms and experimental phantoms compared with the classical model-based methods and 12% improvement for both spatially varying breast-mimicking simulated phantoms and an experimental breast phantom, demonstrating the potential clinical relevance of our work. Moreover, the qualitative comparisons of reconstructed images demonstrate the robust performance of the proposed methods even for complex elasticity structures that might be encountered in clinical settings.
The coherent nature of imaging in synthetic aperture radar (SAR) inevitably gives rise to speckle noise, a challenge exacerbated by the constrained bandwidth and limited look angles. Among the despeckling algorithms, convolutional neural networks (CNNs) have dominated the forefront of SAR image despeckling, showcasing state-of-the-art performance. CNN-based SAR image despeckling methods excel in learning complex features in an image, while they often struggle with a tendency to induce blurring, leading to a loss of texture information in SAR imagery. We hypothesize that this is mainly caused by the inappropriate use of generic loss functions for despeckling that tend to yield a washed-out blurring effect or introduce artifacts during the denoising process. In this article, we propose a new loss function for an existing CNN architecture that is designed to reduce washed-out blurring effects and is capable of enhancing edges and small-scale features while suppressing noise. The loss function we propose is based on the logarithm of discrete cosine transform images, with a perspective to maintain a nuanced equilibrium between high- and low-energy features in an image, meanwhile maintaining acceptable noise suppression. In comparison to established loss functions applied in SAR image despeckling, experimental results show that the CNNs trained with the proposed loss function not only enhance multiple objective metrics but also exhibit considerable advantages in terms of visual effects.
Automated emotion recognition using electroencephalogram (EEG) signals has gained substantial attention. Although deep learning approaches exhibit strong performance, they often suffer from vulnerabilities to various perturbations, like environmental noise and adversarial attacks. In this paper, we propose an Inception feature generator and two-sided perturbation (INC-TSP) approach to enhance emotion recognition in brain-computer interfaces. INC-TSP integrates the Inception module for EEG data analysis and employs two-sided perturbation (TSP) as a defensive mechanism against input perturbations. TSP introduces worst-case perturbations to the model’s weights and inputs, reinforcing the model’s elasticity against adversarial attacks. The proposed approach addresses the challenge of maintaining accurate emotion recognition in the presence of input uncertainties. We validate INC-TSP in a subject-independent three-class emotion recognition scenario, demonstrating robust performance.
Reconstructing a range profile from radar returns, which are both noisy and band-limited, presents a challenging and ill-posed inverse problem.Conventional reconstruction methods often involve employing matched filters in pulsed radars or performing a Fourier transform of the received signal in continuous wave radars.However, both of these approaches rely on specific models and model-based inversion techniques that may not fully leverage prior knowledge of the range profiles being reconstructed when such information is accessible.To incorporate prior distribution information of the range profile data into the reconstruction process, regularizers can be employed to encourage specific spatial patterns within the range profiles.Nevertheless, these regularizers often fall short in effectively capturing the intricate spatial correlations within the range profile data, or they may not readily allow for analytical minimization of the cost function.Recently, Alternating Direction Method of Multipliers (ADMM) framework has emerged as a means to provide a way of decoupling the model inversion from the regularization of the priors, enabling the incorporation of any desired regularizer into the inversion process in a plug-and-play (PnP) fashion.In this paper, we implement an ADMM framework to address the radar range profile reconstruction problem where we propose to employ a Convolutional Neural Network (CNN) as a regularization method for enhancing the quality of the inversion process which usually suffers from the ill-posed nature of the problem.We demonstrate the efficacy of deep learning networks as a regularization method within the ADMM framework through our simulation results.We assess the performance of the ADMM framework employing CNN as a regularizer and conduct a comparative analysis against alternative methods under different measurement scenarios.Notably, among the methods under investigation, ADMM with CNN as a regularizer stands out as the most successful method for radar range profile reconstruction.
Ptychography is a scanning coherent diffractive imaging technique that enables imaging nanometer-scale features in extended samples. One main challenge is that widely used iterative image reconstruction methods often require significant amount of overlap between adjacent scan locations, leading to large data volumes and prolonged acquisition times. To address this key limitation, this paper proposes a Bayesian inversion method for ptychography that performs effectively even with less overlap between neighboring scan locations. Furthermore, the proposed method can quantify the inherent uncertainty on the ptychographic object, which is created by the ill-posed nature of the ptychographic inverse problem. At a high level, the proposed method first utilizes a deep generative model to learn the prior distribution of the object and then generates samples from the posterior distribution of the object by using a Markov Chain Monte Carlo algorithm. Our results from simulated ptychography experiments show that the proposed framework can consistently outperform a widely used iterative reconstruction algorithm in cases of reduced overlap. Moreover, the proposed framework can provide uncertainty estimates that closely correlate with the true error, which is not available in practice. The project website is available here.
Motor imagery (MI) classification based on electroencephalogram (EEG) is a widely-used technique in non-invasive brain-computer interface (BCI) systems. Since EEG recordings suffer from heterogeneity across subjects and labeled data insufficiency, designing a classifier that performs the MI independently from the subject with limited labeled samples would be desirable. To overcome these limitations, we propose a novel subject-independent semi-supervised deep architecture (SSDA). The proposed SSDA consists of two parts: an unsupervised and a supervised element. The training set contains both labeled and unlabeled data samples from multiple subjects. First, the unsupervised part, known as the columnar spatiotemporal auto-encoder (CST-AE), extracts latent features from all the training samples by maximizing the similarity between the original and reconstructed data. A dimensional scaling approach is employed to reduce the dimensionality of the representations while preserving their discriminability. Second, a supervised part learns a classifier based on the labeled training samples using the latent features acquired in the unsupervised part. Moreover, we employ center loss in the supervised part to minimize the embedding space distance of each point in a class to its center. The model optimizes both parts of the network in an end-to-end fashion. The performance of the proposed SSDA is evaluated on test subjects who were not seen by the model during the training phase. To assess the performance, we use two benchmark EEG-based MI task datasets. The results demonstrate that SSDA outperforms state-of-the-art methods and that a small number of labeled training samples can be sufficient for strong classification performance.
Although deep learning-based algorithms have demonstrated excellent performance in automated emotion recognition via electroencephalogram (EEG) signals, variations across brain signal patterns of individuals can diminish the model's effectiveness when applied across different subjects. While transfer learning techniques have exhibited promising outcomes, they still encounter challenges related to inadequate feature representations and may overlook the fact that source subjects themselves can possess distinct characteristics. In this work, we propose a multi-source domain adaptation approach with a transformer-based feature generator (MSDA-TF) designed to leverage information from multiple sources. The proposed feature generator retains convolutional layers to capture shallow spatial, temporal, and spectral EEG data representations, while self-attention mechanisms extract global dependencies within these features. During the adaptation process, we group the source subjects based on correlation values and aim to align the moments of the target subject with each source as well as within the sources. MSDA-TF is validated on the SEED dataset and is shown to yield promising results.
Efficiently identifying sleep stages is crucial for unraveling the intricacies of sleep in both preclinical and clinical research. The labor-intensive nature of manual sleep scoring, demanding substantial expertise, has prompted a surge of interest in automated alternatives. Sleep studies in mice play a significant role in understanding sleep patterns and disorders and underscore the need for robust scoring methodologies. In response, this study introduces LG-Sleep, a novel subject-independent deep neural network architecture designed for mice sleep scoring through electroencephalogram (EEG) signals. LG-Sleep extracts local and global temporal transitions within EEG signals to categorize sleep data into three stages: wake, rapid eye movement (REM) sleep, and non-rapid eye movement (NREM) sleep. The model leverages local and global temporal information by employing time-distributed convolutional neural networks to discern local temporal transitions in EEG data. Subsequently, features derived from the convolutional filters traverse long short-term memory blocks, capturing global transitions over extended periods. Crucially, the model is optimized in an autoencoder-decoder fashion, facilitating generalization across distinct subjects and adapting to limited training samples. Experimental findings demonstrate superior performance of LG-Sleep compared to conventional deep neural networks. Moreover, the model exhibits good performance across different sleep stages even when tasked with scoring based on limited training samples.
Seismic interrogation of the upper mantle from the base of the crust to the top of the mantle transition zone has revealed discontinuities that are variable in space, depth, lateral extent, amplitude, and lack a unified explanation for their origin. Improved constraints on the detectability and properties of mantle discontinuities can be obtained with P-to-S receiver function (Ps-RF) where energy scatters from P to S as seismic waves propagate across discontinuities of interest. However, due to the interference of crustal multiples, uppermost mantle discontinuities are more commonly imaged with lower resolution S-to-P receiver function (Sp-RF). In this study, a new method called CRISP-RF (Clean Receiver-function Imaging using SParse Radon Filters) is proposed, which incorporates ideas from compressive sensing and model-based image reconstruction. The central idea involves applying a sparse Radon transform to effectively decompose the Ps-RF into its underlying wavefield contributions, i.e., direct conversions, multiples, and noise, based on the phase moveout and coherence. A masking filter is then designed and applied to create a multiple-free and denoised Ps-RF. We demonstrate, using synthetic experiment, that our implementation of the Radon transform using a sparsity-promoting regularization outperforms the conventional least-squares methods and can effectively isolate direct Ps conversions. We further apply the CRISP-RF workflow on real data, including single station data on cratons, common-conversion-point (CCP) stack at continental margins, and seismic data from ocean islands. The application of CRISP-RF to global datasets will advance our understanding of the enigmatic origins of the upper mantle discontinuities like the ubiquitous Mid-Lithospheric Discontinuity (MLD) and the elusive X-discontinuity.
The Signal Processing Society is an organization, within the framework of the IEEE, of members with principal professional interest in the technology of transmission, recording, reproduction, processing, and measurement of speech; other audio-frequency waves and other signals by digital electronic, electrical, acoustic, mechanical, and optical means; the components and systems to accomplish these and related aims; and the environmental, psychological, and physiological factors of these technologies.For membership and subscription information and pricing, please visit www.
High-content imaging techniques in conjunction with in vitro microphysiological systems (MPS) allow for novel explorations of physiological phenomena with a high degree of translational relevance due to the usage of human cell lines. MPS featuring ultrathin and nanoporous silicon nitride membranes (µSiM) have been utilized in the past to facilitate high magnification phase contrast microscopy recordings of leukocyte trafficking events in a living mimetic of the human vascular microenvironment. Notably, the imaging plane can be set directly at the endothelial interface in a µSiM device, resulting in a high-resolution capture of an endothelial cell (EC) and leukocyte coculture reacting to different stimulatory conditions. The abundance of data generated from recording observations at this interface can be used to elucidate disease mechanisms related to vascular barrier dysfunction, such as sepsis. The appearance of leukocytes in these recordings is dynamic, changing in character, location and time. Consequently, conventional image processing techniques are incapable of extracting the spatiotemporal profiles and bulk statistics of numerous leukocytes responding to a disease state, necessitating labor-intensive manual processing, a significant limitation of this approach. Here we describe a machine learning pipeline that uses a semantic segmentation algorithm and classification script that, in combination, is capable of automated and label-free leukocyte trafficking analysis in a coculture mimetic. The developed computational toolset has demonstrable parity with manually tabulated datasets when characterizing leukocyte spatiotemporal behavior, is computationally efficient and capable of managing large imaging datasets in a semi-automated manner.
We propose two automatic parameter tuning methods for Plug-and-Play (PnP) algorithms that use CNN denoisers. We focus on linear inverse problems and propose an iterative algorithm to calculate generalized cross-validation (GCV) and Stein’s unbiased risk estimator (SURE) functions for a half-quadratic splitting-based PnP (PnP-HQS) algorithm that uses a state-of- the-art CNN denoiser. The proposed methods leverage forward mode automatic differentiation to calculate the GCV and SURE functions and tune the parameters of a PnP-HQS algorithm automatically by minimizing the GCV and SURE functions using grid search. Because linear inverse problems appear frequently in computational imaging, the proposed methods can be applied in various domains. Furthermore, because the proposed methods rely on GCV and SURE functions, they do not require access to the ground truth image and do not require collecting an additional training dataset, which is highly desirable for imaging applications for which acquiring data is costly and time-consuming. We evaluate the performance of the proposed methods on deblurring and MRI experiments and show that the GCV-based proposed method achieves comparable performance to that of the oracle tuning method that adjusts the parameters by maximizing the structural similarity index between the ground truth image and the output of the PnP algorithm. We also show that the SURE-based proposed method often leads to worse performance compared to the GCV-based proposed method.
Gwen Littlewort合作论文数Machine Perception Laboratory;;Institute for Neural Computation9