We present an efficient method for image segmentation in the presence of strong inhomogeneities. The approach can be interpreted as a two-level clustering procedure: pixels are first grouped into superpixels via a linear least-squares assignment problem, which can be viewed as a special case of a discrete optimal transport (OT) problem, and these superpixels are subsequently greedily merged into object-level segments using the squared 2-Wasserstein distance between their empirical distributions. In contrast to conventional superpixel merging strategies based on mean-color distances, our framework employs a distributional OT distance, yielding a mathematically unified formulation across both clustering levels. Numerical experiments demonstrate that this perspective leads to improved segmentation accuracy on challenging images while retaining high computational efficiency.
We introduce a novel formulation for curvature regularization by penalizing normal curvatures from multiple directions. This total normal curvature regularization is capable of producing solutions with sharp edges and precise isotropic properties. To tackle the resulting high-order nonlinear optimization problem, we reformulate it as the task of finding the steady-state solution of a time-dependent partial differential equation (PDE) system. Time discretization is achieved through operator splitting, where each subproblem at the fractional steps either has a closed-form solution or can be efficiently solved using advanced algorithms. Our method circumvents the need for complex parameter tuning and demonstrates robustness to parameter choices. The efficiency and effectiveness of our approach have been rigorously validated in the context of surface and image smoothing problems.
. In this paper, we employ a Schauder-type estimate method, as developed in the study by Chen et al. [Well-posedness for local and nonlocal quasilinear evolution equations in fluids and geometry, arXiv:2407.05313, 2024] to establish critical well-posedness result for the Fractional Fokker- Planck Equation (FFPE). This equation serves as a fundamental model in kinetic theory and can be regarded as a semi-linear analogue of the non-cutoff Boltzmann equation. We demonstrate that the techniques introduced in this study are not only effective for the FFPE but also hold promise for broader applications, particularly in addressing the non-cutoff Boltzmann equation and the Landau equation. Our results contribute to a deeper understanding of the analytical framework required for these complex kinetic models.
Data scarcity is a fundamental barrier in Electrical Impedance Tomography (EIT), as undersampled Dirichlet-to-Neumann (DtN) measurements can substantially degrade conductivity reconstructions. We address this bottleneck by completing partially observed DtN measurements using a diffusion based generative model. Specifically, we train a conditional diffusion model to learn the distribution of DtN data and to infer full measurement vectors given partial observations. Our approach supports flexible source receiver configurations and can be used as a plug in preprocessing step with off the shelf EIT solvers. Under mild assumptions on the polygon conductivity class, we derive nonasymptotic end to end bounds on the distributional discrepancy between the completed and ground truth DtN measurements. In numerical experiments, we couple the proposed diffusion completion procedure with a deep learning based inverse solver and compare its performance against the same solver with full measurement data. The results show that diffusion completion enables reconstructions comparable to the full data baseline while using only 1
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak long-range speaker memory, and the inability to output timestamps. To address these limitations, we present MOSS Transcribe Diarize, a unified multimodal large language model that jointly performs Speaker-Attributed, Time-Stamped Transcription in an end-to-end paradigm. Trained on extensive real wild data and equipped with a 128k context window for up to 90-minute inputs, MOSS Transcribe Diarize scales well and generalizes robustly. Across comprehensive evaluations, it outperforms state-of-the-art commercial systems on multiple public and in-house benchmarks.
High-order regularization models, especially those involving mean curvature, have gained significant prominence in variational image restoration due to their ability to mitigate the staircase effect and represent smooth intensity variations more accurately than first-order Total Variation models. Despite these advantages, the intrinsic high nonlinearity and non-smoothness of the mean curvature functional present formidable challenges for the development of efficient and stable numerical solvers. In this paper, we propose a novel computational framework for a class of mean-curvature-type models by integrating the Legendre–Fenchel transform within a primal–dual augmented Lagrangian formulation. By leveraging the duality properties of the Legendre–Fenchel transform, we reformulate the original nonlinear problem into a structured saddle-point problem. A distinguished feature of the proposed framework is its streamlined structure, which requires fewer auxiliary variables than traditional variable-splitting-based augmented Lagrangian methods. This reduction leads to enhanced computational efficiency and lower memory requirements. Numerical experiments on various image restoration tasks demonstrate that the proposed algorithm yields competitive performance in terms of both convergence speed and restoration quality, particularly in preserving geometric fidelity.
Operator learning for partial differential equations (PDEs) on arbitrary geometries builds fast neural surrogates for large-scale simulation. Although recent geometry-adaptive neural operators have made substantial progress, they are mainly designed for forward problems in which inputs and outputs share the same spatial domain. This limits their applicability for boundary value problems (BVPs) and inverse problems, where inputs and outputs may live on different domains. We introduce the Geometry-Adaptive Integral Autoencoder (GAIA), an operator learning model that encodes the domain boundary and the interior field distribution into geometry tokens, and conditions integral transform layers on these tokens via cross-attention, allowing the kernel to adapt locally to geometric features. This yields a single architecture for forward (including BVPs) and inverse problems on arbitrary domains in one pass, without retraining, iterative optimization, or graph construction. We evaluate GAIA on seven 2D and 3D benchmarks, four of which are new or substantially extended benchmarks for inverse problems and BVP: electrical impedance tomography, optical tomography, 3D Darcy flow on varying geometries, and a modified setting of Poisson BVP on mechanical components benchmark (MCB). GAIA sets new state-of-the-art results on every inverse and BVP task, reducing median relative L^2 error by 64
Electrical Impedance Tomography (EIT) is a non-invasive medical imaging method that reconstructs electrical conductivity mediums from boundary voltage-current measurements, but its severe ill-posedness renders direct operator learning with neural networks unreliable. We propose the neural correction operator framework, which learns the inverse map as a composition of two operators: a reconstruction operator using L-BFGS optimization with limited iterations to obtain an initial estimate from measurement data and a correction operator implemented with deep learning models to reconstruct the true media from this initial guess. We explore convolutional neural network architectures and conditional diffusion models as alternative choices for the correction operator. We evaluate the neural correction operator by comparing with L-BFGS methods as well as neural operators and conditional diffusion models that directly learn the inverse map over several benchmark datasets. Our numerical experiments demonstrate that our approach achieves significantly better reconstruction quality compared to both iterative methods and direct neural operator learning methods with the same architecture. The proposed framework also exhibits robustness to measurement noise while achieving substantial computational speedup compared to conventional methods. The neural correction operator provides a general paradigm for approaching neural operator learning in severely ill-posed inverse problems.
Discrete representation learning has shown promising results across various domains, including generation and understanding in image, speech and language. Inspired by these advances, we propose MuseTok, a tokenization method for symbolic music, and investigate its effectiveness in both music generation and understanding tasks. MuseTok employs the residual vector quantized-variational autoencoder (RQ-VAE) on bar-wise music segments within a Transformer-based encoder-decoder framework, producing music codes that achieve high-fidelity music reconstruction and accurate understanding of music theory. For comprehensive evaluation, we apply MuseTok to music generation and semantic understanding tasks, including melody extraction, chord recognition, and emotion recognition. Models incorporating MuseTok outperform previous representation learning baselines in semantic understanding while maintaining comparable performance in content generation. Furthermore, qualitative analyses on MuseTok codes, using ground-truth categories and synthetic datasets, reveal that MuseTok effectively captures underlying musical concepts from large music collections.
Image registration is crucial for many medical imaging applications, including longitudinal monitoring and multimodal information fusion. A key challenge is to achieve accurate alignment while strictly preserving topology and invertibility. To address the limitations of traditional penalty-based regularization, which may still permit local folding, this study proposes DTC-Reg, a dynamically learned registration framework that more explicitly enforces diffeomorphic deformation. The framework integrates a homotopy-based control-increment formulation with explicit multiscale geometric constraints. Two parameter-sharing U-Nets first extract multiscale feature pyramids from the input images, after which a symmetric registration module with a sequential temporal cascade network progressively refines the forward and inverse multiscale deformation fields. To further enhance diffeomorphic consistency, this study introduces a Multiscale Folding-aware Deformation Correction (MFDC) module that explicitly detects and geometrically rectifies folding points in the predicted deformation fields. Beyond its integration within DTC-Reg, MFDC can also be readily incorporated into several state-of-the-art registration networks, significantly reducing folding and improving deformation regularity. Extensive experiments on three 3D brain MRI registration tasks demonstrate that the proposed method consistently achieves superior performance over existing approaches in both quantitative and qualitative evaluations.
Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing deep constrained clustering (DCC) methods mainly target hard, expert-agnostic constraints, treating soft labels mostly numerically rather than semantically. We formalize this setting as uncertainty-aware probabilistic constrained clustering (UPCC), defining a canonical aleatoric target through a heterogeneous observation process and analyzing its conditional identifiability. We introduce ProbPair, an angular pairwise objective for probabilistic relations, and build ECI-PP, an estimator–corrector–integrator framework that refines imperfect supervision via belief estimation, correction, and reliability-aware integration. Across challenging probabilistic supervision settings, experiments on diverse benchmarks show that ECI-PP outperforms state-of-the-art DCC methods and remains robust with a shared default configuration.
Dual-energy CT (DECT) exploits attenuation differences across different X-ray spectra to provide richer material information and has been widely used in medical imaging. While sparse-view acquisition can lower radiation exposure, it makes DECT material decomposition even more challenging, as the problem is nonlinear and ill-posed. Existing deep unrolling approaches generally do not explicitly incorporate the Jacobian operator induced by the nonlinear forward model, and their sparsity priors are still mainly built on conventional convolutions, which are insufficient for modeling global structural information. This study addresses the challenge of DECT multi-material decomposition in sparse-view settings by representing it as a sparse-regularized nonlinear least-squares problem. To solve it, we propose an iterative dual-domain refinement network (DECT-DRNet). In each iteration, the filtered back-projection (FBP)-based Jacobian approximation module is used first to generate an intermediate material decomposition result. Here, we characterize the forward process of material decomposition using a nonlinear operator, and then construct a theoretically grounded learnable approximation of the adjoint Jacobian operator by integrating the FBP algorithm with a U-Net into the backward process. In addition, to address the limitation of existing deep learning-based decomposition methods in globally suppressing noise and artifacts, we introduce a learnable sparse dual domain regularization term that incorporates Fourier convolutional residual blocks. This refinement block combines geometric feature extraction in the image domain with noise suppression in the frequency domain, allowing the model to capture both global and local features while maintaining structural details. DECT-DRNet demonstrates its ability to achieve more accurate material decomposition under sparse-view conditions.
Deep-learning techniques have demonstrated significant potential in low-dose computed tomography (LDCT) reconstruction. Nevertheless, supervised methods are limited by the scarcity of labeled data in clinical scenarios, while CNN-based unsupervised denoising methods often result in excessive smoothing of reconstructed images. Although normalizing flows (NFs) based methods have shown promise in generating detail-rich images and avoiding over-smoothing, they face two key challenges: (1) Existing two-way transformation strategies between noisy images and latent variables, despite leveraging the regularization and generation capabilities of NFs, can lead to detail loss and secondary artifacts; and (2) Training NFs on high-resolution CT images is computationally intensive. While conditional normalizing flows (CNFs) can mitigate computational costs by learning conditional probabilities, current methods rely on labeled data for conditionalization, leaving unsupervised CNF-based LDCT reconstruction an unresolved challenge. To address these issues, we propose a novel unsupervised LDCT iterative reconstruction algorithm based on CNFs. Our approach implements a strict one-way transformation during alternating optimization in the dual spaces, effectively preventing detail loss and secondary artifacts. Additionally, we propose an unsupervised conditionalization strategy, enabling efficient training of CNFs on high-resolution CT images and achieving fast, high-quality unsupervised reconstruction. Experimental results across multiple datasets demonstrate that the proposed method outperforms several state-of-the-art unsupervised methods and even rivals some supervised approaches.
Existing level set models employ regularization based only on gradient information, 1D curvature or 2D curvature. For 3D image segmentation, however, an appropriate curvature-based regularization should involve a well-defined 3D curvature energy. This is the first paper to introduce a regularization energy that incorporates 3D scalar curvature for 3D image segmentation, inspired by the Einstein-Hilbert functional. To derive its Euler-Lagrange equation, we employ a two-step gradient descent strategy, alternately updating the level set function and its gradient. The paper also establishes the existence and uniqueness of the viscosity solution for the proposed model. Experimental results demonstrate that our proposed model outperforms other state-of-the-art models in 3D image segmentation.
Image registration is a key technique in image processing and analysis. Due to its high complexity, the traditional registration frameworks often fail to meet real-time demands in practice. To address the real-time demand, several deep learning networks for registration have been proposed, including the supervised and the unsupervised networks. Unsupervised networks rely on large amounts of training data to minimize specific loss functions, but the lack of physical information constraints results in the lower accuracy compared with the supervised networks. However, the supervised networks in medical image registration face two major challenges: physical mesh folding and the scarcity of labeled training data. To address these two challenges, we propose a novel few-shot learning framework for image registration. The framework contains two parts: random diffeomorphism generator (RDG) and a supervised few-shot learning network for image registration. By randomly generating a complex vector field, the RDG produces a series of diffeomorphism. With the help of diffeomorphism generated by RDG, one can use only a few image data (theoretically, one image data is enough) to generate a series of labels for training the supervised few-shot learning network. Concerning the elimination of the physical mesh folding phenomenon, in the proposed network, the loss function is only required to ensure the smoothness of deformation (no other control for mesh folding elimination is necessary). The experimental results indicate that the proposed method demonstrates superior performance in eliminating physical mesh folding when compared to other existing learning-based methods. Our code is available at this link https://github.com/weijunping111/RDG-TMI.git.
Biological vision exhibits exceptional contour perception capabilities. In view of this, research on contour detection guided by biological vision is gradually gaining attention. Inspired by the transmission and processing of visual signals in the primary visual pathway, this study proposes a lightweight contour detection network called the Primary-Visual-Pathway UNet (PVP-UNet), comprising an encoder and a decoder. Primarily, inspiration from the binocular vision mechanism, we transmit the original image through dual deformable convolution modules that simulate the receptive fields of the left and right retinas in the encoder. Utilizing the characteristic of the optic chiasm, the output features of the retinal layers are split, swapped, and merged for further processing. Subsequently, dilated convolution modules and normal convolution modules are involved to simulate the magnocellular (M) and parvocellular (P) pathways of the lateral geniculate nucleus (LGN), respectively. An inhibition module was designed based on the suppression mechanism of classical/non-classical receptive fields (CRF/NCRF) in the primary visual cortex (V1). Additionally, the interconnection pattern of the inhibition modules is deployed by leveraging the aggregation characteristics of simple cells to complex cells in the V1 layer. Inspired by the feedback mechanism of visual information, the feature fusion module is introduced in the decoder to integrate features from different encoding layers in the reverse direction of signal transmission along the primary visual pathway. Our experiments show that the respective Optimal Dataset Scale (ODS) scores achieve 0.811, 0.756, and 0.896 on BSDS500, NYUD, and BIPED datasets. The experimental results demonstrate that the proposed network effectively suppresses background noise and highlights primary contours, exhibiting excellent detection performance on the tested datasets. The code is available at https://github.com/k3chencoco/PVP-UNet.
Pan-sharpening refers to fusing remote sensing multispectral (MS) and panchromatic (PAN) images to generate high-resolution multispectral (HR-MS) images. Recent advancements in deep learning-based pan-sharpening techniques have shown promising results. However, they face the following two issues. On one hand, there is a modality gap between MS and PAN images. Directly fusing them can lead to spectral and spatial distortions. On the other hand, the fusion process is prone to information loss, which can lead to image blurriness. To tackle these issues, we develop a Transformer-based model: FAFormer, which incorporates frequency analysis and focuses on the correlation and specificity of the PAN and MS images. Focusing on correlation can reduce the spectral and spatial distortions while focusing on specificity can reflect the specific information from MS and PAN images in the fusion result. We utilize the Discrete Wavelet Transform (DWT) to obtain the correlate and specific features. We introduce bijective functions based on the Transformer to design an Integrated Attention Block (IAB). As a critical component of the model, it effectively utilizes the correlation and specificity of the two images. In designing the model’s overall framework, we employ a Correlative Feature Attention Module (CFAM) to leverage the correlation between MS and PAN. We utilize a Specific Feature Attention Module (SFAM) to integrate specific information into fused features gradually. Experimental results show that our method improves pan-sharpening performance and has practical value. Codes are available at https://github.com/Xidian-AIGroup190726/FAFormer.
Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks.
Traditional variational models often fail to segment images in the presence of inhomogeneity or weak boundaries, partly due to their reliance on unreliable region metrics that quantify inhomogeneity based on a single mean value or a smoothed image serving as a mean function. The former, such as variance-based methods, are highly sensitive to image inhomogeneity, whereas the latter, such as local convolution-based approaches, lack a global receptive field. To address these issues, we employ an optimal transport-based data fidelity term in our segmentation objective functional. This term accounts for global differences between regions, resolving problems arising from local convolutions. It can also adaptively seek an optimized match between two probability density functions, proving more robust than relying solely on their mean values. Our proposed functional is minimized by gradually performing region merging. Experimental results demonstrate that our model outperforms state-of-the-art variational and deep learning models.
Deep neural networks (DNNs) have recently emerged as effective tools for approximating solution operators of partial differential equations (PDEs) including evolutionary problems. Classical numerical solvers for such PDEs often face challenges of balancing stability constraints and the high computational cost of iterative solvers. In contrast, DNNs offer a data-driven alternative through direct learning of time-stepping operators to achieve this balancing goal. In this work, we provide a rigorous theoretical framework for analyzing the approximation of these operators using feedforward neural networks (FNNs). We derive explicit error estimates that characterize the dependence of the approximation error on the network architecture – namely its width and depth – as well as the number of training samples. Furthermore, we establish Lipschitz continuity properties of time-stepping operators associated with classical numerical schemes and identify low-complexity structures inherent in these operators for several classes of PDEs, including reaction-diffusion equations, parabolic equations with external forcing, and scalar conservation laws. Leveraging these structural insights, we obtain generalization bounds that demonstrate efficient learnability without incurring the curse of dimensionality. Finally, we extend our analysis from single-input operator learning to a general multi-input setting, thereby broadening the applicability of our results.