
Real-world optimization problems often involve complex physical or industrial process constraints that must be met to ensure feasible solutions. Constrained Bayesian Optimization (CBO) provides an efficient framework for such optimization tasks that relies on uncertainty-aware surrogate models to approximate both the objective function and the constraints. Existing approaches for modeling the constraints often assume homoscedastic Gaussian noise and utilize Gaussian processes (GPs), which can lead to miscalibrated uncertainty estimates, particularly in the presence of model mismatch. To address these challenges, we propose a novel framework that improves both the calibration and adaptivity of constraint modeling. First, we use conformal prediction (CP) to construct prediction intervals with guaranteed calibration, independent of the noise distribution. Next, we improve local adaptivity by modeling residuals with a nonparametric kernel density estimator, enabling the intervals to adjust dynamically to heteroscedastic and non-Gaussian noise. Finally, we use a distance-based uncertainty heuristic to detect distribution shifts, improving robustness in regions with sparse or out-of-distribution samples. Our framework is model-agnostic and can be applied to any prediction model. We validate our method on a synthetic CBO task and on real-world data by modeling constraints for carbite-free bainitic steel optimization [1]. Our results show that standard GP models often produce miscalibrated uncertainty estimates, while our method yields improved calibration, and reduced prediction intervals.
Phase retrieval is a fundamental inverse problem that arises in many scientific and engineering disciplines, where the goal is to reconstruct a signal from intensity-only measurements. In this work, we develop prNet, a novel phase retrieval framework that stochastically refines initial reconstructions with learned denoising and model-based updates. Our framework combines Langevin dynamics-based posterior sampling, adaptive noise schedule learning, warm-start initialization from classical solvers, and a progressive training strategy inspired by algorithm unrolling. By considering the perception-distortion tradeoff, our method also mitigates the over-smoothing effects commonly observed in prior approaches and enables reconstructions with fine details while minimizing artifacts. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in both efficiency and reconstruction quality.
Over the last decade, deep learning models have been widely used for automatic feature extraction and classification in various Brain-Computer Interface (BCI) tasks. However, their performance and generalization capabilities are often not adequately assessed, as these models are frequently trained and tested under flawed setups and / or influenced by spurious correlations. Recently, these limitations have also been observed in the training and evaluation of Large Brainwave Foundation Models (LBMs). In this work, we employ causal reasoning and careful consideration for task-discriminative artifacts in various EEG datasets covering diverse BCI paradigms and propose a benchmarking protocol to properly evaluate the decoding performance and generalization capabilities of LBMs. Utilising a subject-independent cross-validation approach for each curated benchmark dataset, we showcase that LBMs achieve marginal performance gains over conventional deep learning baselines.
Exploiting large-scale pre-trained models has proven effective in many domains, enabling limited data downstream tasks to be solved more effectively by leveraging upstream representations. However, transfer learning remains underexplored in passive sonar, where models are often trained from scratch using vision-oriented architectures on engineered features. We present the first systematic benchmark comparing transformer and CNN models across multiple passive sonar classification datasets. We evaluate six models across four sonar datasets, employing various fine-tuning strategies, including parameter-efficient Low-Rank Adaptation. Our findings demonstrate that full fine-tuning with audio pre-trained transformers performs most consistently. Fine-tuning on just 10 % of the data can outperform training from scratch on the full dataset. Importantly, we introduce standardised training and evaluation protocols across multiple datasets, making conclusions more robust and replicable compared to prior work. This will inform both research and practice in passive sonar recognition.
Image restoration seeks to reconstruct high-quality images from low-quality observations, with tasks such as superresolution, deblurring, and inpainting. Recently, diffusion models have emerged as powerful tools for image restoration, generating diverse and high-quality samples. However, their high computational demands, including multiple denoising steps, hinder deployment on resource-constrained devices like smartphones. While efforts to optimize sampling trajectories have been made, the large parameters of DMs remain a challenge. Model quantization, a common solution for reducing latency and memory usage, faces channel-wise variance when applied to DMs. Thus, an incoherence processing is needed for the outlier suppression to produce approximately Gaussian distribution. Futhermore, the uniform bitwidth allocation also turns into the bottleneck of quantization due to the ignorance of layer-wise sensitivity. To address the two main challenges, we propose MPRDiff, a mixed-precision quantization method that uses a computationally efficient randomized Hadamard transform technique to eliminate channel-wise incoherency and enable W3.8A4 layer-wise quantization of the DMs. Additionally, we use integer linear programming for optimal mixed bitwidth allocation and incorporate sampling step correction techniques. MPRDiff is evaluated on CelebA and ImageNet, demonstrating a minimal increase in FID while significantly reducing model size.
Computer simulations across various scientific domains often necessitate efficient calibration methods to ensure accurate representation of real-world phenomena. Traditional approaches rely heavily on expert intervention and high computational resources, which can be limiting. In this paper, we propose a calibration framework based on the model-bridge (MB) paradigm, which addresses these problems. Our MB method has two components: a surrogate and a bridge. Our surrogate is a Hankel correlation matrix, which is built by treating simulation systems as signals, representing their frequency-amplitude information compactly. This surrogate allows us to flexibly deal with different time lengths and simulation behaviors, and provides a straightforward geometric interpretation enabled by the underlying matrix manifold geometry. For the bridge model, we propose a kernel ridge regression approach and multiple kernel functions based on matrix manifold metrics. Through experimental validation on synthetic signal and fluid dynamics simulations, we show that the proposed Hankel framework outperforms existing calibration methods.
Most existing works on Continual Learning (CL) tend to assume exclusivity or dissimilarity among learning tasks, consequently these methods usually require constantly accumulating task-specific knowledge in memory for each task. This results in the eventual prohibitive expansion of the knowledge repository if we consider learning from a long sequence of tasks. In this work, we introduce a paradigm where the continual learner gets reoccurring tasks. We propose a framework that uses a task identity detection function that does not require additional learning, with which we analyze whether the current task is reoccurring to a specific task in the past. We can then reuse previous knowledge to slow down parameter expansion, ensuring that the system expands the knowledge repository sublinearly relative to the number of learned tasks. Our experiments show that the proposed framework performs competitively on widely used benchmarks such as CIFAR10, CIFAR100, EMNIST, and TinyImageNet, from which we create sequences of 10 to 100 tasks.
Multispectral object detection based on visible and infrared images leverages the complementary advantages of both modalities for robust object detection across both day and night conditions. However, existing methods only employ single-stage fusion paradigm, which lack hierarchical information interaction for deep feature fusion. To address this, we propose a novel multispectral coarse-to-fine fusion approach for object detection (MCFF-Det). MCFF-Det features dual fusion stages developed based on YOLOv8 and transformer, performing coarse and fine fusion, respectively. Specifically, the fine fusion stage introduces a transformer-based fusion refinement (TFR) module to refine the coarse-fused features on a patchwise basis and reintroduce the original modalities to complete the cross-modal fusion. Extensive experiments conducted on the LLVIP dataset demonstrate that MCFF-Det outperforms state-of-the-art methods quantitatively and qualitatively in terms of pedestrian detection across diverse scenarios.
Accurate prediction of pathloss radio maps is essential for the design and optimization of next-generation indoor wireless communication systems. Incorporating sparse pathloss measurements as auxiliary information has demonstrated significant potential in improving prediction accuracy. In this paper, we propose a novel sampling-assisted indoor pathloss prediction method (SAIPP-Net). First, we design a UNetbased neural network with variable-channel inputs to adapt to different levels of sampling availability. Second, we introduce a sampling-aware training strategy that employs tailored training schemes for low and high sampling rates, respectively. Finally, we develop a prioritized hybrid sampling strategy that jointly considers the transmitter distance and signal gradient to guide the selection of informative sampling locations. SAIPP-Net was evaluated in the context of MLSP 2025 The Sampling-Assisted Pathloss Radio Map Prediction Data Competition, achieving a weighted root mean squared error of 4.67 dB on the test set and securing 1st place in the competition.
Historical manuscripts suffer from various forms of degradation such as ink bleed, fading, and physical damage, which impairs their readability and usability. This study proposes a novel image restoration framework using Diffusion Probabilistic Models (DPMs) enhanced with OCR-guided finetuning. By integrating a two-stage architecture-comprising a NAFNet-based Initial Predictor and a Diffusion-Based De-noiser-this system restores both low- and high-frequency document details. Evaluations on benchmark datasets demonstrate that the proposed method outperforms traditional techniques in terms of PSNR, SSIM, and OCR accuracy, making it a scalable and robust solution for cultural heritage preservation.
In this paper, we introduce IRM-Net, a novel variant of the U-Net architecture designed for indoor radio map estimation. IRM-Net incorporates two principal improvements over the standard U-Net. First, we replace conventional convolutional layers with a cascaded combination of a Detail Enhancement Block (DEB) and a Detail Enhancement Attention Block (DEAB), which enhances the model's ability to capture finegrained features. Second, we implement dense connections in both the encoder and decoder, facilitating multi-level semantic interactions that mitigate information loss more effectively than traditional serial connections. IRM-Net was trained and evaluated on the benchmark dataset provided by the Sampling-Assisted Pathloss Radio Map Prediction Data Competition. Experimental results demonstrate that our approach can reliably predict path loss distributions in previously unseen indoor environments.
Due to the widespread use of drones in an urban environment, drones present an increased risk to the safety of urban life. Reliable detection of drones becomes crucial for countering the hazard introduced by drones. However, drones are difficult to detect because of their size and customization. This paper introduces DDL, a dataset aimed at drone sound detection, classification, and localization via a specially constructed set of microphones. As a baseline, we propose a deep uncertainty-aware framework implementing Conformer for joint drone classification and localization. We employ heteroscedastic loss functions that jointly estimate means and variances for spatial localization to model prediction uncertainty. Experiments on the DDL dataset demonstrate a classification accuracy of 99.9 % and a Euclidean distance mean absolute error (MAE) of approximately 16 meters. The uncertainty estimates are well-calibrated, with coverage closely matching the expected confidence intervals ($68 \%, 95 \%$, and 99.7 %) as defined by the empirical rule, suggesting DDL as a benchmark dataset for audio-based drone localization.
The classical Multi-armed Bandit (MAB) framework relies on immediate access to reward feedback, yet in many practical settings rewards are inaccessible or severely delayed. To address this gap, inverse MAB seeks to recover the latent reward function by observing the actions of one or more demonstrators. In this work, we study the problem of learning from multiple noisy demonstrations in a stochastic MAB environment, in the absence of reward information. Four algorithms for learning from noisily optimal demonstrators, that is, demonstrators who choose the optimal arm with fixed probability and a suboptimal one otherwise, are considered; three adapted from existing literature and one newly proposed. Their performance is evaluated both theoretically, through regret bounds, and empirically, via numerical tests.
We propose a new online frequency estimation algorithm in a local differential privacy (LDP) setting with adaptive randomized response mechanisms. The method is based on two components: (i) Online expectation-maximization (EM) for parameter estimation; (ii) using an adaptive randomized response mechanism to provide LDP. The purpose of the adaptation is to increase the utility of the randomized responses. We present numerical results that show the benefit of adapting. The paper also includes several extensions of the current work that are to be explored.
Manual segmentation of liver vessels is a challenging and time-consuming task for radiologists, prone to variability and misclassification, which can affect treatment planning for conditions like hepatocellular carcinoma. To address this, the study utilized the nnU-Net framework as a baseline and introduced several enhancements, participating in the VEELA 2025 challenge. The VEELA 2025 challenge comprises 20 train and 20 test hepatic and portal vein labeled computed tomography angiography (CTA) images for segmentation and classification. The methodology involved training on a combination of CTA and contrast-enhanced CT images from datasets including the Medical Segmentation Decathlon, LIRCAD, and the VEELA 2025 challenge data, due to the limited availability of labeled CTA images. Five different models were tested, primarily nnU-Net based, each with a different loss function incorporating elements like Dice Similarity Coefficient (DSC), cross-entropy loss, centerline Dice loss (clDice) to promote topological continuity, and edge map Dice loss to improve boundary precision. The final model, which employed a weighted loss strategy prioritizing pixel-wise classification while reinforcing spatial consistency and structure preservation, achieved the top ranking in the VEELA 2025 challenge. Key contributions include using centerline regression cross-entropy loss for minor vessel classification, weighted losses for balanced optimization, and leveraging contrast-enhanced CT data to overcome CTA data scarcity.
Range-based anomaly detection in multivariate time series is critical for real-world applications such as industrial monitoring, where anomalies often span intervals and exhibit distinct types. We propose a transformer-based framework designed for highly imbalanced, multiclass anomaly detection at the range level. To improve prediction stability and coherence, we introduce two post-inference strategies: (i) a majority voting mechanism to consolidate overlapping multi-step predictions, and (ii) a transition-aware masking scheme that applies domain-specific rules to constrain class transitions. We evaluate our method on the Exathlon benchmark, extending its original binary protocol to support multiclass assessment. Under binary evaluation, our approach outperforms standard forecasting- and reconstruction-based baselines, achieving up to a 24% improvement in F1 score.
Developing a highly accurate automatic license plate recognition system (ALPR) is challenging due to environmental factors such as lighting, rain, and dust. Additional difficulties include high vehicle speeds, varying camera angles, and low-quality or low-resolution images. ALPR is vital in traffic control, parking, vehicle tracking, toll collection, and law enforcement applications. This paper proposes a deep learning strategy using YOLOv8 for license plate detection and recognition tasks. This method seeks to enhance the performance of the model using datasets from Ontario, Quebec, California, and New York State. It achieved an impressive recall rate of 94
The prediction of spectrum maps has received significant attention in recent years as a result of its critical role in resource management and signal detection. However, contemporary spectrum maps often exhibit rapid temporal variations, and historical spectrum data are frequently incomplete or contaminated by anomalies. Existing research on spectrum map prediction remains limited in addressing missing data processing capabilities. To address this challenge, we propose DSwinLSTM-I, an encoder-decoder framework that integrates imputation and prediction of spectrum maps into a unified process, thus streamlining the traditional two-step im-putation-prediction workflow. DSwinLSTM-I incorporates a dedicated imputation unit, allowing the model to infer missing values while simultaneously predicting future spectrum maps. Experimental results demonstrate that the proposed DSwinLSTM-I achieves superior performance in spectrum map prediction with respect to both accuracy and robustness.
Empathetic dialogue generation aims to produce responses that demonstrate a thorough grasp of both conversational semantics and emotion features. Unlike conventional methods that prioritize the emotion prediction through topic modeling or enhance the dialogue generation via reinforcement learning, these methods often fail to address essential features or components of the emotional empathy. This paper copes with this issue by using the disentangled representation learning to consolidate the emotion components of a dialogue. In particular, this study presents a new contrastive learning to estimate the valence-arousal-dominance to refine emotion attributes for empathetic generation. These attributes are integrated with soft prompts in large language models (LLMs), ensuring the responses which are contextually relevant and emotionally appropriate. Experimental results on the EmpatheticDialogues dataset show that the proposed method enhances the emotion sensitivity in text generation by using existing LLMs. Significant gains are seen in both automatic and human assessments, reflecting superiority in overall performance for empathetic response.
Spike-based temporal messaging enables SNNs to efficiently process both purely temporal and spatio-temporal time-series or event-driven data. Combining SNNs with Gated Recurrent Units (GRUs), a variant of recurrent neural networks, gives rise to a robust framework for sequential data processing; however, traditional RNNs often lose local details when handling long sequences. Previous approaches, such as SpikGRU, fail to capture fine-grained local dependencies in event-based spatio-temporal data. In this paper, we introduce the Convolutional Spiking GRU (CS-GRU) cell, which leverages convolutional operations to preserve local structure and dependencies while integrating the temporal precision of spiking neurons with the efficient gating mechanisms of GRUs. This versatile architecture excels on both temporal datasets (NTIDIGITS, SHD) and spatio-temporal benchmarks (MNIST, DVSGesture, CIFAR10DVS). Our experiments show that CS-GRU outperforms state-of-the-art GRU variants by an average of 4.35