
When a system operates under impulsive disturbances and conventional state estimation methods become ineffective, robust peak error reduction is required. In this paper, the error-to-error transfer function approach is used to compute the bias correction gain $\mathbf {K}$ for a robust recursive $\mathcal {L}_{2}$-to-$\mathcal {L}_\infty$ filter applied to discrete-time state-space models based on the backward Euler method. The disturbance is treated as a Gauss-Markov sequence, and the gain $\mathbf {K}$ is computed using the energy-to-peak lemma, a newly formulated theorem, and a linear matrix inequality. The performances of the $\mathcal {L}_{2}$-to-$\mathcal {L}_\infty$, $H_\infty$, unbiased finite impulse response (UFIR), and Kalman filters are compared numerically based on the quasi periodic harmonic model in terms of mean square error, robustness, and estimation quality. An experimental verification is provided for air quality monitoring in a significantly polluted city area. It is shown that the gain $\mathbf {K}$ of the $\mathcal {L}_{2}$-to-$\mathcal {L}_\infty$ filter obeys the previously formulated rule of thumb, i.e. it ranges between the gains of the UFIR and Kalman filters. In both numerical and experimental examples, the $\mathcal {L}_{2}$-to-$\mathcal {L}_\infty$ filter exhibited the smallest peak errors.
Light Field (LF) cameras simultaneously capture both intensity values and directional information of light rays in a single exposure, providing a unique perspective for computational photography and 3D geometry perception. However, existing LF cameras are constrained by sensor resolution, limiting their ability to capture high spatial and angular resolutions simultaneously. To mitigate these issues, various learning-based methods have been proposed to increase the angular resolution of captured LFs, known as LF View Synthesis (LFVS). Many of these methods either neglect essential geometric cues or rely on neural networks with large receptive fields, which restrict their ability to accurately exploit LF structural characteristics. In response to these challenges, this paper introduces a dual representation-based LFVS method that employs deformable convolutional and Deep Residual Channel Attention (DRCA) networks. The proposed method includes two main modules: (i) Coarse Light Field View Synthesis (CLFVS) for initial LFVS, and (ii) Coarse-To-Fine Refinement (CFR) for final quality enhancement. The CLFVS module relies on deformable convolutions to adaptively extract LF features using two parallel networks: (i) Spatial Feature Extraction (SPFE) network using a depth-dependent LFVS approach, and (ii) Angular Feature Extraction (AFE) network using a non-depth-dependent LFVS approach. The CFR module refines the CLFVS output using a DRCA network, which employs dense residual connections between residual groups instead of conventional convolutional layers. The DRCA network employs Residual Channel Attention Blocks (RCABs) to model inter-channel dependencies, selectively enhancing meaningful features while suppressing irrelevant ones. The proposed method achieves state-of-the-art performance on synthetic and real-world LF benchmarks.
Display-recapture attacks pose a critical threat to the integrity and authenticity of digital document images, particularly by concealing tampering traces through rephotographing displayed content. Existing document presentation attack detection (DPAD) methods often struggle to distinguish forensic artifacts (e.g., chromatic distortions and moiré patterns) from the natural textures inherent in documents with complex backgrounds. To address this texture confusion challenge, we propose a dual-stream LC&DF framework that integrates Local Chromaticity (LC) features with a Masked Attention mechanism and a Discriminative Frequency (DF) branch enhanced via a Frequency-domain Moiré-Aware Adapter (FMAAda). This architecture jointly models local chromatic distortions and global frequency cues to robustly isolate recapture-induced artifacts from genuine document content. Extensive evaluations demonstrate the superiority of our method. Under the cross-dataset protocol, LC&DF achieves state-of-the-art performance. In challenging in-the-wild scenarios evaluated on the ROD_M&F, SRDID162, DLC2021, and KID34K benchmarks, our method consistently outperforms existing baselines, achieving an AUC of 0.9105 and reducing the Equal Error Rate by up to 15.72 percentage points on challenging document benchmarks. Furthermore, we conduct a zero-shot evaluation of Multimodal Large Language Models (MLLMs), revealing that they lack the sensitivity to subtle forensic artifacts. Visualizations of challenging samples further confirm that our claims on distinguishing the forensic artifacts from the document textures. The source code of this work will be available upon acceptance.
This paper describes the structure of a reconfigurable digital beamforming network, which offers significant improvements in implementation complexity and cost, while maintaining a high degree of flexibility. This architecture supports a multi-beam array antenna and allows complete flexibility in steering each individual beam produced by this antenna, achieving computational complexity comparable to that of the Fast Fourier Transform as the number of beams and antenna elements change. The proposed architecture is built upon the theory behind the Non-Uniform Fast Fourier Transform and is demonstrated for linear array geometries with regular or non-regular element positions. Fixed-point simulations, with a view to real-time hardware implementation, demonstrate effective reduction in the computational complexity of the reconfigurable digital beamforming network and its full flexibility.
When dealing with a time series dataset, a key challenge lies in grasping the dynamic signals. An integral task involves representing the time series through line segments, typically done as a preprocessing step to capture patterns and signals. This paper, however, centers on optimizing the time series segmentation through segmented linear regression models, specifically addressing large-scale datasets. We first formulate the optimization of Multi-Segment Linear Regression (MSLR), which aims to minimize the global square error of the segmented linear representations. The major contribution is that we introduce an optimal MSLR algorithm, denoted as OMSLR, featuring a 2-level dynamic programming structure. We also demonstrate the optimality and the low time complexity of O($kn^{2}$), where $k$ is the number of non-overlapping segments and $n$ is the length of the time series. The results demonstrate that the designed OMSLR matches the optimal global squared error of the classical $O(kn^{3})$ dynamic programming baseline and brute-force search to floating-point precision, while running roughly two orders of magnitude faster at $n = 10,000$ and producing lower reconstruction error at every $k$ than the bottom-up, top-down, Sim-Piece, and Mix-Piece baselines. Both effectiveness and efficiency are demonstrated on synthetic series, reference benchmarks, and the UCR time series archive.
This paper introduces the topology-independent distributed multichannel Wiener filter (TI-dMWF), a novel algorithm for distributed node-specific signal estimation in wireless acoustic sensor networks (WASNs) with unconstrained topologies. The TI-dMWF enables each node in the network to compute its centralized multichannel Wiener filter solution by exchanging only low-dimensional fused signals, without requiring iterative estimation, unlike state-of-the-art approaches such as the topology-independent distributed adaptive node-specific signal estimation (TI-DANSE) algorithm. The TI-dMWF is proven optimal when each source is observed by either all nodes or only one node. Theoretical analysis and numerical simulations confirm that it achieves centralized estimation performance in a single run. Its latency as a function of the pruned-tree depth and its computational complexity are also analyzed. Its robustness is assessed in reverberant-room simulations under estimated second-order statistics, various network topologies, and deviations from the assumed observability model.
Efficient and accurate direction of arrival (DoA) estimation is important in radar, wireless communication, and integrated sensing and communication systems. We propose a novel gridless DoA estimation algorithm within the variational Bayesian (VB) framework, tailored for antenna arrays equipped with few-bit analog-to-digital converters. The proposed approach facilitates the derivation of variational distributions for the spatial frequencies of the incoming paths and their associated path gains. Simulation results reveal that the proposed gridless-VB algorithm achieves superior normalized mean-squared error performance for DoA estimation compared to MUSIC, variational line spectral estimation (VALSE), and VALSE expectation propagation (VALSE-EP) algorithms, while maintaining comparable computational complexity. Additionally, the gridless-VB algorithm achieves the best MSE among the considered algorithms when evaluating the angle, gain, and phase estimation under different ranges of SNRs and bit depths. The impacts of key parameters such as the number of antennas, the number of paths, and the bit-depth of the quantizers are also comprehensively evaluated, offering valuable insights for algorithm design and optimization.
Attributing synthetic images to the source that generated them is a difficult problem, particularly in data-scarcity conditions requiring the adoption of few-shot or zero-shot learning strategies. In this paper, we tackle this problem by introducing a training-free attribution method based on image resynthesis. Our method works by first automatically generating a textual prompt which describes the target image, and then using it to resynthesize the image with each candidate generator. Attribution is obtained by matching the image to the model whose resynthesis is most similar to the original one within an appropriate feature space - chosen based on a preliminary experimental analysis - among several possibilities offered by modern deep learning image analysis tools. Experiments using state-of-the-art few-shot models and other baselines show that our resynthesis-based method outperforms existing techniques when only a limited number of images are available for training or fine-tuning. Particular attention is given to assessing the robustness of the proposed method against common image processing operators and its adversarial robustness. In particular, the experimental results show that our method is robust to post-processing and that its adversarial robustness is significantly higher than that of baselines.
Chirp waveforms are widely used in radar, sonar, acoustics, and other sensing systems due to their desirable pulse-compression behavior. However, standard chirp signals like linear frequency modulation (LFM) waveforms inevitably produce sidelobes that limit detection performance. In this paper, we propose a unified allpass-based chirp design (ABCD) framework that generates chirp-like waveforms with the perfect pulse compression (PPC) property, i.e., their autocorrelation reduces exactly to an impulse. The key idea is to design an allpass IIR system whose impulse response resembles a chirp. The allpass structure guarantees zero sidelobes, while its phase response closely approximates the ideal LFM. Band-limited chirp waveforms and long-duration, high-time-bandwidth signals can be obtained via interpolation and cascading, while preserving the perfect pulse-compression property. Compared with existing windowing or nonlinear FM techniques, the proposed method achieves exact sidelobe elimination together with flexible control of waveform characteristics, offering clear advantages for high-resolution and high-dynamic range sensing applications.
Fisheye cameras are used for coverage in transportation systems and vehicles, yet distortions pose challenges to object detection algorithms. In this paper, Radial-Aware LoRA (RA-LoRA) is proposed, a distortion-aware parameter-efficient adaptation method that modulates low-rank updates based on radial position to address the spatially varying nature of fisheye distortion. RA-LoRA fine-tuning of Florence-2-large with only 0.5% additional parameters achieves a 181% relative improvement in mAP@50 over zero-shot inference, outperforming fully trained CNN detectors in fisheye imagery. Detection performance is assessed in all five object classes of the FishEye8K dataset (bike, bus, car, pedestrian, and truck). The results show that Florence-2-large with RA-LoRA achieves a mean average precision (mAP) of 0.724, outperforming YOLO26x (mAP: 0.581) and YOLO26l (mAP: 0.605). Although the model requires approximately 4.5× more memory than YOLO26, the RA-LoRA adapters represent only 0.5% of the total parameters and can be merged into the base model at inference time with no additional overhead. This makes the proposed approach particularly suitable for server-side processing, offline traffic analysis, post-incident auditing, and high-accuracy surveillance applications where detection reliability outweighs the need for real-time embedded deployment.
Reliable instance segmentation of cells in microscopy images is essential for quantitative imaging analysis, yet estimating segmentation quality on unseen data, though fundamental, remains largely overlooked. In practice, performance assessment typically requires manual ground-truth annotations, which are costly and impractical at scale. We address this evaluation gap by introducing QANet, a neural framework for post-hoc quality assessment of instance segmentations. Unlike segmentation models, QANet does not predict masks; instead, it receives an image and a segmentation produced by any method and estimates a quantitative quality score that approximates standard evaluation metrics, without requiring ground-truth annotations at inference time. QANet is model-agnostic and formulated as a regression task over image-mask pairs. It is built on the RibCage architecture, which performs multi-scale comparison between image content and segmentation structure, enabling sensitivity to both global shape consistency and fine boundary errors. Training is performed using synthetically perturbed segmentations with known quality scores, enabling supervision across controlled error modes while eliminating the need for large annotated failure datasets. We evaluate QANet on 2D and 3D datasets from the Cell Segmentation Benchmark and demonstrate accurate prediction of both overlap-based measures and boundary-sensitive metrics across multiple segmentation methods.
Monocular depth estimation in complex, dynamic environments remains challenging due to rapid object motion, texture repetitions, occlusions, and strong geometric constraints inherent in structured scenes. To address these challenges, we propose DepthPlay-Agent (DEAN), an agentic multi-model fusion framework that builds upon SOTA depth estimator backbones such as UNet++, HybridDepth, DINOv3, and ZoeDepth under a learnable Mixture-of-Experts controller. This controller embeds scene context through transformers and uses task priors to dynamically assign pe-pixel fusion weights. A domain-aware refinement module further enforces geometric consistency using planar and semantic segmentation cues. Beyond static fusion, DEAN introduces an agentic inference layer that dynamically regulates expert contributions and refinement strategies, enabling adaptive and interpretable decision-making. Experiments across four benchmarks (SoccerNet-Depth, KITTI, MPI-Sintel, NYU Depth V2) demonstrate that DEAN consistently achieves SOTA performance, improving AbsRel by up to 5% and reducing SILog by 8% over strong baselines. By coupling multi-modal intelligence with structured geometric reasoning, DEAN establishes a new paradigm for adaptive, context-aware depth estimation in dynamic real-world domains.
Distributed active noise control (DANC) systems, typically employing the filtered-x least mean squares algorithm, are widely used for spatial noise reduction in acoustic networks. However, due to acoustic coupling among the nodes, vulnerability to impulsive interference, and degradation in non-Gaussian noise environments, the performance of these systems degrades. Additionally, most of the existing DANC frameworks neglect practical limitations in communication, such as transmission delays and interruptions, which limits their scalability in real-world deployments. To address these challenges, an adaptive improved distributed stochastic gradient (AIDSG) algorithm grounded within the primal-dual optimization framework is introduced in this paper. The proposed approach initiates an informed primal-dual update across nodes with instantaneous dual feedback, which results in resilience against colored and impulsive noise, accelerates convergence, and increases adaptability to dynamically changing acoustic environments. Four types of algorithm variants are presented, which address the trade-off between convergence speed and steady-state error. Simulation results on a five-node network validate the superiority of the AIDSG algorithm over conventional DANC and diffusion-based schemes while demonstrating better residual reduction performance, faster mean square convergence, and a better mean noise ratio under additive white Gaussian noise, colored noise, and impulsive noise conditions.
Electrocardiography (ECG) is a widely used, non-invasive tool for assessing cardiac function, but conventional disease-centric models do not fully capture overall cardiovascular health. Recent work has introduced the concept of ECG age: a neural network–predicted age derived from ECG signals. Its difference from chronological age, known as delta age ($\Delta$Age), has emerged as a surrogate marker of cardiovascular well-being. While deep learning approaches have shown promise for ECG age estimation, their computational complexity and lack of interpretability limit deployment in compute-constrained clinical environments. Kolmogorov–Arnold Networks (KANs) offer parameter efficiency and improved interpretability, yet existing variants remain compute-heavy, underexplored for regression tasks, and unable to disentangle contributions from individual ECG leads. To address these challenges, we propose LeadKAN, a lightweight and explainable KAN architecture for ECG age estimation. LeadKAN is built on LoRKAN layers, a novel layer design that replaces fully connected layers with low-rank bilinear mixing followed by an RBF-kernelized top, significantly reducing parameter count and computation. LeadKAN achieves ECG age estimation performance (MSE $\approx$ 112; MAE $\approx$ 8.25 years) comparable to state-of-the-art models, while requiring 16× fewer parameters and 45× fewer multiply–add operations. Additionally, lead-specific encoders enable attribution analysis, thereby enhancing clinical interpretability. These results position LeadKAN as an efficient and explainable framework for ECG age estimation, with strong potential for deployment in real-world, compute-limited settings.
Camera Image Quality Metrics (IQMs) are widely used to characterise imaging systems for human as well as machine vision applications, yet their relationship to the latter remains insufficiently understood. Recently introduced image information metrics have been proposed as alternatives to traditional IQMs, but their behavior under controlled imaging degradations and relevance to vision tasks require further study. This work analyzes the behavior of traditional camera IQMs and image information metrics under controlled blur and noise. We collected a dataset consisting of two parts. The first consists of laboratory captures of four objects, created with systematic camera variations in defocus, noise level, exposure value, and camera-to-object distance. The images included a test chart to enable direct measurement of imaging metrics from each scene. For the second part, a physicsbased imaging pipeline simulation was used to generate synthetic images with independently controlled blur and noise levels. Across both laboratory and simulated data, image information metrics exhibited greater sensitivity to the combined effect of optical and noise degradations compared to traditional IQMs. Multiple object detection networks were evaluated and qualitative relationships between traditional IQMs, image information metrics, and detection performance were examined. Although detection performance variations were modest, certain image information metrics, particularly the ideal observer signal-to-noise ratio, exhibited qualitatively more consistent monotonic trends than traditional IQMs. Overall, the results indicate that image information metrics show greater sensitivity to combined blur and noise degradations than traditional IQMs, while also highlighting the limitations of imaging metrics in predicting task-level performance, which remains content dependent.
Deep neural networks have achieved impressive performance across a wide range of tasks, but this success often comes with substantial computational and storage costs due to large-scale training data. Dataset distillation addresses this challenge by constructing compact yet informative datasets that enable efficient model training while maintaining downstream performance. However, most existing approaches primarily emphasize matching data distributions or downstream training statistics, with limited attention to preserving high-level semantic information in the distilled data. In this work, we introduce a semantic-aware perspective for dataset distillation by leveraging Contrastive Language-Image Pretraining (CLIP) as a semantic prior for post-sampling. Our goal is to obtain distilled datasets that are not only compact but also semantically class-discriminative and diverse. To this end, we design three semantic scoring functions that quantify class relevance, inter-class separability, and intra-set diversity in a pretrained semantic space. Based on image pools generated by existing distillation methods, we further develop a two-stage strategy for effective sampling: the first stage filters semantically discriminative samples to form a reliable candidate set, and the second stage performs a dynamic diversity-aware selection to reduce redundancy while preserving semantic coverage. Extensive experiments across multiple datasets, image pools, and downstream models demonstrate consistent performance gains, highlighting the effectiveness of incorporating semantic information into dataset distillation.
Event-to-video (E2V) reconstruction has gained significant attention recently for its advantages in enabling high dynamic range and fast motion capture capabilities. However, event data encodes only relative brightness changes, lacking the absolute intensity information necessary for accurate reconstruction. Recent methods incorporate previously reconstructed images to provide intensity references but process them in the spatial domain where low- and high-frequency components are highly coupled. This spatial processing typically leads to the degradation of fine details and introduces artifacts such as over-smoothing, blurring and low contrast reconstruction. To address this, we propose a deep spatio-temporal and frequency guided fusion network for E2V reconstruction (DSTFN-E2V), featuring a dual-path architecture with two key components: i) a prior frequency decomposition module (PFDM), and ii) a spatio-temporal event-driven feature extraction module (STEM). The PFDM decouples low- and high-frequency information from previously reconstructed images and current event voxel grid via a 2D discrete wavelet transform, processing the low-frequency subband through residual blocks to preserve structural coherence and intensity references, while an edge-detail refinement module (ERM) enhances edge and texture details from high-frequency subbands. The frequency-specific features from PFDM and the spatio-temporal features from STEM are then integrated through the proposed event-image fusion blocks (EIFBs) that apply cross-attention across three encoder stages, enabling simultaneous structural preservation and detail recovery. Experiments on four real-world datasets demonstrate that DSTFN-E2V achieves state-of-the-art results with 12% SSIM improvements while being 50% faster than recent attention-based methods, with superior edge fidelity and reduced artifacts.
Accurate spatial acoustic characterization is crucial for immersive audio applications, such as virtual and augmented reality, which require low-latency, high-fidelity rendering of sound fields. Traditional approaches, including physics-based simulations and centralized deep learning models, face significant challenges in scalability, data efficiency, and robustness under sparse measurement conditions, and typically require retraining or fine-tuning whenever new measurements are acquired. This paper introduces DAME, a Distributed Acoustic Mixture-of-Experts framework for large-area spatial acoustic modeling. DAME decomposes the global sound field reconstruction task into localized sub-problems, each handled by a compact neural expert trained on region-specific data, while a lightweight gating network combines the experts' predictions based on spatial queries. Unlike state-of-the-art methods, the framework supports incremental expansion on the receiver side. Specifically, new experts can be added to cover previously unsampled regions, and only the gating network is updated, avoiding full retraining of the model and preserving previously learned behavior. Evaluated on both simulated and real-world higher-order Ambisonics room impulse responses, DAME consistently outperforms state-of-the-art parametric, kernel-based, and physics-informed baselines in terms of reconstruction accuracy, directional estimation, and energy decay preservation. Computational analysis demonstrates its suitability for low-latency operation, and a perceptual evaluation confirms superior audio quality and spatial fidelity. The framework therefore provides a scalable, data-efficient alternative for robust spatial audio reconstruction in challenging sparse-data regimes.
Traditional hybrid narrowband active noise control (HNANC) systems typically rely on fixed or energy-based variable step-size adaptation. These adaptation methods sometimes result in suboptimal convergence and poor trade-offs between convergence speed and steady-state performance in the presence of nonstationary tonal disturbances and secondary-path modelling errors. These methods often necessitate persistent external references and face challenges in managing feedforward-feedback loop interactions, which impairs robustness in practical narrowband noise environments. This paper proposes a temporal coherence-driven adaptive control mechanism for HNANC systems, in which the instantaneous coherence between a predefined reference signal and the residual error is used as a tracking-confidence metric to regulate adaptive step-sizes. The temporal coherence emphasizes structured narrowband components while suppressing updates driven by incoherent residuals, thereby improving adaptation behavior under nonstationary conditions. The proposed approach integrates a supporting control filter with coherence-based step-size scheduling to coordinate feedforward and feedback adaptation using normalized least mean square algorithms. Simulation and experimental results under multiple narrowband noise scenarios demonstrate faster retracking and improved stability compared to state-of-the-art HNANC methods, with a modest increase in computational load.
We present a novel zero-shot video editing framework that extends pre-trained text-to-image (T2I) models with optical flow coherence enforcement to achieve temporally consistent and high-quality video edits. By seamlessly integrating optical flow into the attention mechanism, our method propagates edits across frames, effectively addressing challenges such as motion and deformation while minimizing artifacts common in frame-wise editing. A key innovation of our approach is its ability to condition edits on both text prompts and reference images, enabling precise and flexible control over style and content. It supports diverse editing tasks, including text-based editing, image-to-video style transfer, and a combination of the two. Compared to existing (text-only) video editing tools, it improves temporal coherence and visual fidelity significantly, achieving state-of-the-art results in zero-shot settings.Our code is available at https://github.com/AviadDahan/MFF.