
Sound event detection (SED) requires simultaneous event classification and precise temporal localization. Existing methods face three orthogonal challenges: frequency‑axis translation equivariance in standard convolution, difficulty in capturing events of vastly different durations, and imbalance between local and long‑range dependency modeling in recurrent or Transformer models. Additionally, strong label scarcity limits generalization. To address these issues, we propose a convolutional recurrent neural network (CRNN)‑based model, Slide‑ATST‑FDy‑Hybrid‑SED (Audio Teacher‑Student Transformer, Frequency Dynamic Convolution, and Hybrid‑minGRU), with four synergistic innovations: an ATST frame‑wise module providing 40‑ms resolution features; frequency dynamic convolution that adapts kernel weights to frequency positions, breaking the physically invalid equivariance; a slide‑window module fusing global and local temporal contexts for multi‑scale perception; and a hybrid‑minGRU module that parallelizes lightweight minGRU with self‑attention to capture both local patterns and long‑range dependencies efficiently. A mean teacher semi‑supervised framework with two‑stage training leverages weakly labeled and unlabeled data. On the DESED 2023 Task 4 dataset, our model achieves an Event‑based F1 of 0.636, PSDS1 and PSDS2 (Polyphonic Sound Detection Score) of 0.589 and 0.824—relative improvements of 46.9
This paper presents a floating three-terminal CFOA-based memtransistor emulator and investigates its suitability for artificial synaptic plasticity applications. The proposed three-terminal emulator is realized using six current-feedback operational amplifiers (CFOAs), one analog multiplier, eight resistors, and a capacitor, and is formulated to reproduce gate-controlled, memory-dependent drain–source behavior under a near-zero gate-leakage assumption. A theoretical analysis is developed from the constitutive mem-element framework to derive the terminal relations of the emulator and to clarify the role of the internal state path in establishing memory-dependent conductance. PSpice simulations show that the proposed circuit exhibits the pinched hysteresis behavior expected from a memtransistor emulator. To demonstrate functional relevance beyond terminal characterization, the proposed emulator is further employed in an artificial synaptic plasticity application.
Nowadays, the diffusion algorithms have been widely used in processing large-scale data because of the advantage of parameter estimation of multiple nodes at different locations. Considering the poor convergence of existing diffusion algorithms in the errors-in-variables (EIV) model containing generalized Gaussian noise, the diffusion generalized maximum total correntropy (DGMTC) algorithm is proposed in this paper. This algorithm significantly enhances the robustness against generalized Gaussian noise by introducing the generalized maximum correntropy criterion (GMCC) into the diffusion-based distributed adaptive estimation utilizing gradient-descent total least-squares (DGDTLS). In addition, the DGMTC algorithm is analyzed for local mean stability and steady state mean square performance. Finally, the excellence of the DGMTC algorithm compared with other algorithms and the correctness of the theoretical analysis are demonstrated by simulation.
The development of the discrete fractional Hankel transform (DFRHT) depends fundamentally on the generation of orthonormal eigenvectors of the symmetric kernel matrix of the discrete Hankel transform (DHT). For the DFRHT to approximate its continuous counterpart, namely, the fractional Hankel transform (FRHT), the eigenvectors in question should approximate samples of the eigenfunctions of the Hankel transform (HT), which are the products of generalized Laguerre polynomials, a Gaussian function, and a power function. Consequently, the target eigenvectors will have the descriptor “Laguerre-Gaussian-power-like (LGPL)”. The recently developed techniques for generating optimal eigenvectors of any unitary kernel matrix can be classified as either indirect (in the sense of first generating initial eigenvectors as a prerequisite for generating the final optimal ones) or direct (in the sense of not requiring initial eigenvectors). Moreover, the direct techniques can be viewed as either batch (in the sense of generating a partition of the modal matrix corresponding to a single distinct eigenvalue as a whole) or sequential (in the sense of generating the columns of the partition one after another). The present paper aims to assess the performance of the recently developed direct batch and sequential generation algorithms of optimal orthonormal eigenvectors when applied to the real symmetric kernel matrix of the DHT. This research endeavor has not been undertaken before. The assessment will demonstrate the relative merits of each algorithm.
Nighttime image flare removal is a challenging task critical for autonomous driving and surveillance. Unlike global degradations, flare exhibits locality, diversity, and nonuniformity, with halos, streaks, and diffuse artifacts often co-occurring and coupling with background scenes. Existing methods suffer from incomplete removal and noticeable artifacts. To address these issues, this paper proposes a Multi-Scale Spatial-Frequency Synergistic Network (MSFNet), which adopts an explicit parallel architecture. A spatial-frequency synergistic module adaptively separates low-frequency halos from high-frequency details, enabling synergistic fusion with spatial features. A differential gated fusion module, driven by feature differences, integrates shallow details and deep semantics via spatial gating and channel attention, effectively suppressing artifacts. Extensive experiments on Flare7K++ validate that MSFNet achieves a PSNR of 27.71 dB on real-world images and 31.33 dB on synthetic images, outperforming the next-best method by 0.12 dB and 0.39 dB, respectively. It surpasses state-of-the-art algorithms in terms of both global and local metrics, which demonstrates the effectiveness of spatial-frequency synergistic learning for flare removal and provides a new technical pathway for nighttime image restoration.
Underwater image enhancement is complicated by spatially uneven optical degradation. Current methods often struggle to reconcile the dual goals of restoring local textures and removing distant haze within a single global mapping. To address this, we propose a depth-guided dual-stream network that explicitly decouples the enhancement process. Rather than applying uniform processing, our framework utilizes depth priors to extract pure geometric contexts and generate soft attention masks via a feature refinement mechanism, explicitly differentiating between foreground and background regions. Adopting a divide-and-conquer strategy, one stream adaptively restores high-frequency details in the near field, while the other employs physical transmission cues to correct non-linear scattering and color distortion in the far field. Furthermore, to rectify global color casts, we introduce an illumination-color refinement stage that embeds a color constancy mechanism during feature decoding to simulate the stability of human vision. Experiments across multiple datasets demonstrate superior performance in both quantitative metrics and visual quality, effectively restoring natural color balance and vivid details. Additionally, such high-quality restoration provides a solid foundation for downstream visual applications. This provides compelling evidence that explicitly decoupling spatial heterogeneity via depth priors is an effective strategy for addressing complex underwater image degradation.
This paper investigates the qualitative behavior of semilinear Hilfer-nabla fractional discrete time systems characterized by inherent memory effects. The considered model employs the Hilfer-type fractional difference operator, which interpolates between the Riemann-Liouville and Caputo nabla operators and provides additional flexibility through the order and type parameters. By transforming the system into an equivalent fractional Volterra summation equation, an explicit representation of the solution is obtained via discrete multi-parameter Mittag-Leffler functions using the Picard’s successive approximation method. Based on this representation, sufficient conditions ensuring the existence, uniqueness, finite-time stability, and attractive stability of the system are derived. Finally, two numerical examples are presented to validate the theoretical findings and illustrate the effectiveness of the proposed stability criteria.
In this paper, an adaptive correlated Gaussian approximation filter (A-CGAF) is introduced for nonlinear state estimation problems when there are correlated noises of process and measurement. Classical statistics-fixed filtering techniques have to rely on a priori knowledge of static noise correlation; otherwise, their performance will be degraded. This paper seeks to improve the problem by developing a new algorithm that takes advantage of a neural network to estimate the noise cross-correlation in an online fashion based on a correlated Gaussian approximation filter structure. The effectiveness of the approach is demonstrated through simulations conducted under static, sudden, and gradually changing noise cross-correlation environments. Simulation results show that the proposed A-CGAF has superior estimation accuracy, robustness, and fast convergence performance compared with the fixed statistics-based competitors. More specifically, it delivers a 20.43
In this paper, we introduce the working of signals, systems, and signal processing on totally weighted graphs. We discuss some real-world systems on graphs by considering sensing points as vertices, connection between sensing points as edges, assigning signal values as vertex weight, measure of connection between vertices as edge weight, and try to ensure the efficiency of the system by assigning a signal condition on each vertex. By suggesting a matrix representation for this system, we try to connect the matrix energy as a measure of the energy of the signal and system. In addition, signal processing is represented on totally weighted graphs considering a different type of edge weight. Also, discussed are the sampling of analog signal and representation of the digitalized signal on totally weighted graph. We also consider the energy of signal, system, and signal processing matrix as a measure of efficiency.
The growing deployment of unmanned aerial vehicles in sensitive fields such as military operations demands high levels of image security and authentication. To address the increasing risk of tampering, this study proposes a robust watermarking and encryption framework designed to ensure both the integrity and confidentiality of unmanned aerial vehicles images. The proposed hybrid technique integrates non-subsampled shearlet transform, QR decomposition and multi-level singular value decomposition for robust and imperceptible watermark embedding. To further secure the data, double random phase encoding is applied for encryption, followed by compression using compressive sensing to enable efficient storage and transmission. The framework is evaluated using standard image quality and security metrics under various attack conditions. Experimental results demonstrate that the method preserves the visual quality of unmanned aerial vehicles images after decryption and offers strong resilience to common signal processing and noise attacks. The watermark remains reliably detectable, with minimal loss of image fidelity. The proposed framework ensures secure transmission and authentication of unmanned aerial vehicles images while maintaining high image quality. Its layered architecture shows significant potential for deployment in real-time unmanned aerial vehicles—based applications, especially where security and data integrity are critical.
Robust adaptive filtering is a key technique for sparse system identification. However, traditional algorithms struggle to balance robustness, convergence speed, and steady-state accuracy under impulsive and heavy-tailed noise. The logarithmic Student’s t-based filtering method offers good robustness to impulsive noise due to its heavy-tailed error modeling, but its fixed zero-centered symmetric assumption introduces steady-state bias under nonzero-mean noise and does not fully leverage the system’s sparsity prior. To address this, we propose the Logarithmic Student’s t-based Generalized Maximum Variable-Center Correntropy (LSGMVCC) algorithm, which dynamically adjusts the error distribution center using a sliding-window median to compensate for nonzero-mean bias. Additionally, we introduce a sparsity-enhanced variant, SN-LSGMVCC, incorporating a sparsity-aware regularization term for improved convergence and steady-state accuracy in sparse system identification. Simulation results demonstrate that the proposed algorithms achieve faster convergence and lower steady-state mean square deviation under various non-Gaussian noise environments, significantly enhancing the robust identification performance of sparse systems.
Digital signal processors predominantly use Carry Select Adder (CSLA) arithmetic units. This paper presents the design and implementation of multiply and accumulate (MAC) units through five proposed multistage hybrid CSLAs to reduce computational complexity and to improve performance metrics. Five different CSLA structures are proposed with novel power, area optimized carry generation and propagation blocks, i.e., Proposed Block-1 and Proposed Block-2. Based on these optimized blocks, two multistage square-root (SQRT)-based CSLAs, two SQRT-based hybrid CSLAs with fast add-on multiplexing (FAM), and one combined hybrid CSLA are developed. The proposed architectures are functionally verified using a System Verilog test bench in Xilinx-Vivado and implemented in the Cadence system design tool using 90 nm technology to obtain performance parameters such as area, delay, and power. In comparison with the existing state-of-the-art CSLA architectures, Proposed-I, III, and V achieved area reductions of up to 35–53
Maximum-likelihood (ML) direction-of-arrival (DOA) estimation for vector hydrophone arrays provides strong statistical performance. However, multidimensional nonlinear search still causes a heavy computational burden. To address this problem, this paper proposes an ML-DOA estimation method based on an improved velvet worm optimization (IVWO) algorithm. The original velvet worm optimization (VWO) algorithm balances global exploration and local exploitation, but its direct application to complex ML-DOA objective functions exhibits insufficient initial search coverage, an abrupt exploration–exploitation transition, and limited late-stage local search. Accordingly, IVWO introduces randomized Hammersley initialization, proposes an iteration-dependent three-stage smooth quota allocation strategy, and incorporates logarithmic-spiral local search to improve initial coverage, transition smoothness, and local search, respectively. Simulations under different SNRs, numbers of snapshots, population sizes, and numbers of sources show that IVWO-ML achieves faster convergence than the Improved Secretary Bird Optimization Algorithm (IMSBOA), Improved Invasive Weed Optimization (IIWO), Genetic Algorithm (GA), Particle Swarm Optimization (PSO), the original VWO, and Grey Wolf Optimizer (GWO), while also attaining lower RMSE and better run-to-run stability, especially in closely spaced multi-source scenarios. Compared with the traditional Multiple Signal Classification (MUSIC) and Alternating Projection Maximum Likelihood (AP-ML) baselines, IVWO-ML also provides higher estimation accuracy and stronger robustness under low-SNR and closely spaced multi-source scenarios.
In supervised speaker diarization, speaker misclassification errors remain a significant challenge despite advances in embedding representations and deep learning architectures. This work introduces a voting-based centroid–distance post-processing framework that refines diarization outputs by reassigning ambiguous speaker embeddings through consensus across multiple distance measures. Specifically, Mahalanobis, geodesic, fast dynamic time warping (DTW), and Euclidean distance measures are used to determine cluster affiliation, enabling more reliable correction of misclassified segments. The proposed approach is evaluated using ECAPA-TDNN and x-vector embeddings in conjunction with CNN, Bi-LSTM, and Conformer models, with ECAPA-TDNN consistently outperforming x-vectors across all configurations. On the CallHome dataset, agglomerative hierarchical clustering (AHC) yields a diarization error rate (DER) of 29.28
Accurate direction of arrival (DOA) estimation in impulsive noise remains a critical challenge in array signal processing. To address this issue, this paper proposes TSUC-Net, a novel two-stage deep learning framework that integrates a U-Net architecture with a convolutional neural network (CNN). In the first stage, a U-Net-based denoising network is designed with a learnable complex soft-clipping (LCSC) module placed at the input front-end. Unlike fixed nonlinear transformations, the LCSC module adaptively suppresses alpha-stable distributed impulsive noise by learning an optimal amplitude threshold while preserving phase information critical for spatial processing. Since the LCSC limiter’s parameters are fully adaptively learned, the burden of parameter tuning associated with traditional methods is eliminated. The denoised signals are then fed into the second-stage CNN for robust DOA estimation. Simulation results demonstrate that the proposed framework outperforms existing algorithms, including traditional subspace methods and recently proposed nonlinear transformation-based techniques, in terms of accuracy and generalization capability under various impulsive noise conditions.
Reconstructing a signal from a magnitude-only short-time Fourier transform (STFT) remains challenging because missing phase information introduces audible and structural artifacts. This paper proposes Time-domain Enhanced Spectrogram Alignment (TESA), a gradient-based framework that directly optimizes temporal samples so that the STFT magnitude of the reconstruction matches a specified target spectrogram. Unlike Griffin–Lim that updates phase implicitly through alternating projections between time and time–frequency domains, TESA performs sample-level optimization using a closed-form gradient that maps magnitude mismatch to an inverse STFT-based update, and employs adaptive moment estimation optimizer to stabilize and accelerate convergence. We conduct two experiments under controlled oracle conditions by providing the clean target spectrogram to isolate phase-retrieval capability from magnitude estimation errors. We additionally report performance when the target spectrogram is estimated from the observation to quantify practical impact. Across denoising and source separation benchmarks, TESA improves reconstruction quality relative to representative baselines (Griffin–Lim, optimal-transport alignment, deep neural network, and alternating direction method of multipliers), yielding average gains of 1.94 dB signal-to-noise ratio and 3.86 dB signal-to-distortion ratio under oracle condition, with consistent advantages in time–frequency mean square error. Runtime–resolution analysis further shows that TESA supports reduced execution time through lower-resolution STFT settings and fewer iterations. These results indicate that TESA provides a flexible optimization framework for magnitude-constrained reconstruction, with performance that remains competitive when moving from oracle to non-oracle evaluations.
Neural networks have gained more attention for their success in various applications using image, and speech. NNs heavily rely on the arithmetic operators, especially adders during the training and validation process. In Convolutional Neural Networks (CNNs), the adders, multipliers and shifters play the major role. The convolution operation relies heavily on Multiply-and-Accumulate (MAC) computations in which adders are one of the important computational units. Hence it is essential to study the performance of various adders under different number format representations. In this paper, the performances of the adders are compared and analyzed with FPGA implementation. The adders considered are Ripple Carry Adder (RCA), Carry Select Adder (CSLA), Carry Skip Adder (CSA), Carry Look Ahead Adder (CLA), Kogge Stone Adder (KSA), Brent Kung adder (BKA), Han Carlson Adder (HCA), 16-bit, and 32-bit Floating Point Adders (16FPA and 32FPA). Each adder is implemented for 4, 8, 16 and 32-bits of inputs and their outputs (carry and sum) are verified. All the adders are synthesized and simulated in Vivado 2024.2 tool using Verilog code and implemented on Xilinx Artix-7 FPGA (XC7A35T-1CPG236C). Performances of adders in terms of delay, power (dynamic and static), Energy-Delay Product and area utilization using Look Up-Tables (LUTs) are compared. It is found that RCA utilized lowest resources. CSLA consumes highest power due to its high resource utilization. The computation speeds of CLA and KSA are higher and they produce lower delays as compared to the other adders considered.
To address the long-standing trade-off between invisibility and robustness in deep-learning–based digital watermarking, this paper proposes a robust algorithm based on the Unilateral M-Net (UM-Net) architecture. The method incorporates a discrete wavelet transform (DWT) to extract low-frequency components while reducing image resolution. These low-frequency components are then concatenated with feature maps of corresponding sizes within the network. This design alleviates detail loss caused by convolution and pooling operations and improves reconstruction quality. In the extraction stage, a Dense Dilated Convolution (DDC) module is introduced to enhance multi-scale feature extraction by expanding the receptive field and improving extraction accuracy. Furthermore, a BCH error-correction code is employed to lower the bit-error rate (BER) and strengthen overall robustness. Experimental results demonstrate that the proposed method achieves an average PSNR of 37.61 and an average SSIM of 0.9729 on the test set. Under various noise attacks, it maintains strong robustness while preserving high invisibility.
Approximate multipliers are crucial in error-resilient image processing, machine learning, and high-performance computing systems. These systems often use ASICs and FPGAs to achieve better speed and energy efficiency compared to general-purpose systems. It focuses on saving hardware resources while maintaining acceptable performance and accuracy, particularly in applications such as digital signal processing and machine learning, where resource constraints are crucial. Arithmetic operations are accelerated by fast hardware multiplication, especially in algorithms that compute large-scale numbers. FPGAs are unable to use ASIC-based techniques presented in the literature due to architectural differences. Unlike conventional multiplier architectures that generate all partial products followed by multiple reduction and addition stages, the proposed LUT-based design (LBD) reduces hardware complexity at the partial product generation (PPG) stage itself by exploiting LUT sharing and absorption of the ‘11’ input pair, thereby reducing the number of generated partial products and subsequent addition hardware. Building on this principle, this work proposes one accurate and three approximate unsigned multipliers using novel LUT-based decoder architectures for both partial product generation (PPG) and partial product reduction (PPR), effectively leveraging FPGA LUT and carry-chain resources for improved area efficiency. The proposed 8 × 8 approximate multipliers utilize 41.67
Elliptic curve cryptography (ECC) over binary finite fields relies heavily on polynomial multiplication, making the efficiency of the underlying multiplier architecture a critical factor in cryptographic hardware design [23]. Among various approaches, Karatsuba multiplication is widely adopted due to its reduced arithmetic complexity; however, aggressive parallelism and improper resource reuse often lead to excessive power consumption and thermal infeasibility on FPGA platforms [11, 26]. This paper presents a comprehensive architectural exploration of seven 256-bit Karatsuba-based binary polynomial multiplier architectures over Galois Field (GF) (2256), each employing a different resource-reuse strategy ranging from fully parallel to fully sequential implementations. All architectures are implemented on a Xilinx Artix-7 FPGA using identical synthesis and power estimation settings to ensure a fair comparison in terms of area, latency, power consumption, and junction temperature. The study reveals that fully parallel Karatsuba architectures, although fast, are thermally impractical due to excessive switching activity, while fully sequential designs suffer from high latency. Based on these insights, a balanced resource-reuse architecture is identified, which selectively reuses multiplier blocks from 16-bit to 128-bit levels while retaining parallelism at lower levels. Compared to the fully sequential Karatsuba architecture, the proposed balanced design reduces computation latency by more than 26 × (121 cycles vs. 3280 cycles) while maintaining thermally safe operation (73.0 °C) and consuming only 18