
In this paper, we investigate a class of weak dual frames generated by nonhomogeneous sub-wavelet systems on the half-line [Formula: see text]. Unlike classical wavelet frame theory on [Formula: see text] or [Formula: see text], the lack of a group structure on [Formula: see text] prevents the direct application of standard Fourier analysis. To address this, we work within the framework of the Walsh-Fourier transform associated with the operation “⊕”. We introduce a class of nonhomogeneous sub-wavelet systems and establish the Walsh-Fourier domain equation for their (weak) oblique dual frames in [Formula: see text]. Furthermore, we present a mixed oblique extension principle (MOEP) for such dual frames.
Enhancing visibility and reducing noise in low-light images significantly improves the usability of images in surveillance, medical imaging, and autonomous driving, where clarity in poor lighting is crucial for accurate interpretation. Many low-light image enhancement approaches struggle to effectively balance brightness enhancement and detail preservation, often resulting in either overexposed regions or loss of fine details. To overcome this issue, this manuscript proposes enhancing visibility and reducing noise in lowlight images through disentangled frequency-guided processing and progressive graph convolutional networks (LLI-DFP-PGCN). Initially, input images are gathered from the Smartphone Image Denoising Dataset (SIDD) and undergo initial enhancement using the Confidence Partitioning Sampling Filtering (CPSF) method, which improves visibility by selecting high-confidence pixel regions, filtering noise, and guiding the enhancement process. The enhanced images are then processed using the Advanced Disentangled Frequency Paradigm for Low-Light Image Enhancement (AFD-LLIE), which separates lower and higher frequency components of an image to independently enhance illumination and preserve fine details. Finally, the frequency refined images are passed to the Progressive Graph Convolutional Networks (PGCN), which models complex spatial-frequency relationships to restore realistic colour, preserve texture, and eliminate residual noise. The performance of the proposed LLI-DFP-PGCN approach is evaluated using Structural Similarity Index Measure (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Learned Perceptual Image Patch Similarity (LPIPS), Natural Image Quality Evaluator (NIQE), Mean Absolute Error (MAE) and Computational time. The proposed approach attains high PSNR, and high SSIM compared to existing techniques like Denoising diffusion post-processing for lower-light image enhancement (DDPLLIE-CNN), a Self-supervised network for low-light traffic image enhancement depending on deep noise and artifacts removal (LTIE-DNAR-GAN) and A Joint Network for Low-Light Image Enhancement Depending upon Retinex (LLIEDCNN) methods respectively.
Image classification is a fundamental task in computer vision that entails categorizing images into pre-defined classes according to their visual features. Traditional classification approaches mostly relied on handcrafted feature extraction techniques, which often struggled to handle complex patterns, high variability and large-scale diverse datasets. Recent advancements in deep learning have highly impacted image classification tasks by enabling automatic hierarchical feature learning directly from raw image data. This review presents a comprehensive analysis of advanced deep learning architectures designed for image classification, focusing on architectural innovations, feature representation, transfer learning strategies and model interpretability. Different network architectures, such as residual networks, recurrent attention-based networks, fully convolutional networks, capsule networks and region-based convolutional networks, are critically reviewed based on their design and performance across different application domains. The review further explores the role of pre-trained models and knowledge transfer techniques in addressing challenges, data scarcity and training complexity. In addition, it discusses widely used development frameworks that facilitate efficient model implementation and deployment. By integrating both theoretical knowledge and recent developments, this review provides systematic guidance for researchers and practitioners to develop efficient, scalable, and robust image classification systems for both scientific research and industrial applications.
Transformer-based image super-resolution has become a critical technique for enhancing visual quality in digital photography, remote sensing and video surveillance. However, the quadratic computational complexity of self-attention has led to a surge of token compression strategies, where an inherent conflict arises between coarse-grained token representations and high-fidelity image reconstruction. Existing methods typically overlook the guiding role of semantic and structural priors in token dependency modeling, resulting in suboptimal attention allocation and limited token compression ratios. To address these issues, we propose a novel Semantic-Aware and Structural-Focused lightweight Super-resolution Network. Specifically, a semantic-aware token interaction module is designed to efficiently associate semantically similar yet spatially distant regions, facilitating complementary information aggregation from low-resolution inputs. Subsequently, pre-modeled structural priors queried from semantic cues are leveraged to guide the reconstruction of fine-grained details. Moreover, we design a structural-focused token aggregation module that adaptively focuses on similar structures, significantly reducing redundant computations among irrelevant tokens. In the feedforward process, group-wise global-local interaction mechanisms are employed to enable efficient information propagation while maintaining lightweight computation. Extensive experiments demonstrate their effectiveness and efficiency. Compared with the strongest baseline, our method achieves up to 0.44dB improvement in PSNR while delivering 1.6 & times; faster inference speed.
To address the demand for efficient detection of ballast geometry degradation in China’s railways, this paper proposes a fast train-borne inspection method based on point cloud diff-graph analysis method (DGAM). The method converts train-borne LiDAR point clouds into diff-graphs, integrates a standard ballast template, and employs a diff-graph projection module (DGPM) to extract geometric deviations. It further utilizes Michelson-like contrast (MLC) enhancement to suppress false defects caused by template mismatch. Subsequently, a defect detection module (DDM) is applied to achieve precise defect localization. Experimental results demonstrate that DGAM achieves high precision, with IoU exceeding 80.70%, and high speed, supporting inspection at 160[Formula: see text]km/h in the simulation tests, outperforming traditional contrast enhancement methods. It has been integrated into a high-speed inspection vehicle, providing an effective technical solution for railway ballast maintenance.
Handwritten Telugu scripts in a developing stage provide substantial opportunities for further research and analysis. Optical character recognition (OCR) is considered an optical-based technology that includes a machine to automatically understand digitized and scanned characters. Yet, limited research studies have been conducted in recent times for Handwritten Character Recognition (HWCR) models for Telugu text. Hence, this research study implements an effective Telugu HWCR using an adaptive deep learning network model. Initially, the raw images are taken from diverse benchmark datasets for further processing. Subsequently, the gathered images are fed into the developed recognition network, named dual attention vision transformer-based adaptive residual dilated dense network (DViT-ARDDNet). Here, the DViT technique initially splits the words into characters from the input handwritten images, allowing the RDDNet model to effectively recognize the Telugu characters. Furthermore, hyperparameters are selected optimally by Statistically Improved Botox Optimization (SIBO). Therefore, the outcome of the model elucidates that the proposed network has the ability to appropriately identify the individual Telugu characters.
Over the past few years, nuclear norm-based matrix regression (NMR) has achieved superior performance in the field of image recognition. By employing the nuclear norm as the loss metric to preserve the two-dimensional structural information of error images, NMR exhibits strong robustness against structural noise. However, NMR fails to effectively exploit the category information of training data and ignores the local structural relationships among samples. To overcome the above limitations, we design a new classification method named locality-sensitive competitive matrix regression (LSCMR). First, a competitive representation term is introduced into the objective function of NMR, enforcing training samples from distinct categories to competitively characterize the test sample. To further boost the discriminative ability of representation coefficients, LSCMR incorporates the local structural information between samples, which encourages samples of the correct class to contribute more to the representation process. The alternating direction method of multipliers (ADMM) is utilized to solve the objective function of LSCMR. Experimental results on several public datasets verify the superiority of LSCMR. Specifically, on subset 4 and subset 5 of the Extended Yale B dataset, LSCMR achieves recognition accuracy of 97.37% and 62.18%, respectively, which are 9.21% and 20.02% higher than those of NMR. The source code of LSCMR is available at https://github.com/yinhefeng/LSCMR.
Geomagnetic-induced currents in power networks and transmission lines pose challenges for so many nations as they produce many unexpected effects. This paper contributes to the modeling and monitoring of geomagnetic induced currents in a power network, with an aim of minimizing the energy in the signals using the optimal control method, specifically, the Hamiltonian method. This paper specifically employs the Haar wavelet transform to investigate the optimal energy in selected Namibian power network substation’s geomagnetic induced current’s data during the two geomagnetic disturbance events in 2013 and 2015. The results reveal some significant phenomena during specific times of the day, which suggest the need for appropriate network disturbance mitigation strategies for the selected stations. Another major contribution is that the control managed to keep the energy as low for most of the period, illustrating that optimal control technique presents a good tool for geomagnetically induced current modeling.
Quaternion matrix completion (QMC) aims to accurately recover multi-channel data exhibiting inter-channel dependencies from incomplete observations. However, in the scenarios with large singular value disparities, traditional nuclear norm-based methods often fail to capture the true low-rank structure, which results in recovery deviation. Meanwhile, the noncommutativity and structural complexity of quaternion algebra pose significant challenges for quaternion low-rank modeling and optimization. Despite their empirically promising performance, existing QMC methods generally provide convergence analysis but lack a bound of the convergence. To address these challenges, we propose a novel quaternion completion method based on a smooth log-determinant surrogate with a dynamic parameter termed QCSLD, which effectively captures the low-rank structure of hypercomplex multi-channel data even under severe singular value disparities. To ensure numerical stability and theoretical convergence, we develop an efficient iterative algorithm and provide an explicit quadratic bound on the convergence behavior under mild assumptions, which ensures predictability and stability for QMC. Experiments on synthetic and real datasets further validate that QCSLD achieves superior reconstruction accuracy compared to general comparison methods.
This paper addresses 2-frames and 2-Riesz bases based on a 2-inner product induced by a Hilbert space inner product. We establish some links among frames, Riesz bases, 2-frames and 2-Riesz bases; and investigate the existence and characterizations of 2-Riesz bases.
The discrete Gabor analysis behaves excellently in digital signal processing. Now that defined on the real line & Ropf; has seen great achievements, but that defined on half real line & Ropf;(+) has not. Due to & Ropf;(+) being not a group under usual "+", the discrete Gabor analysis on & Ropf;(+) differs from that defined on & Ropf;. Due to & Ropf;(+) being an addition group under a new addition "circle plus", this paper addresses discrete periodic Gabor analysis on & Ropf;(+) associated with such addition. Using "circle plus"-based discrete Zak transform, we characterize a class of discrete periodic Gabor frames (Riesz bases, orthonormal bases) on & Ropf;(+) and their weak Gabor duals. Several examples are also provided to illustrate our results.
Representation-based classification (RC) methods have been extensively studied in visual recognition tasks. However, some early methods were typically based on the mean squared error (MSE), which is highly sensitive to gross corruption. Although various robust strategies have been proposed, they still suffer from insufficient robustness and are sensitive to outliers. Consequently, achieving robust modeling in complex scenarios with gross corruption and outliers remains a significant challenge. To address this concern, this paper proposes a novel RC method, called joint outlier detection and representation learning (JODRL), which integrates outlier detection and representation learning into a unified framework. By mutually boosting each other, JODRL can effectively reduce the impact of outliers and further improve robustness. Furthermore, we propose an efficient alternating optimization algorithm based on the alternating direction method of multipliers (ADMM) and half-quadratic (HQ) theory. Experimental results on five representative benchmark datasets demonstrate that JODRL achieves significantly better performance than existing methods under various complex noise and outlier interference scenarios, fully validating the model's effectiveness and robustness.
Accurate neuron stitching across large-scale electron microscopy volumes is crucial for reconstructing complete neural circuits. We propose TransStitch, a distributed Transformer-based framework that addresses these challenges by integrating multimodal feature fusion with topology-aware self- and cross-attention mechanisms to model global structural dependencies across adjacent electron microscopy blocks. To refine uncertain predictions, a dynamic 1-nearest-neighbor strategy progressively converts the probabilistic connectivity matrix into discrete associations without relying on a fixed threshold. Additionally, a mapping-based lazy relabeling strategy reduces merging complexity from voxel to fragment level, significantly improving computational efficiency and scalability. Extensive experiments on public electron microscopy datasets (SNEMI3D, CREMI-C, FIB25) with ground truth demonstrate superior stitching accuracy of proposed method compared to baseline, while qualitative evaluation on a large-scale, self-collected zebrafish whole-brain dataset confirms coherent 3D reconstruction across tens of thousands of sections. These results highlight TransStitch as an accurate and scalable solution for large-scale connectomics reconstruction.
Face Recognition is the process of identifying people by extracting their facial features, and it is widely utilized in several applications, including authentication, healthcare, and security. The traditional approaches faced troubles in providing better accuracy and computational efficiency due to the lack of identifying the facial patterns. Therefore, a Root Cause Analysis (RCA) is essential in a face recognition system to prevent failures in recognizing faces. Hence, the Channel and Spatial Attention-based Explainable Convolutional Network (CSA-ECNet) model is proposed to enhance the face recognition results through detecting the defects and analyzing the root causes. The incorporation of an explainable technique helps to provide insights to the CSA-ECNet model in detecting the root causes, thereby increasing the performance of the CSA-ECNet model in recognizing faces without any failures. The incorporation of the Channel and Spatial Attention (CSA) facilitates increasing the accuracy by enabling the CSA-ECNet model to selectively concentrate on the vital spatial regions and feature channels, which strengthens the model's ability to handle various aspects, including poor lighting. Experimental results demonstrate the exceptional performance of the CSA-ECNet model, reporting the high sensitivity of 97.79%, specificity of 98.44%, and accuracy of 98.11% for 90% of training data on Face Recognition Dataset.
In [Deepshikha and A. Samanta, On spectrally optimal duals of frames generated by graphs, preprint (2024), arXiv:2406.00776], authors studied spectrally optimal dual frames for 1-erasure and 2-erasures of frames generated by graph. In this paper, we study spectrally optimal dual frames for r-erasures. We show that the spectral radius of the error operator of unitary equivalent frames is same with respect to their respective canonical dual frames. We prove that if a frame is generated by a connected graph, then its canonical dual frame is a unique spectrally optimal dual frame for r-erasures. Further, we show that the canonical dual of frames generated by disconnected graphs is a non-unique spectrally optimal dual frame for r-erasures.
Recognizing historical documents is vital in protecting cultural heritage by facilitating access, search and analysis of important archival materials. Nonetheless, current techniques face difficulties due to factors like poor image quality, diverse handwriting styles, erased or missing words, damaged documents and intricate page designs. These challenges affect precise text extraction and reduce the overall efficiency of automated recognition systems. In this work, a Pine Makeup Optimization-enabled Convolutional Generative Transformer Network (PMO_CGTN) is proposed for missing character recognition in historical documents. First, an input historical document image is applied for image enhancement using the Multi-scale Gray World Algorithm. Then, segmentation of each text line and segmentation of each word within the lines is performed using the Semantic Text Segmentation Network (STSN). Finally, missing character recognition and filling of missed characters are accomplished using CGTN. Here, a Convolutional Neural Network (CNN) model is modified by incorporating a Generative Pre-Trained Transformer (GPT) layer to form CGTN, which is trained using a Pine Makeup Optimization (PMO), and is a merging of Pine Cone Optimization Algorithm (PCOA) and Makeup Artist Optimization Algorithm (MAOA). Lastly, an Optical Character Recognition (OCR) document is obtained as the output.
Speech recognition from noisy speech signals remains a significant challenge in the field of human-computer interaction. Speech signals are often degraded by reverberation and background noise, leading to reduced recognition accuracy. Conventional methods typically rely on large volumes of training data and fail to generalize effectively under noisy conditions. Moreover, these approaches often exhibit limited robustness and reduced performance in adverse environments. In this work, a new technique called Fractional Super Bilterling Optimization with Random Multimodal Deep Learning-Convolutional Neural Network (FrSBO_RMDL-CNN) is proposed for speech recognition from unclear speech signals. At first, the unclear speech signal collected from the database is denoised by employing WaveNet denoising. Next, the voice enhancement is performed utilizing the Speech Enhancement and Language Model (SELM), which is trained using the Super Bilterling Optimization (SBiO). Here, the SBiO is developed by the integration of Bilterling Fish Optimization (BFO) and Superb Fairy-wren Optimization Algorithm (SFOA). After that, the speech word is segmented employing the Attentional Encoder-Decoder. Last, the speech recognition is performed using the RMDL-CNN, which is tuned by Fractional Super Bilterling Optimization (FrSBO). Furthermore, FrSBO_RMDL-CNN computed the maximum Negative Predictive Value (NPV), recognition accuracy and Positive Predictive Value (PPV) of 97.146%, 97.798% and 96.888%, respectively.
The wavelet neural networks (WNNs) have emerged as one of the most powerful computational models for solving complex nonlinear problems across various domains. They are thus currently a focus for many researchers. WNNs combine neural networks (NNs) from artificial intelligence with the robust mathematical framework of wavelets. This review paper aims to explore and present the current status of research in the field of WNN. To start with, a brief description of NNs is provided along with the history, construction and types of wavelets. Based on the literature review, we have categorized WNNs into three categories, namely, wavelet activation NNs (WANN), wavelet synapse NNs (WSNN) and wavelet transformation NNs (WTNN). These three types of WNNs are discussed in detail, including their construction and applications. The pros and cons of each type of WNN are also presented. Another focus of this paper is to identify the types of wavelets that have majorly been used till now and also the wavelets which are less common in WNN structures. Despite significant progress, several research gaps remain, it has been found that only a few of the first-generation wavelets, namely the Morlet, Gaussian and Mexican-hat wavelets, have been used in WNNs hitherto. Along with the literature and gaps, this research paper demonstrates that the incorporation of wavelets into a NN algorithm improves the accuracy and enhances the capability of capturing physical features in the model. It can be used as a guide to further the work in the field of WNNs.
We present a weighted competitive nonnegative representation (WCNR) method. Specifically, WCNR introduces a competitive term, which leverages the competitive ability of training data from distinct classes to strengthen the connection between the representation and classification phases. In addition, WCNR incorporates a weight constraint. The weight constraint of each class imposed on the representation coefficients endows similar classes with more representation contributions, which boosts the discriminative power of the representation coefficients. To assess the classification performance of WCNR, extensive experiments are carried out on benchmarking datasets. Experimental results confirm that WCNR exceeds the state-of-the-art representation-based classification methods (RBCM) and also surpasses several deep learning approaches. The MATLAB code of WCNR is available at https://github.com/yinhefeng/WCNR.
In this paper, we consider designing a parallel-form filter with tunable bandwidths while ensuring its stability. The tunable filter is implemented by parallelizing multiple second-order (2nd-order) blocks, leading to a parallel structure for high-speed signal processing. Since those blocks have variable coefficients, the parallel structure has continuously tunable frequency-bandwidths. To guarantee the stability of the parallelized 2nd-order blocks, we present a new function for transforming the block parameters so as to satisfy the stability condition. By exploiting the parallel structure alongside the transformations, we further develop a two-phase procedure for designing the parallel-form bandwidth-tunable filter. As compared to the existing designs based on general-form structures, the design using this parallel structure can obtain a parallel-form bandwidth-tunable filter, leading to high-speed frequency-bandwidth tuning during signal processing. The parallel structure also facilitates modular hardware implementation. This is because the entire structure parallelizes only the 2nd-order blocks, and thus those blocks act as the basic modules for implementing the whole structure. Verification design is included for exemplifying both stability and high performance of the resulting parallel-form bandwidth-tunable filter. The design results confirm that the bandwidth-tunable filter maintains its stability during frequency-bandwidth tuning in signal processing. Furthermore, the bandwidth-tunable filter also exhibits high performance.