Diffusion-based face restoration that adjusts the sampling trajectory of pre-trained diffusion models has achieved remarkable progress. However, existing approaches provide insufficient constraints during reverse diffusion, causing identity-related structural drift and degraded fidelity under severe degradations. To address this, we propose WaveFreqAnchor, a training-free framework based on Wave-Structural Anchoring and Frequency Correction Diffusion. Specifically, Anchor-Space Wave-Structural Guidance (ASWG) constrains facial structures through anisotropic wave-response consistency, while Multi-scale Wavelet-Fourier Injection (MWFI) aligns the predicted low-frequency subband with the observation by replacing its phase, correcting inconsistencies accumulated during reverse diffusion. For real-world scenes, we further introduce Subband High-Frequency Enhancement (SHE), which performs bounded, spatially masked refinement on the predicted high-frequency subbands to recover fine facial details under unknown compound degradations. Together, these designs effectively preserve facial identity while restoring sharp and realistic facial details. Extensive experiments show that our method consistently outperforms existing methods, achieving high-quality and high-fidelity face restoration.
Accurate detection of pesticide residues in agricultural products is crucial for food safety and control strategies. Conventional chromatographic methods suffer from complex operation, high cost, and poor suitability for realtime monitoring. To overcome these limitations, this study investigates a low-cost approach integrating Visible-Near Infrared (Vis-NIR) reflectance spectroscopy with machine learning and deep learning algorithms to detect imidacloprid residues in spinach, and systematically evaluates 19 preprocessing pipelines to select the optimal strategy for different models. We acquired 1600 Vis-NIR spectra from spinach leaves and applied 19 preprocessing pipelines, including interpolation, Savitzky-Golay (SG) smoothing, spectral derivative, wavelet-threshold denoising, and mean centering, to enhance signal quality and reduce noise. Three classification models-Support Vector Machine (SVM), Random Forest (RF), and 1D-Convolutional Neural Network (1D-CNN)-were developed using these datasets for comparative analysis. Under optimal preprocessing, SVM and RF achieved 99.27 % accuracy, F1-scores above 0.9950, and AUC approaching 1.0. The 1D-CNN achieved 97.81 % accuracy, F1-score of 0.9854, and AUC of 0.9940, but showed the lowest sensitivity to preprocessing variations (performance fluctuation of 1.29 %) and higher confidence in identifying positive samples (0.7445). Overall, RF and SVM excel in simple classification, while 1D-CNN offers superior adaptability and generalization in spectral analysis.
BACKGROUND:Quantitative determination of total phosphorus (TP), an indirectly absorbing aquatic indicator, using near-infrared (NIR) spectroscopy is challenged by high-dimensional, noisy, and nonlinear spectral data. Furthermore, traditional data-driven models tend to neglect underlying physical principles, resulting in overfitting and physically inconsistent predictions. METHOD:We propose PICSEN, a Physics-Informed Convolutional-Sequential Dual-Branch Fusion Network. Its architecture synergistically fuses global representations, extracted by a CNN from PCA features, with localized sequential dependencies captured by a GRU from key spectral sequences. To enhance physical consistency, a specialized regularization term is introduced. Unlike traditional methods, it learns an effective absorption proxy to reconstruct the original spectra, thereby embedding implicit physical constraints tailored for TP's indirect optical response within an end-to-end training framework. SIGNIFICANT FINDINGS:Through rigorous repeated validation and statistical testing, PICSEN achieved an average R2 of 0.9380 ± 0.0191, demonstrating competitive and robust performance across all benchmarks (p<0.05). Ablation studies confirmed the critical contributions of both the dual-branch architecture and the physics constraint, with the latter serving as a primary driver for model stability. The model demonstrated high stability across random seeds and enhanced resilience to Gaussian noise. SHAP analysis and saliency maps further validated that PICSEN aligns with known physicochemical absorption regions, indicating strong physical consistency within the studied aquatic matrix. While the current findings are based on a specific river basin (N=235), the adaptable nature of the effective absorption proxy provides a robust framework for regional water quality monitoring, with promising potential for recalibration across diverse hydrological environments.
The restoration of painted cultural relics is crucial for cultural heritage conservation, but existing deep learning methods are mostly oriented towards general images, making it difficult to meet their special requirements of faithfully restoring materials, colors and cultural semantics. To this end, we have constructed the first benchmark for painted cultural relic restoration, and proposed the RelicFormer method, which guides restoration by distilling the global semantic prior of the pre-trained Vision Transformer (ViT). On the test set containing 889 murals, this method significantly outperforms existing technologies in metrics such as PSNR (29.87dB) and SSIM (0.9287). In particular, it exhibits stronger structural fidelity and semantic coherence in moderately and severely damaged areas, providing an effective solution and evaluation benchmark for high-fidelity digital restoration of cultural relics.
Generating background music highly aligned with video content and emotional tone remains a core challenge in cross-modal multimedia computing. Existing methods often rely on shallow feature concatenation, failing to capture the non-linear mappings and long-range temporal dynamics between visual semantics and musical emotions. To address this challenge of deep emotional alignment, we propose VES-BGM (Video-Emotion Synergy Background Music), a cross-modal deep fusion framework driven by emotional trajectories. We establish a synergistic mechanism for audiovisual synchronization through three novel components: (1) an emotion entropy enhancement strategy that reconstructs discrete probability distributions into high-density semantic embeddings for fine-grained modeling; (2) a bidirectional temporal emotion encoder that leverages Bi-LSTM to explicitly model the continuous evolution of video emotions, resolving logical discontinuities caused by abrupt emotional shifts; and (3) an emotion-guided cross-modal modulation mechanism via multi-layer cross-attention, which forces visual features to dynamically focus on relevant semantic cues, achieving a paradigm shift from simple content matching to deep emotional resonance. Experiments on the MuVi-Sync dataset demonstrate that VES-BGM significantly outperforms state-of-the-art baselines such as Music Transformer and CMT in both chord prediction accuracy and subjective audiovisual coherence. Our results validate that an emotion-trajectory-driven modulation architecture effectively bridges the semantic gap in cross-modal generation, providing a robust paradigm for intelligent music scoring.
Biochemical oxygen demand (BOD) is an important indicator that can directly reflect water bodies' degree of organic pollution, Real-time monitoring of water BOD is significant in water resource protection and water environment improvement. The traditional BOD measurement method will consume a lot of human and material resources, and the measurement cycle is long, which can not quickly reflect the changing conditions of the water body, and can not realize the timely and effective early warning of sudden water pollution events. With the wide application of machine learning in the field of water monitoring, to solve the problem of difficulty in obtaining the input variables of the machine learning model and the existence of missing values. we further combine the hyperspectral technology to realize the accurate and rapid estimation of the BOD content of the water body. The raw spectral data of ten BOD standard liquids with different concentrations were collected, and 100 sets of transmission spectral data were obtained by whiteboard correction. A noise reduction technique based on PCA transmission spectra reconstruction is proposed, which utilizes the PCA algorithm to extract the principal component eigenvectors of the original transmission spectra and then reconstructs the whole dataset by using the first part of the principal component eigenvectors whose cumulative variance contribution rate reaches a certain percentage. The first 2, 10, and 15 principal component feature vectors were used in the experiment to reconstruct the transmission spectral data and compared with the traditional noise reduction methods for spectral data, We combined the SVM model and BP neural network model to establish a model for estimating the BOO content of water bodies, The results showed that the BPNN model was superior to the SVM model regarding regression accuracy and degree of fit, and the noise reduction effect was more significant. The model using the first 2 feature vectors reconstructed for noise reduction did not fit as expected, probably due to the loss of information, The BPNN model with the first 10 feature vectors reconstructed for noise reduction performed the best with an RMSE of 0.040 6 and an R of 0.980 3. The reconstruction of the first 15 feature vectors did not improve the noise reduction effect, probably because more than 10 feature vectors added redundant information. The experiments verified the feasibility of noise reduction using PCA reconstruction of transmission spectra and provided a new idea for estimating the BOD content of water bodies.
Existing pavement crack inspection methods heavily rely on a large amount of annotated samples and require overburdened computational power, which is not affordable on edge devices. Recent methods mostly focus on spatial features, which cannot effectively capture the long, continuous, tender, and thin road features. The masked image modeling (MIM) is an effective way to rebuild the crack primitive features by masking strategy on unlabeled data to reduce the dependence on annotated data, and the convolution on the frequency domain provides a flexible and efficient path to capture continuous and tender road features. Inspired by the thoughts of masked frequency modeling (MFM), we proposed a masked discrete cosine transform (DCT)-domain modeling strategy, named crack masked DCT-domain modeling (CrackMDM), for efficient pavement crack segmentation. Specifically, we propose a DCT-domain masked modeling method in the CrackMDM model, which combines the advantages of separable convolutions and spectral convolutions (SP-Convs) in the DCT domain to extract continuous and tender crack structures. Additionally, we introduce the self-supervised pretraining with a masking strategy in the DCT domain using unlabeled crack samples to build crack primitives and to fine-tune the encoder and decoder parameters on labeled crack data to refine crack features in the fine-tuning phase. The CrackMDM model is evaluated on three public benchmarks: CFD, YCD, and GAPs, and achieves state-of-the-art (SOTA) performance with superior inference speed. Codes are available at https://github.com/Jyuan357/CrackMDM
The accurate detection of pollution levels in water bodies using transmission spectrum data and fusion algorithms has become crucial for safeguarding water resources, Inaccurate predictions and detection frequently result from the high dimension of transmission spectrum data and model instability. The Yangtze River water body's total phosphorus concentration content is predicted in this study, and an accurate and environmentally friendly approach is suggested to achieve this goal. In particular. maxi min normalization and mean centering are two preprocessing operations carried out on the Yangtze River's measured water quality transmission spectrum data. These operations remove noise while eradicating differences between different data magnitudes, guaranteeing the consistency and reliability of the data. In addition, to solve the problem of the high dimension the transmission spectrum data, the KPCA method is used to reduce the dimension of the data and extract the features. The KPCA method is used to select the top 6 principal components that represent 99.42% of the information content of the original data for subsequent prediction model training by finding a classification plane in a high dimension space, Then, the foundation of the initial particle swarm algorithm, the particle initialization rule, multiple swarm competition strategy, parameter adaptive update strategy, population diversity guidance strategy, and particle variation mechanism are added to improve the particle warm's capacity for optimization and prevent particles from trapping in the local optimal solution. Additionally, the improved article swarm algorithm optimizes the initialized weights and parameter values in the BP neural network to accelerate the convergence of the network and improve prediction performances, Finally, the total phosphorus content of the samples in the test set was predicted using the IMCPSO-BPNN model. The experimental results showed an R-2 of 0.975 786, an RMSE of 0.002 242, and an MAE of 0.001 612. The IMCPSO-BPNN model suggested in this work has a better fitting effect and better Accuracy in forecasting the total nitrogen concentration in the Yangtze River water body when compared to other models such as The RF model, the BPNN model, and the PSO-BPNN model. It offers fresh concepts and viewpoints for studying and applying predictive modeling using transmission spectrum data and fusion algorithms to protect water resources and environmental management.
Developing an accurate and efficient comprehensive water quality prediction model and its assessment method is crucial for the prevention and control of water pollution. Deep learning (DL), as one of the most promising technologies today, plays a crucial role in the effective assessment of water body health, which is essential for water resource management. This study models using both the original dataset and a dataset augmented with Generative Adversarial Networks (GAN). It integrates optimization algorithms (OA) with Convolutional Neural Networks (CNN) to propose a comprehensive water quality model evaluation method aiming at identifying the optimal models for different pollutants. Specifically, after preprocessing the spectral dataset, data augmentation was conducted to obtain two datasets. Then, six new models were developed on these datasets using particle swarm optimization (PSO), genetic algorithm (GA), and simulated annealing (SA) combined with CNN to simulate and forecast the concentrations of three water pollutants: Chemical Oxygen Demand (COD), Total Nitrogen (TN), and Total Phosphorus (TP). Finally, seven model evaluation methods, including uncertainty analysis, were used to evaluate the constructed models and select the optimal models for the three pollutants. The evaluation results indicate that the GPSCNN model performed best in predicting COD and TP concentrations, while the GGACNN model excelled in TN concentration prediction. Compared to existing technologies, the proposed models and evaluation methods provide a more comprehensive and rapid approach to water body prediction and assessment, offering new insights and methods for water pollution prevention and control.
Road features in remote sensing images typically exhibit characteristics such as thinness, continuity, and varying textures. Accurate road extraction requires effective modeling of multi-scale contextual information and high-frequency boundary details, but existing methods face limitations in this regard.Traditional attention modules struggle to balance low-frequency semantics and high-frequency structures, resulting in blurred edges and poor connectivity; additionally, multi-scale information fusion relies on a global attention mechanism with a unified dimension, which increases computational burden. To address these issues, this paper proposes a remote sensing road segmentation network based on dynamic spectrum sensing and mixed attention. The network combines the Frequency-Adaptive Spatial Attention (FASA) module and the Mixed Frequency Agent Attention (MFAA) module, enhancing the model’s perception ability in two aspects: frequency-space fusion and multi-scale modeling. The FASA module improves the fusion of details and global semantic information through a spatial-frequency dual-branch architecture, while the MFAA module adopts a lightweight design to optimize multi-scale information interaction in skip connections, reducing computational overhead and enhancing contextual modeling capabilities.By integrating multi-scale feature representations from both spatial and frequency domains, the network can more effectively capture road structures at different frequency levels. Compared to existing methods, the proposed network significantly improves the F1 score, accuracy, recall, and mean intersection-over-union (mIoU) on the Massachusetts and DeepGlobe public road datasets while maintaining minimal parameter size and computational complexity.
Convolutional neural networks (CNNs) have demonstrated strong capabilities in hyperspectral image (HSI) classification. However, it is still a challenge to adaptively adjust the size of the receptive fields (RFs) of CNNs base on the information of different scales in HSI to achieve adaptive selection of spectral–spatial features. In the paper, we modify the convolutional block attention module (CBAM) and propose a modified-CBAM-based network (MCNet) to adaptively select spectral–spatial features for HSI classification. In particular, the modified CBAM not only enables the model to adjust its RF size according to the information of different scales in HSI, but also enables the model to achieve a joint focus on important spectral and spatial features. This is very important to adaptively select more descriptive and discriminative spectral–spatial features. The proposed MCNet is compared with currently popular methods on Indian Pines, Kennedy Space Center, University of Pavia, and Botswana HSI datasets. The results show that MCNet has better classification results than other methods on overall accuracy, average accuracy, and Kappa.
Background: With the increasing severity of global water pollution, accurate prediction models of water pollution content are critical for effective environmental management. However, traditional methods often exhibit low prediction accuracy for pollutant concentrations when data samples are limited and do not adequately address data noise. This study focuses on predicting total phosphorus (TP) concentrations in the Yangtze River Basin by integrating data augmentation and denoising methods with spectral technology and deep learning, using water samples collected from Wuhan to Anhui, China. Method: The study utilized an improved Conditional Generative Adversarial Networks (CGAN) for data augmentation, increasing dataset diversity and training effectiveness. Adaptive threshold wavelet denoising is applied to reduce noise and improve data quality. A Convolutional Neural Network (CNN) with a coordinate attention (CA) mechanism is used to extract key spectral features linked to TP concentration prediction. Significant Findings: This study introduces an innovative approach that combines advanced CGAN-based data augmentation, adaptive threshold wavelet denoising, and a CNN model incorporating a CA mechanism, achieving high accuracy in TP concentration prediction. The proposed model outperforms traditional methods, achieving R2 = 0.9805, RMSE = 0.0019, and MAE = 0.0009. This novel method significantly enhances prediction performance, providing an effective solution particularly in scenarios with limited data samples.
Deep learning has demonstrated significant advantages in managing nonlinear relationships within high-dimensional spectral data, making it widely applicable in water quality monitoring. However, the variety of model selection and construction strategies has resulted in substantial fluctuations in predictive performance, particularly with high-dimensional data. This study constructs an integrated deep learning framework for predicting water pollutant concentrations, incorporating several key modules including data preprocessing, frequency decomposition, feature enhancement, sample augmentation, and decoder regression prediction. In the established model, an improved wavelet transform algorithm is first employed to address the issue of original data being unable to effectively distinguish detailed features, thereby accurately extracting the periodicity and volatility characteristics of the data. Secondly, an encoder module based on the Informer architecture enhances various frequency domain features and further improves the quality of features and their correlation with labels through distillation techniques. Subsequently, an improved generative adversarial network is introduced to tackle the problem of small sample data by effectively augmenting the limited dataset, thereby enhancing the overall quality of the dataset. Finally, a decoder module combining an optimization algorithm and an improved convolutional neural network (IMCPSO-RCNN) effectively addresses the shortcomings of traditional models in hyperparameter optimization and predictive performance, achieving efficient and accurate regression prediction of pollutant concentrations. A case study in the middle and lower reaches of the Yangtze River shows that this model outperforms others in prediction accuracy, achieving coefficients of determination (R-2) of 0.9785, 0.9733, and 0.9741 for TN, COD, and TP, respectively. The root mean square error (RMSE) values are 0.0601, 0.6248, and 0.0023, while the mean absolute error (MAE) scores are 0.0252, 0.2810, and 0.0006, respectively. The necessity and effectiveness of each model component are validated through ablation experiments. This research offers an efficient and unified deep learning solution for monitoring water pollutants. Synopsis: This deep learning framework enhances water quality monitoring by accurately predicting pollutant concentrations, informing environmental policy and water system management.
Convolutional Neural Network (CNN) has been widely used in precipitation downscaling. However, it is unclear whether the predictor region size significantly influences the precipitation downscaling results using CNN for predictor-predictand mapping. In this paper, we perform sensitivity experiments on various predictor areas in CNN-based precipitation downscaling. Specifically, we select the middle reaches of the Yellow River (MRYR) as the study area (predictand region). For the predictor areas, we expand , , , and in four directions based on the MRYR, respectively. These sensitivity experiments indicate that the predictor region size significantly affects the precipitation downscaling results. The result of the precipitation downscaling expanded to over the MRYR performs best in Root Mean Square Error (RMSE) of spatial-temporal distribution. Specifically, on mean precipitation, it reduces RMSE by 5.71% and 12.77% relative to the Base experiment (predictor area is the MRYR) in space and time, respectively. For extreme precipitation, RMSE decreases by 2.12% (4.27%) and 12.9% (14.02%) in space and time compared to the Base experiment for R95P (R99P), respectively. Then, the downscaled precipitation results deteriorate when continuing to expand the predictor area. That is mainly because the thermodynamic and dynamic variables near the study area significantly affect local precipitation. When the predictor area is continuously expanded without restriction, the complexity of nonlinear relationships amongst climate variables may markedly increase, resulting in many redundant features during downscaling, thereby reducing downscaling performance. Therefore, our results suggest that appropriately expanding the predictor aera may positively influence the downscaling of regional precipitation.
The middle reaches of the Yellow River (MRYR), located in northern China, are the transition zone between semi-arid and semi-humid climates. As one of the climate-sensitive regions in China, MRYR has a fragile ecological environment and serious soil loss, which leads to geological disasters such as landslides, collapses, and mudslides caused by extreme precipitation. However, scarceness of high-resolution precipitation data over MRYR limits assessment of the environmental impacts caused by climate change, especially for extreme precipitation. In this article, we design a Residual-in-Residual Dense Block based Network (RRDBNet) model for the statistical downscaling of precipitation in MRYR, and compare the proposed RRDBNet with a generalized linear regression model (GLM) and two popular deep-learning-based models. The multi-level residuals and dense connectivity strategies introduced in RRDBNet help it to learn more abstract features and complex nonlinear relationships among climate variables to improve downscaling performance. The results show that the proposed RRDBNet has good performance in precipitation simulations, which can reproduce the spatial-temporal characteristics of high-resolution precipitation well. RRDBNet reduces the root-mean-squared error (RMSE) by 19% and improves the Pearson correlation coefficient (CC) by 6% relative to GLM for climatology mean precipitation. Especially, RRDBNet has substantial improvements in extreme precipitation compared with other models. It reduces RMSE by 58% (79%) and improves CC by 38% (145%) relative to GLM for R95P (R99P), where R95P and R99P represent extreme precipitation and very extreme precipitation, respectively. For the probability density function of daily precipitation, it is further demonstrated that RRDBNet performs better as regards extreme precipitation frequency. Our results suggest that statistical downscaling based on RRDBNet may be an effective tool for historical and future climate simulations from global climate models. A Residual-in-Residual Dense Block based Network (RRDBNet) model is designed for the statistical downscaling of regional precipitation. The evaluation shows that RRDBNet has substantial improvements in precipitation and extreme precipitation compared with other two deep-learning models. For probability density function of daily precipitation, it is further demonstrated that RRDBNet performs better as regards extreme precipitation frequency. Our results suggest that statistical downscaling based on RRDBNet may be an effective tool for historical and future climate simulations from global climate models. image
This study focuses on the detection of safety helmets at construction sites, aiming to improve detection accuracy and robustness. The algorithm design is based on the YOLOv5 framework, with various optimizations made to the original structure. Firstly, the CA (Coordinate Attention) mechanism is incorporated to enhance the model's learning and utilization of crucial features. This attention mechanism helps improve the network's resistance to interference in complex backgrounds, thus making the model more suitable for safety helmet detection tasks in real construction environments. Additionally, the SIoU (Spatial Intersection over Union) loss function is introduced to replace the traditional CIoU loss function, aiming to measure spatial overlap more accurately in object detection, thereby enhancing the model's precision and robustness.
Extraction of road from remote sensing (RS) images confronts a dual challenge: the inhomogeneous intensity and inconsistent contrast. Historically, conventional approaches predominantly relied on spatial-domain convolutional neural networks (CNNs), which were constrained by their local receptive fields, limiting their ability to effectively capture extensive and intricate road features. In addition, self-attention mechanisms with global awareness have a high computational load, and recent image frequency-domain representation learning makes up for this shortcoming. Based on this, the paper presents a road segmentation network, specifically designed using a U-Net architecture, called lightweight spatial-frequency domain learning U-Net (LSFDLU-Net). In this network, spatial-frequency domain learning module is used as the basic component. Firstly, a spatial-channel Fourier neural operator (SC-FNO) based on frequency-domain learning is proposed. SC-FNO employs gate-controlled filtering to coarsely process the frequency-domain representation of road features in RS images, globally suppressing interference information related to the road. Subsequently, a multi-scale feature fusion module based on spatial-domain learning is devised, consisting of multi-scale depth-wise separable convolutions (MSDSC), feature separation (FS), and attention-guided fusion (AGF). It finely processes RS images by separating the original, high-frequency, and low-frequency spatial domain representations, and employs MSDSC and AGF to effectively cap-ture slender road features. The network significantly improved Fl scores, accuracy, recall, and mean Intersection over Union (mIoU) for the Massachusetts and DeepGlobe public road datasets while maintaining a smaller number of parameters and computations compared to related methods.
Unsupervised embedding learning aims to learn highly discriminative features of images without using class labels. Existing instance-wise softmax embedding methods treat each instance as a distinct class and explore the underlying instance-to-instance visual similarity relationships. However, overfitting the instance features leads to insufficient discriminability and poor generalizability of networks. To tackle this issue, we introduce an instance-wise softmax embedding with cosine margin (SEwCM), which for the first time adds margin in the unsupervised instance softmax classification function from the cosine perspective. The cosine margin is used to separate the classification decision boundaries between instances. SEwCM explicitly optimizes the feature mapping of networks by maximizing the cosine similarity between instances, thus learning a highly discriminative model. Exhaustive experiments on three fine-grained image datasets demonstrate the effectiveness of our proposed method over existing methods. (c) 2024 SPIE and IS&T
Atmospheric turbulence can often introduce phase errors into a propagating light field, thus resulting in anisoplanatic and temporally varying blur and distortion of images. Restoring such images degraded by atmospheric turbulence is extremely ill-posed, due to multiple plausible solutions for a given input image. Most methods offer a deterministic estimation of clean images and require high-computational costs. To address these challenges, this article proposes a fast turbulence mitigation network (FTMNet). It is a lightweight model for atmospheric turbulence mitigation. Differing other methods, it does not employ a strategy for producing a single deterministic reconstruction. Instead, it leverages the Monte Carlo method to enhance restoration performance and produces a different and reasonable set of reconstructed images for a given input. As a result, FTMNet effectively mitigates atmospheric turbulence effect while maintaining low-inference time and computational resource requirements. Experimental results demonstrate that FTMNet shows high-inference speed, reaching 90 fps, and outperforms the state-of-the-art peers.