
With the rapid advancement of generative artificial intelligence (AI), the visual fidelity of synthesized images has increased dramatically, posing serious challenges to the verification of digital content authenticity. Existing AI-generated image detection methods often suffer from limited generalization and robustness, particularly when confronting unknown generative models or complex post-processing perturbations. To attenuate such deficiency, we propose an AI-generated image detection scheme. Leveraging a frozen contrastive language–image pre-training with Vision Transformer as visual backbone, the network extracts and stacks multi-scale intermediate features from the transformer modules to effectively capture both low-level and high-level forensic fingerprints. Based on this representation, we introduce an improved convolutional block attention module, which adopts a cascaded design by first applying channel-wise attention and then spatial attention. This design enables the network to adaptively select informative feature hierarchies while strengthening the representation of local generative artifacts. To further optimize the feature space structure, we propose a hard-sample-aware contrastive learning loss. It dynamically mines hard samples to enhance intra-class compactness and inter-class separability. In addition, we construct a mixed-source training image dataset named mixed-source AI-generated image dataset, which covers diverse generative paradigms. Large-scale testing results show that our proposed scheme ranks first on average across seven benchmark datasets, with accuracy 82.96% and the area under receiver operating characteristic curve 93.43%, demonstrating its outstanding generalization ability. Code is publicly available at https://github.com/multimediaFor/MIDNet.
Despite the success of deep learning in remote sensing (RS) image classification, substantial domain shifts—stemming from heterogeneous sensors and diverse environmental conditions—frequently compromise model reliability. Although source-free unsupervised domain adaptation (SFUDA) has emerged as a critical paradigm to bypass data privacy and storage constraints, existing methods remain fragile in complex RS scenes where noisy pseudo-labels often trigger catastrophic semantic drift. We propose PromptRefine, a white-box SFUDA framework designed to anchor target adaptation through cross-modal intelligence. Specifically, we leverage the zero-shot semantic priors of large-scale vision–language models (e.g., contrastive language–image pre-training) to rectify source-biased predictions via a dynamic prompt fine-tuning mechanism. The framework executes a three-stage alternating optimization strategy that integrates cross-modal semantic alignment, hard-sample mining via sliced Wasserstein discrepancy, and fine-grained prompt evolution. Evaluated across 18 cross-domain tasks on five benchmarks (UCM, WHU-RS19, AID, RSSCN7, and NWPU-RESISC45), PromptRefine consistently achieves remarkable performance. Specifically, it outperforms the leading vision transformer-based baseline (VisTA) by 0.42% and 0.47% on the two groups of cross-domain tasks and surpasses the top ResNet-based baseline (SRKT/DFENet) by 2.94% and 5.02% in average classification accuracy. The proposed method also outperforms other SFUDA and unsupervised domain adaptation methods on all 18 tasks, demonstrating its superior adaptation capability. Our approach provides a robust, privacy-preserving, and computationally efficient solution, setting a benchmark for scalable RS scene characterization.
To address the long-tailed distribution problem in particleboard surface defect detection, we propose YOLOv8-Particleboard Long-Tailed (PLT). First, we design the Wavelet-Enhanced Feature Propagation Block , which adopts a ”global decomposition followed by dual-path refinement” strategy based on discrete wavelet transform to alleviate insufficient feature extraction for tail-class defects. Second, we construct the Frequency-Spatial Collaborative Attention Block by introducing a learnable frequency-domain processing path into the Mobile Vision Transformer architecture, enhancing the model’s global modeling capability through frequency-spatial collaborative augmentation. Furthermore, we propose the tail-aware task-aligned assigner, which dynamically adjusts positive sample allocation to strengthen the perception of tail-class defects. Finally, we design the Class-Adaptive Balanced Focal Loss loss function, which achieves automatic balancing of long-tailed distributions through dynamic parameter adjustment, online performance feedback, and a learnable threshold mechanism. Extensive experiments on our custom Particleboard-LT dataset demonstrate that YOLOv8-PLT improves the overall mAP@0.5 by 8.0% compared with the baseline. Notably, it achieves a +21.7% AP gain on the tail-class sand mark while safely maintaining head-class accuracy above industrial deployment thresholds. Moreover, extended evaluations on NEU-DET-LT, a long-tailed dataset reconstructed from the public NEU-DET benchmark, further validate the robust cross-domain generalizability of the proposed method. Without architectural modifications, YOLOv8-PLT effectively counters data scarcity, boosting the AP of the tail-class Sc by 6.4% while simultaneously preserving the head-class performance, proving its effectiveness for imbalanced industrial inspection scenarios.
Low-light object detection remains fundamentally challenging due to the intrinsic misalignment between physical imaging characteristics and detection-oriented feature representations under extreme illumination degradation. This misalignment originates from the inconsistency between sensor-level signal formation and downstream representation learning, leading to unstable feature distributions and degraded detection performance. Existing approaches either rely on enhancement in the sRGB domain or directly learn from RAW data, yet both struggle to effectively bridge this gap. In this paper, we propose IDAM-RAW, a unified RAW-domain framework that bridges physical imaging processes and detection-oriented representation learning through task-driven end-to-end optimization. Specifically, DetISP maps RAW measurements to detection-friendly features without relying on fixed hardware ISP processing. The Residual Illumination Decoupling Module progressively reduces illumination-related variations in feature space and stabilizes optimization, whereas the Adaptive Feature Modulation Module suppresses interference propagation and enhances target-related responses across multiscale features. To support evaluation, we construct CR7-RAW, a real-world low-light bimodal dataset with spatially paired RAW and RGB observations, providing a new benchmark for RAW-based perception tasks. Extensive experiments on LOD, CR7-RAW, and BDD-Night demonstrate that IDAM-RAW achieves consistent performance gains across the evaluated datasets. Averaged over three independent random seeds, IDAM-RAW improves mAP@50 from 40.3 to 72.7 on LOD while maintaining consistent improvements on CR7-RAW and BDD-Night, supporting its effectiveness and cross-dataset robustness.
Guest Editors Xin Ding, Jialie Shen, Xiaoyan Luo, Yinqiang Zheng, and You Yang introduce the Special Section on Computational Imaging and Intelligent Image Processing.
In the domain of unmanned aerial vehicle (UAV) aerial imagery, objects frequently exhibit dense and nonuniform distribution patterns, often resulting in false positives and missed detections. To overcome these challenges, we propose SIG-YOLOv8s, an advanced object detection architecture built upon the YOLOv8s framework. The nomenclature reflects three core enhancements integrated into the model: the small object detection head, the Inner-DS-IoU loss, and the feature gather-and-distribute (FGD) module. First, the FGD module is utilized to process multiscale feature information via dilated convolutions. By optimizing cross-layer recursive information fusion, this module effectively mitigates the adverse effects of varying shooting angles and complex backgrounds. Second, the Inner-DS-IoU loss function is employed to account for bounding box shape and scale during regression, thereby enhancing robustness against complex environments and accelerating convergence. In addition, a dedicated prediction head for small objects is incorporated to capture tiny targets that are typically challenging to discern, significantly improving detection accuracy. Finally, extensive empirical evaluations validate the efficacy of the proposed method, demonstrating that SIG-YOLOv8s achieves superior performance in object detection and recognition tasks.
Real-time recognition of abnormal behaviors in rail transit surveillance demands both high accuracy and low latency, yet existing methods struggle to balance these requirements under complex backgrounds, illumination variations, and occlusions. A real-time multimodal network (RTM-Net) is proposed, fusing RGB video and skeleton keypoints through four synergistic modules. Skeleton-guided cropping eliminates background noise via adaptive ROI selection. Dynamic motion-aware sampling prunes redundant frames based on motion energy. A lightweight adaptive spatio-temporal network captures spatiotemporal dependencies through low-rank decomposition and multiscale dilated convolutions. A dynamic weight fusion strategy further adapts modality contributions based on skeleton quality. A dedicated Rail Transit Video Dataset (RTVD) containing 2536 clips across six action categories is also constructed. Experiments on RTVD, Something-Something V1/V2, and NTU RGB+ D demonstrate that RTM-Net achieves competitive accuracy with only 31.51 GFLOPs (∼32 FPS on RTX 3090), providing an effective solution for real-time safety monitoring in resource-constrained scenarios.
Light field imaging technology enables digital refocusing with “shoot first, focus later” capability, but acquiring its high-dimensional data involves an inherent trade-off between spatial and angular resolution. To address the challenge of refocusing under sparse angular sampling, we present a comparative analysis of two representative technical approaches: end-to-end methods based on deep learning and multi-stage methods based on traditional image processing. Experimental results on the Stanford Light Field Dataset demonstrate that RefocusNet, with its end-to-end architecture, achieves inference speeds in milliseconds, making it well-suited for resource-constrained scenes. In contrast, methods based on bokeh rendering (BR) and super-resolution deliver superior refocusing quality and support post-capture adjustment of the depth of field, making them more appropriate for practical applications. We provide a reference framework for selecting sparse light field refocusing algorithms for different applications.
Smartphone screens are susceptible to micro-scale defects such as tiny cracks and scratches during the production process, and these defects often have extreme aspect ratios and complex background interference, which are usually difficult to identify. However, existing detectors encounter a clear bottleneck when dealing with such issues: high-precision models typically suffer from excessive parameter counts, while lightweight models lack the sensitivity required to detect micro-scale defects. To bridge this gap, a lightweight detection model called HRI-YOLO is proposed. It comprises three key innovations: constructing the hybrid-resolution interactive network, using high- and low-resolution stream interaction design and residual aggregation paths to reduce spatial information loss; the C2f_HGA module is designed to achieve simultaneous extraction of static texture features and geometric topology features using HGAConv; the C2iRMB module, which combines an inverted residual moving block with the window-based self-attention mechanism, is designed to effectively suppress the background interference. Experimental results on the Smartphone Screen Glass Dataset for Defect Detection demonstrate that HRI-YOLO achieves an mAP@0.5, mAP@0.5:0.95, Precision, and Recall of 60.4%, 27.7%, 63.6%, and 59.8% with only 3.52 M parameters, yielding improvements of 6.5%, 1.6%, 12.4%, and 1.2% over YOLO11. This proves the superiority of HRI-YOLO in smartphone screen glass detection, providing a new solution in the field of automated optical inspection for intelligent manufacturing.
Face imaging and recognition are ubiquitous in daily applications, yet transmitting biometric face data to untrusted servers introduces critical privacy risks. Although various privacy-preserving face recognition (PPFR) methods have been proposed, they often suffer from significant degradation in both privacy performance and recognition accuracy under resource constraints. To address these challenges, we propose WHFR, an end-to-end PPFR framework integrating the discrete wavelet transform (DWT). By transforming the original image into the frequency domain via DWT, we first discard the low-frequency sub-band to obfuscate visual information. To thwart reconstruction attacks, we randomly construct high-frequency residuals, which naturally form an underdetermined system, and further combine them with stochastic sign flipping, together yielding a dual-randomization defense. To focus the downstream face recognition model on discriminative features within the perturbed residuals, we introduce a high-frequency enhancement module that employs a task-customized convolutional cosine-similarity attention mechanism, thereby preserving recognition accuracy. Experiments conducted on several benchmark datasets demonstrate that WHFR effectively safeguards visual privacy and defends against adversarial reconstruction. The accuracy dropped by only 4.84% compared with the unprotected baseline—significantly lower than the 9.34% to 12.65% accuracy loss observed in state-of-the-art methods. Moreover, WHFR substantially reduces computational overhead, enabling efficient PPFR.
Video summarization plays a crucial role in efficient long video understanding and question answering. However, existing deep-learning-based approaches suffer from two major limitations. On the one hand, they lack interpretability, often leading to a mismatch between the generated summaries and users’ actual interests. On the other hand, as video sequences grow longer, conventional transformer architectures struggle to maintain long-range contextual dependencies while remaining computationally efficient. To address these challenges, we propose a human–computer interactive video summarization system driven by visual scanpaths. First, we design a visual-question-answering-based scanpath generator that simulates human gaze behavior in a task-oriented manner, effectively filtering out irrelevant content. Second, based on the generated scan paths, we develop an adaptive frame prioritization strategy to extract top-K keyframes consistent with human attention and interest perception. To efficiently preserve the contextual dependencies among these distant keyframes, we further introduce a token compression module, enabling compact visual representations and efficient information propagation with fewer tokens. Finally, we integrate a user-centric, interactive visual question answering system that dynamically customizes video summaries according to user queries. Extensive experiments on multiple datasets demonstrate that our interactive, gaze-driven summaries align better with user preferences than conventional methods.
Current continuous sign language recognition (CSLR) methods often struggle to capture fine-grained spatio-temporal dynamics. Although widely adopted, 3D convolutions and their variants process spatial appearance and temporal motion jointly. This joint processing introduces representational ambiguity, blurring subtle motion cues at sign boundaries and causing the loss of detailed spatial information. Moreover, the standard CTC loss provides supervision exclusively at the final output. This causes severe gradient attenuation in shallow layers, forcing the model to rely on a sparse set of discriminative frames. Consequently, the model overfits to a single dominant alignment path. To address these issues, we propose STFNet, an RGB-only framework comprising a Spatio-Temporal Fusion (STF) module and a Hierarchical Semantic Alignment (HSA) module. STF adopts a factorized fusion strategy. It first extracts multiscale temporal features through lightweight depthwise convolutions while preserving the original spatial features in a separate branch. The two pathways are then fused via a learnable structured mechanism that enforces channel-wise pairing between appearance and motion cues. This enables adaptive local fusion, preserving fine-grained spatial details while injecting temporal motion cues in a learnable manner to produce highly discriminative joint representations. HSA injects auxiliary CTC supervision at multiple intermediate stages, acting as a subnetwork regularizer. It directly backpropagates gradients to earlier layers to counteract gradient attenuation, encourages the exploration of diverse alignment paths, and mitigates the CTC peak-collapse problem. All auxiliary losses share the same classifier, which ensures semantic consistency across scales. Moreover, HSA incurs no additional inference cost, as it is active only during training. Extensive ablation studies and visualizations validate the effectiveness of our approach. Results on PHOENIX14, PHOENIX14-T, and CSL-Daily demonstrate that STFNet achieves state-of-the-art performance using only RGB frames. Code is available at https://github.com/zhanglong908/STFNet.
Accurate segmentation of thyroid nodules in ultrasound images is essential for thyroid cancer risk assessment and computer-aided diagnosis, yet remains challenging due to ambiguous boundaries and significant shape variations. Existing convolutional neural network (CNN)-based methods effectively capture local features but are limited in modeling long-range dependencies and complex boundary structures. To address these limitations, we propose SAFM-Net, a dual-branch CNN–GNN Network that integrates synergistic attention and frequency-domain modulation. The network adopts a dual-branch encoder, where a graph-based branch leverages a synergistic-attention dynamic graph convolution (SA-DGC) module to adaptively model global relationships among feature nodes, enhancing structural and boundary representation. In parallel, a CNN branch captures local textures and fine-grained details. To fuse complementary features, a frequency-domain modulation (FDM) module is introduced to enable cross-branch interaction and hierarchical integration, improving feature representation capability. Extensive experiments on the DDTI and TN3K datasets demonstrate the effectiveness of the proposed method. Compared with GED-Net, SAFM-Net achieves improvements of 0.54%, 1.00%, 1.68%, and 0.73% in terms of Accuracy, Dice, IoU, and Precision, respectively, on the DDTI dataset, and improvements of 0.17%, 0.43%, 0.67%, and 1.48% in terms of Accuracy, Dice, IoU, and Precision, respectively, on the TN3K dataset. These results indicate that SAFM-Net provides accurate and robust segmentation performance under challenging ultrasound imaging conditions.
Object detection in low-light environments is severely degraded by photon starvation, leading to feature concealment, noise amplification during feature fusion, and motion blur caused by long exposure. To address the limitations of the conventional “enhancement-then-detection” pipeline, this paper proposes CIE-Det, an end-to-end lightweight object detection framework that jointly optimizes illumination enhancement and feature learning. A lightweight cascaded illumination enhancement module (CIE module) with 0.0017 M parameters is first embedded at the network input to restore image contrast via downsampled illumination estimation and cascaded nonlinear curves with a dynamic gating mechanism. To alleviate cross-layer semantic inconsistency, a dynamic semantic alignment fusion operator is introduced to replace static feature concatenation with learnable adaptive weighting, enabling effective noise suppression and multiscale feature alignment. Furthermore, an orthogonal asymmetric large-kernel detection head (OAK-Detect) is designed using separable 1×5 and 5×1 convolutions to enhance receptive fields and improve robustness to motion blur. Extensive experiments on the ExDark dataset demonstrate that CIE-Det achieves 73.1% mAP@0.5 (mean average precision at an intersection over union threshold of 0.5), outperforming YOLOv11s by 2.5%, while reducing model complexity to 9.12 M parameters. The proposed method achieves a favorable trade-off between detection accuracy and efficiency in low-light scenarios.
Handwritten mathematical expression recognition (HMER) aims to convert images of handwritten mathematical expressions into structured LaTeX sequences. Although recent encoder-decoder models have achieved strong performance, they still struggle with visually similar symbols, Greek letters, special symbols, and structural tokens such as fractions, radicals, superscripts, and subscripts, especially when expressions contain complex spatial layouts. To address these issues, we propose a Counting-based Position-aware Transformer (CP-Former) that integrates symbol-counting supervision and position-aware structural modeling into a CoMER-based recognition framework. Instead of treating counting and position information as isolated cues, CP-Former jointly learns global symbol occurrence statistics and relative structural positions from LaTeX annotations, without requiring additional symbol-level labels. Experiments on CROHME 2014, 2016, and 2019 show that CP-Former consistently improves over the CoMER baseline, with absolute ExpRate gains of 3.00%, 0.96%, and 2.00%, respectively. Compared with recent state-of-the-art methods such as PosFormer, CP-Former achieves competitive exact-match performance and obtains comparable or better results on several error-tolerant metrics, particularly under the ≤2 and ≤3 settings. These results suggest that the proposed counting-based position-aware modeling improves structural robustness for complex handwritten expressions, whereas the fair ablation study further shows that the gains cannot be attributed to counting supervision alone, but mainly arise from its integration with position-aware structural modeling.
Due to complex underwater illumination and scattering effects, underwater images commonly suffer from color distortion, low contrast, and structural blur. These distortions not only reduce visual perceptual quality but also severely limit the performance of downstream object detection tasks. Although existing underwater image enhancement methods have made notable progress in improving human visual perception, their optimization objectives mainly focus on subjective visual quality while overlooking detection-friendly feature representations. This results in a distributional mismatch with detector features, ultimately degrading object detection performance. To address these issues, we propose a global-local feature collaborative underwater image enhancement network. The network adopts a three-branch collaborative architecture to model local structural details, cross-color channel dependencies, and global contextual information, enabling joint optimization of global consistency and local discriminability. It provides more stable and discriminative feature representations for downstream object detection. To enhance the perception of small-scale targets and fine-grained structures, we design a structural detail enhancement module (SDEM) that combines multibranch and dilated convolutions to capture edge and texture features at different scales. The detail recalibration attention module (DRAM) combines local feature refinement with channel-spatial attention to adaptively reweight multibranch features, enabling fine-grained fusion of heterogeneous enhanced features and reinforcing key structural cues in the enhanced images. The experimental results show that our method significantly improves detection accuracy on the DUO and URPC2020 datasets. It also achieves strong PSNR and SSIM performance on the UIEB dataset. These results demonstrate that the proposed method can balance the detection performance and the visual quality.
Underwater three-dimensional reconstruction is inherently challenged by light scattering and wavelength-dependent absorption, causing color distortion, contrast degradation, and limited visibility in recovered geometries. Existing underwater three-dimensional (3D) Gaussian splatting methods rely solely on pixel-wise losses [L1/structural similarity index measure (SSIM)], which produce over-smoothed textures in scattering environments, while their direction-only medium modeling assumes spatially homogeneous water properties that contradict real-world depth-dependent turbidity. To address these limitations, we present an enhanced framework for underwater 3D reconstruction integrating progressive perceptual supervision with spatial coordinate encoding. A progressively weighted learned perceptual image patch similarity (LPIPS) loss recovers fine texture details without compromising geometric accuracy, while sinusoidal pixel-coordinate encoding enables self-supervised modeling of spatially varying turbidity without depth sensors. Experimental evaluation on the SeaThru-NeRF dataset demonstrates substantial perceptual quality improvements, with the progressive LPIPS alone achieving up to 33.9% LPIPS reduction on high-turbidity scenes, while the full framework maintains photometric accuracy (peak signal-to-noise ratio/SSIM). Cross-dataset evaluation on the out-of-distribution, deep-sea Submerged3D dataset further confirms that these perceptual gains generalize beyond a single benchmark. The method achieves interactive rendering performance (15 to 35 frames per second) and efficient training (∼8 to 13 min), providing a practical solution for real-time underwater 3D reconstruction in robotic applications.
Occlusions and far distance pedestrians pose significant challenges for pedestrian detection, often leading to insufficient feature representation learned by models, which in turn results in degraded detection accuracy and a high miss rate. To address this issue, we propose a three-stage fusion multispectral pedestrian detection network named TFNet. The network employs a three-stage fusion strategy: first, the multiscale feature fusion attention module selectively enhances critical features within each modality, effectively focusing on discriminative regions of occluded and distant pedestrians. Subsequently, the transformer interactive fusion module establishes long-range dependencies and enables dynamic interaction across modalities, achieving deep semantic alignment and complementary information exchange. Finally, the pixel-adaptive feature fusion (PAFF) module performs pixel-level refinement and adaptive fusion of the interacted features, generating a more discriminative unified representation. Extensive experiments conducted on the public multispectral pedestrian detection KAIST dataset show that TFNet outperforms existing state-of-the-art algorithms. In particular, it achieves a steady and substantial reduction in miss rate on the occlusion subset and multiscale subset. Meanwhile, it obtains higher average precision on the FLIR and LLVIP datasets. All experimental results fully demonstrate that the proposed network can effectively alleviate missed detections of occluded and far-distance pedestrians, and is expected to present high application value in the field of pedestrian safety monitoring in public scenarios.
Depression detection focuses on individuals' long-term and stable states, whereas emotional expression is typically a short-term, context-sensitive transient response. Direct fusion of emotional features with depressive features tends to misclassify "momentary emotional fluctuations" as "long-term depressive signals," thereby leading to false detection and poor generalization. To address this issue, we propose a lightweight uncertainty-aware fusion framework, namely, A2F-Net. Our core idea is straightforward: first, a depression branch generates a reference judgment indicating whether the current sample is more likely depressive or nondepressive. Subsequently, emotional cues are examined moment-by-moment to verify their consistency with this reference. For segments where emotional evidence is significantly inconsistent with the reference, their weights are reduced to retain only emotional information that is more conducive to depression assessment. Finally, one-way residual learning is adopted to supplement the filtered emotional information into the depression representation, preventing short-term emotional noise from interfering with the long-term depression representation. Experiments on LMVD, D-vlog, and DAIC-WOZ demonstrate that A2F-Net consistently improves depression detection performance, delivering consistent improvements over representative methods (up to 1.39% in F1-score) and exhibiting stronger robustness under intense emotional fluctuations.
Multiview stereo (MVS) aims to recover a scene's 3D geometry from multiple calibrated images by establishing reliable cross-view correspondences. Existing deep-learning-based MVS methods construct cost volumes and perform multiscale fusion, but often suffer from insufficient information integration that introduces noise and mismatches, limiting reconstruction accuracy and robustness. To address these issues, we propose an Attention-Driven Cost Volume Enhancement method for MVS (AVE-MVS) to improve reconstruction quality and efficiency. Concretely, we (1) design a contextual cost-volume attention module that leverages context-aware mechanisms to extract more discriminative visual features and to filter and strengthen cost-volume signals during construction, effectively suppressing redundant information; (2) introduce a cross-scale attention fusion module that uses attention weights to select salient regions and guide multiscale cost-volume fusion, thereby enhancing discriminative features across scales; and (3) adopt a gated aggregation unit to iteratively update fused depth information together with geometric priors, avoiding expensive 3D convolutions while preserving performance. Experiments on DTU and Tanks and Temples demonstrate that AVE-MVS outperforms several state-of-the-art methods in reconstruction accuracy, efficiency, and memory usage, showing good practicality and scalability.