Infrared image generation is essential in scenarios with low illumination or complex environments, but scarcity of aligned visible-infrared data and the lack of physical realism in generated results remain challenges. However, existing generative models often overlook the thermodynamic principles underlying infrared imaging, resulting in synthetic images that are visually plausible yet physically inaccurate. In paper, we propose Infrared Physics-guided Latent Diffusion (IPLD), a novel framework that integrates physics guided modeling into a latent diffusion process for high-fidelity synthesis of infrared images. Central IPLD is the Temperature-Emissivity-Environmental Radiance (TeR) decomposition, which decomposes thermal signals into temperature, emissivity, and environmental radiance components, governed by the laws blackbody radiation. To enhance the environmental radiance modeling, we introduce Environmental Radiance Map Estimation (ERME), a hybrid local-global estimation mechanism that preserves both spatial detail thermal consistency. Furthermore, a novel Skip Connection Diffusion Transformer (SCDT) is proposed strengthen and balance semantic structure and fine-grained details during image reconstruction. Extensive experiments on public datasets demonstrate that IPLD outperforms state-of-the-art Generative Adversarial Network (GAN)-based and diffusion-based methods, achieving superior results in Structural Similarity Measure (SSIM), Peak Signal-to-Noise Ratio (PSNR), Learned Perceptual Image Patch Similarity (LPIPS), Fr & eacute;chet Inception Distance (FID) metrics. Ablation studies validate the complementary value of TeR, ERME, and SCDT in improving radiative realism. Our approach establishes a new paradigm for physically grounded image translation, offering enhanced generalization and reliability for downstream perception tasks such target detection.
In this paper, a model is developed to investigate the dynamic response of electromagnetic rail launcher rails subjected to a moving magnetic pressure. When the armature velocity reaches the critical velocity, destructive resonant regimes occur in the rails, significantly affecting launch accuracy and rail service life. The dynamic response of the rails is modeled using the Bernoulli–Euler beam theory, treating the rails as beams resting on an elastic foundation and subjected to a moving load. An analytical expression for critical velocity is derived using the characteristic equation method, and its accuracy is validated through finite element simulations implemented in ANSYS. Subsequently, simulation analyses are conducted to examine how varying design parameters influence rail dynamic response and critical velocity. The results demonstrate that critical velocity is dependent on rail geometric and material parameters, as well as the stiffness of the elastic foundation. The discrepancy between the analytical solution and simulation results presented in this study is within 5
Digital screen exposure has become ubiquitous, yet its causal impact on self-regulation traits remains poorly understood due to methodological limitations in observational research. This study develops a methodological innovation by integrating Mendelian randomization (MR) with advanced multivariable modeling and machine-learning instrument selection. Unlike conventional MR applications primarily in biomedical research, our framework extends genetically informed causal inference into computational social systems. Specifically, we propose a multivariable MR approach enhanced by least absolute shrinkage and selection operator (LASSO)-based instrumental variable selection, which disentangles the independent causal effects of multiple, genetically correlated digital exposures. Using genome-wide association summary data from over 450 000 individuals from the UK Biobank, we investigate the behavioral effects of mobile phone use, computer use, computer gaming, and television viewing on self-regulation traits. The results show that mobile phone and computer use increase risk-taking tendencies, television viewing is associated with lower risk-taking scores and higher perseverance scores, possibly reflecting the passive and sequential nature of this activity, while computer use and gaming are linked to diminished endurance. These findings not only uncover causal pathways between digital activity patterns and psychological traits but also demonstrate a robust methodological advance for applying MR in complex behavioral systems. By demonstrating how screen exposure causally shapes core personality traits, this study provides actionable insights for technology designers to develop personality-aware digital interfaces, for policymakers to establish evidence-based screen time guidelines differentiated by media type, and for parents and educators to implement personalized digital engagement strategies that account for individual differences in self-regulation capacity.
UAV detection in surveillance is difficult due to small target size, cluttered backgrounds, and inter-class confusion with birds. We propose CPCR-DEIM, built on the DEIM framework, with two lightweight modules. First, a Coordinate Pyramid Channel-guided Recalibration (CPCR) block is inserted at an intermediate backbone stage to improve target-background separability via multi-scale directional reweighting before further spatial downsampling. Second, a Feature Fusion Enhancement Pyramid Network (FFEPN) replaces standard neck fusion nodes with FFE operators that use deep-layer features to compute per-channel weights, reducing the influence of clutter-dominated channels during top-down aggregation. Experiments on Air-UAV (visible light) and Anti-UAV (infrared) show that CPCR-DEIM achieves $80.4\pm 0.2\%$ $\text{mAP}_{50}$ and $44.9\pm 0.2\%$ $\text{mAP}_{50:95}$ at 25.1 GFLOPs and 10.2 M parameters, with 5.99 ms latency (167.07 FPS) on the measured GPU platform. Among the compared YOLO-family and transformer-based detectors, CPCR-DEIM gives the best strict-IoU localization accuracy on Air-UAV and the best accuracy on Anti-UAV.
Object detection in real-world environments remains challenging due to low illumination, cluttered backgrounds, and modality-specific noise. Dual-modality fusion methods such as RGB-Infrared (RGB-IR) and RGB-Polarization (RGB-Pol) provide complementary information, but existing solutions often suffer from high computational cost, limited scalability, and modality interference. To address these issues, we propose a unified and efficient dual-modality detection framework built around a Dual-Modality Fusion Module (DMFM), which performs lightweight early alignment followed by deep symmetric fusion. DMFM integrates a Swapping operation to realign shallow features. It then applies an Efficient Symmetric State-space Modality Fusion (ESSMF) unit powered by the proposed Improved 2D Selective Scan (IS2D). IS2D models long-range dependencies via sparse directional scanning with linear complexity. In addition, a Mix attention module refines multi-scale fused features by emphasizing informative spatial and channel cues. The framework supports both RGB-IR and RGB-Pol input streams without any structural modification. Extensive experiments on four public benchmarks demonstrate that our method achieves state-of-the-art performance across all datasets, with consistent gains in detection accuracy, precision, and recall. Notably, it outperforms existing mainstream methods while maintaining real-time speed. These improvements hold across varying modalities, object types, and environmental conditions, including low light, occlusion, and small targets. Overall, the proposed method offers a scalable, interference-resistant solution for real-time dual-modality detection and provides new insights into the design of efficient visual state-space fusion.
Industrial automation has accelerated the adoption of multi-permanent-magnet synchronous motor (PMSM) drive systems in large servo presses and gantry-stage platforms, where high dynamic speed tracking and precise synchronization are both required. This paper proposes a unified-model-based, non-cascaded predictive speed-synchronization framework for a three-PMSM system. A multivariable unified model is established to capture inter-motor coupling, and a compact finite-control-set model predictive controller (FCS-MPC) is developed by embedding both speed-tracking and synchronization errors into a single cost function. Unlike conventional multi-motor PMSM schemes based on parallel combinations of single-motor models, the proposed method coordinates current regulation, speed tracking, and synchronization control within one predictive framework. A Luenberger load-torque observer is further introduced for online disturbance estimation and compensation, improving robustness under differential-load conditions. Simulations and experiments show that, compared with deviation-coupling control (DCC), the proposed method reduces the peak speed-tracking error from 110.25 rpm to 53.92 rpm, shortens the recovery time from 1.46 s to 0.82 s, and decreases the peak synchronization error from 38.84 rpm to 31.32 rpm. It also achieves the lowest tracking-error IAE/ITAE (39.90 rpm & centerdot;s, 211.47 rpm & centerdot;s & sup2;) and synchronization-error IAE/ITAE (12.47 rpm & centerdot;s, 66.83 rpm & centerdot;s & sup2;), demonstrating superior dynamic recovery and disturbance rejection.
Lung cancer, ranking among the most lethal malignancies globally, poses a critical challenge in early diagnosis due to the difficulty of precisely detecting pulmonary nodules—its key early imaging feature. In CT imaging, challenges such as small target size, low contrast, and poor discriminability from surrounding tissues significantly hinder accurate nodule detection. To address these issues, this paper proposes an Improved Scaled YOLOv11n (ISE-YOLO) detection framework for pulmonary nodule localization. Diverging from standard YOLOv11 architectures, ISE-YOLO integrates three core enhancements: an Iterative Attention Fusion Module (iAFF) to enhance feature discrimination, a Slim-neck architecture to streamline feature propagation, and an Efficient Up-Convolution Block (EUCB) to optimize scale-aware feature recovery. These modifications collectively strengthen the model's representation capability, boost detection accuracy, and reduce computational redundancy. Experimental evaluations on a benchmark dataset demonstrate that ISE-YOLO achieves significant performance gains over baseline YOLOv11, with mAP@0.5 and mAP@(0.5:0.95) improving by 4.7% and 2.9%, respectively. The results validate that ISE-YOLO is better tailored for the specific demands of pulmonary nodule detection in CT imagery.
Detection of small infrared targets holds significant importance in fields such as emergency rescue operations and autonomous driving. However, the low contrast and high noise sensitivity nature of infrared images often lead to poor detection performance. To address these limitations, this paper proposes ISTD-DETR, a novel approach that integrates super-resolution preprocessing with an improved Real-Time Detection Transformer (RT-DETR) model. First, this approach significantly enhances the detail information of infrared images by introducing EDSR super-resolution techniques and extensive image augmentation methods. Second, the backbone of the ISTD-DETR model is designed using an Efficient Multi-Scale Attention (EMA) structure and state space model, improving the model’s ability for feature extraction and long-range dependency modeling. Third, a new S2 feature layer and micro-target detection EncoderHead are added to the model to achieve better feature fusion and target localization, making ISTD-DETR well-suited for infrared small target detection. Additionally, the method incorporates an SPD-EMA module into the feature fusion segment, further enhancing the model’s ability to detect small-scale and low-resolution targets. Experimental results on the public datasets demonstrate that ISTD-DETR significantly outperforms existing methods, reducing false positives and false negatives while maintaining strong real-time performance.
In this study, an adaptive disturbance rejection control strategy with a dynamic performance constraint mechanism is proposed to address the risks, including system singularity caused by random perturbations and control instability caused by unstructured perturbations in multi-rotor probing and fighting vehicles. While retaining the high cost-effectiveness ratio, the vehicle innovatively enables the dynamic adjustment of the control boundary and the synchronous compensation of compound disturbances. This innovation is realized by constructing a synergistic architecture of the time-varying prescribed performance function and the disturbance observer. Compared with the traditional prescribed performance control method, this strategy effectively enhances the disturbance suppression capability of the system under the premise of ensuring the transient response.
Accurate detection of small object clusters in modern security surveillance systems is critical for threat prevention and response, directly impacting monitoring efficiency and risk management efficacy. However, prevailing object detection algorithms struggle with high-density small target scenarios, suffering from slow inference speeds, insufficient precision, frequent false positives, and high miss rates. These limitations undermine real-time reliability and responsiveness to emergencies. To address these challenges, we propose SIEP-YOLO, an enhanced YOLOv11-based model integrating three novel components: the SDI-iAFF feature fusion module, EUCB upsampling module, and a specialized small object detection layer P2. This architecture boosts representational capacity and detection performance while simplifying computational complexity, achieving a balance between high accuracy and lightweight design. Experimental results demonstrate that SIEP-YOLO outperforms the original YOLOv11 by 3.2% and 2.0% in mAP@0.5 and mAP@(0.5:0.95), respectively, on benchmark datasets. The proposed model thus emerges as a superior solution for small object cluster detection, enabling more efficient and reliable surveillance in complex environments.
Military aircraft target detection remains a critical challenge in modern defense systems, particularly for remote sensing imagery with complex environmental interference. Existing approaches often exhibit limitations in maintaining high detection fidelity across varying resolutions and cluttered backgrounds. To overcome these constraints, this study presents SR-YOLOv11, a multi-stage framework integrating Super-Resolution (SR) reconstruction and hierarchical feature optimization. Initially, the Enhanced Deep Super-Resolution (EDSR) network is deployed to refine input image quality, ensuring precise preservation of critical aircraft morphological features. Subsequently, the YOLOv11 architecture is systematically enhanced through three key innovations: 1) Replacement of the native C3k2 module with a C3k2_AdditiveBlock to amplify discriminative feature learning; 2) Integration of an Adaptive Downsampling (ADown) layer for computationally efficient multi-scale context aggregation; 3) Implementation of an Auxiliary detection (AuxDetect) head mechanism with cross-layer feature fusion, significantly boosting localization accuracy for occluded targets. Comprehensive evaluations on the MAR20 dataset demonstrate the framework’s superiority, achieving 98.7% mAP@50 and 80.9% mAP@(50:95). The proposed architecture demonstrates enhanced robustness in complex environments while maintaining real-time processing efficiency, validating its operational viability in aerial surveillance scenarios.
A key component of solar photovoltaic (PV) modules, the frame plays a crucial role in panel installation and service life. This study investigates the performance degradation of floating PV frames in harsh offshore environments. Environmental monitoring systems with optimized sensor networks were deployed to record temperature, humidity, salt spray, and corrosion levels. Failure modes were analyzed using discipline-specific equations and margin modeling techniques. Corrosion and environmental data collected under marine conditions were processed to establish a performance degradation model, with systematic identification of uncertainty sources. An updated margin framework was employed to enable reliability quantification and performance prediction under combined temperature, humidity, and salt spray conditions. The proposed methodology provides a foundation for enhancing the environmental adaptability and operational reliability of floating PV systems and offers theoretical guidance for the sus tainable development of offshore PV installations.
To address the challenges of low resolution, substantial scale variation, sample imbalance, and small object detection in UAV aerial imagery, this paper proposes an enhanced YOLOv11-nano (YOLOv11) detection algorithm, termed YOLOv11n-SPS. The SPD-Conv module is integrated into the backbone network to replace the conventional strided convolution and pooling layers found in traditional CNN architectures. This substitution mitigates the loss of fine-grained features typically caused by standard downsampling operations, thereby preserving critical information in low-resolution inputs. In the detection head, an improved C3k2_PPA module is introduced. This module employs a multi-branch strategy to capture object features at multiple scales, and incorporates an attention mechanism to adaptively enhance relevant features, emphasizing critical details of small targets. To address sample imbalance, the Slide Loss function is adopted as the loss function. By assigning higher weights to difficult samples, the model is guided to focus more on challenging instances during training, enhancing overall learning effectiveness. Experimental results demonstrate that, compared with the baseline model, YOLOv11n-SPS achieves improvements of 2.5% in precision, 2.8% in recall, and 2.6% in mean average precision (mAP) on the VisDrone2019 dataset, indicating strong detection performance and robustness in complex aerial scenarios.
In target state estimation using sensor networks, conventional methods often face difficulties. These include nonlinear target dynamics, high communication overhead, and weak robustness against disturbances. This paper introduces a distributed estimation framework with heterogeneous filters. It combines the strengths of different filters to suit various operating conditions. The framework builds a multi-layered cluster of nodes. These nodes include standard Kalman filters (KF), extended Kalman filters (EKF), and particle filters (PF). Within each homogeneous filter group, the system applies a covariance-based fusion strategy. Between heterogeneous groups, it uses an adaptive fusion method driven by residuals. To reduce communication load, the framework avoids transmitting raw data. It also introduces a node activation mechanism based on trajectory nonlinearity. This mechanism decides which filter groups should run at each time. Simulation results show clear improvements. The method enhances accuracy and robustness. It also lowers computation and communication costs. These results show promise for real-world use and large-scale deployment.
Short-wave irregularities induced by locomotive impacts and periodic loads during railway transportation threaten operational safety. Detection methods based solely on vibration signals are often compromised by noise, limiting their ability to accurately characterize the geometry and severity of such irregularities. This study introduces a detection approach integrating rail profile data and axle box vibration signals. Rail profiles, acquired via high-precision two-dimensional point clouds, enable detailed geometric analysis, while vibration data provide dynamic assessments. Feature extraction and dimensionality reduction are performed using a stacked auto-encoder for vibration signals, and Principal Component Analysis (PCA) is applied to profile data following Dynamic Time Warping (DTW)-based wear sequence generation. The proposed model fuses both modalities at the decision level. Experimental validation demonstrates that this integrated approach significantly enhances the accuracy of short-wave irregularity detection and severity assessment, surpassing methods utilizing a single data source, and thereby contributes to improved railway safety.
In urban safety, intelligent transportation, and smart security applications, robust pedestrian detection is paramount. Methods that rely solely on visible light imaging struggle in low-light or adverse weather conditions. To address these challenges, we propose dual-modality pedestrian detection (DMPD)—a novel dual-modality pedestrian detection framework that fuses visible and infrared imaging through innovative fusion strategies. The method integrates a modal alignment module to reduce pixel-level misalignment, a differential modal fusion module to effectively combine complementary features while suppressing noise, and a mix module that enhances multiscale feature extraction via integrated convolution and self-attention mechanisms. Furthermore, the enhanced YOLOv7 is used to further boost feature representation and detection accuracy. Experimental results on the public dataset demonstrate that DMPD achieves a detection $mA{{P}_{50}}$ of 97.1% and a real-time speed of 118 FPS, outperforming state-of-the-art methods under both normal and adverse conditions, including fog, rain, and snow. These results confirm the effectiveness of the proposed fusion strategy in harnessing the complementary strengths of visible and infrared modalities, thereby offering a highly robust and scalable solution for pedestrian detection in complex urban environments.
Unmanned Aerial Vehicle (UAV) detection is a key technology for Anti-UAV systems, which has high requirements for accuracy and real-time performance. The existing detection algorithms based on You Only Look Once (YOLO) struggle to strike a balance between detection accuracy and speed, making them unsuitable for practical Anti-UAV systems. To address these problems, a UAV detection algorithm based on the cross-weighted pixel reconstruction and feature selection network is proposed. First, a novel feature reconstruction module for parallel processing of multi-dimensional information is proposed, namely Cross-weighted Pixel Reconstruction module. This module designs a spatial segmentation strategy, through which the pixels of feature maps are reconstructed from three dimensions. The purpose is to highlight the object details and eliminate the interference from background and redundant information. Second, an innovative feature selection mechanism guided by semantic information is developed. This is the adjacent feature selection mechanism, which makes full use of the semantic information contained in the high-level feature maps to guide the screening of important information in the low-level feature maps. Furthermore, based on the adjacent feature selection mechanism, an improved feature fusion method, feature selection fusion, is constructed. This method can exploit semantic information and multi-scale information, enhancing the network's detection ability of small objects and similar objects. The experimental results show that the proposed algorithm has better comprehensive performance than YOLOv5, YOLOv8, YOLOv9 and YOLOv10 in UAV detection. The video of the UAV detection results based on the proposed method have been uploaded to https://github.com/mengmeng1023/CFS-YOLO.
Short-wavelength rail irregularities (SWRIs) are localized geometric deviations that evolve under periodic wheel loads and can compromise operational safety. Vibration-only inspection is prone to noise and response ambiguity, whereas profile-only inspection lacks temporal dynamics of wheel--rail interaction. We propose a multi-modal framework that fuses axle-box vibration sequences with high-resolution rail profile data. On the vibration side, features are learned via stacked autoencoders (SAE); on the profile side, measured contours are registered to a standard template, transformed into wear sequences via Dynamic Time Warping (DTW), and compressed with Principal Component Analysis (PCA). Two kernel SVMs (KSVMs) trained on the respective modalities are aggregated through a class-dependent confidence fusion. On a 100~m laboratory rail with simulated grinding (F1), spalling (F2), abrasion (F3), and normal (F4), we collect 200 paired samples and show that the fusion model markedly reduces misclassification, achieving accuracies of 87.5\% (F1), 100\% (F2), 93.3\% (F3), and 94.1\% (F4), with overall gains ranging from 7.1\% to 47.7\% over single-source baselines. The approach provides a principled path to robust, intelligent rail condition monitoring.
In the field of foreign object detection on transmission lines, advanced detection technologies are crucial. Existing methods have limitations in detecting multiple types of foreign objects, leading to a decrease in detection accuracy. To address these challenges, this paper introduces a new attention-augmented detection head method for foreign object detection, which uses the YOLOv8 model and incorporates Mixed Local Channel Attention (MLCA) and Auxiliary Detect Head (DetectAux) modules. These modules efficiently utilize image data to perform multi-object detection tasks, thereby improving detection performance under various environmental conditions. An integrated RepBlock module further enhances feature extraction. The proposed solution significantly improves the accuracy of foreign object detection in different scenarios, providing a more robust and adaptable approach to transmission line foreign object detection. The experimental results demonstrate the superiority of the proposed method in this paper for high-precision and highspeed recognition under complex background conditions.
Baoming Li (栗保明)合作论文数Nanjing University of Science & Technology4