Infrared imaging technology demonstrates significant advantages in aviation safety monitoring due to its exceptional all-weather operational capability and anti-interference characteristics, particularly in scenarios requiring real-time detection of aerial objects such as airport airspace management. However, traditional infrared target detection algorithms face critical challenges in complex sky backgrounds, including low signal-to-noise ratio (SNR), small target dimensions, and strong background clutter, leading to insufficient detection accuracy and reliability. To address these issues, this paper proposes the AFK-YOLO model based on the YOLO11 framework: it integrates an ADown downsampling module, which utilizes a dual-branch strategy combining average pooling and max pooling to effectively minimize feature information loss during spatial resolution reduction; introduces the KernelWarehouse dynamic convolution approach, which adopts kernel partitioning and a contrastive attention-based cross-layer shared kernel repository to address the challenge of linear parameter growth in conventional dynamic convolution methods; and establishes a feature decoupling pyramid network (FDPN) that replaces static feature pyramids with a dynamic multi-scale fusion architecture, utilizing parallel multi-granularity convolutions and an EMA attention mechanism to achieve adaptive feature enhancement. Experiments demonstrate that the AFK-YOLO model achieves 78.6% mAP on a self-constructed aerial infrared dataset—a 2.4 percentage point improvement over the baseline YOLO11—while meeting real-time requirements for aviation safety monitoring (416.7 FPS), reducing parameters by 6.9%, and compressing weight size by 21.8%. The results demonstrate the effectiveness of dynamic optimization methods in improving the accuracy and robustness of infrared target detection under complex aerial environments, thereby providing reliable technical support for the prevention of mid-air collisions.
Semi-supervised anomaly detection (SSAD) methods have demonstrated their effectiveness in enhancing unsupervised anomaly detection (UAD) by leveraging few-shot but instructive abnormal instances. However, the dominance of homogeneous normal data over anomalies biases the SSAD models against effectively perceiving anomalies. To address this issue and achieve balanced supervision between heavily imbalanced normal and abnormal data, we develop a novel framework called AnoOnly (Anomaly Only). Unlike existing SSAD methods that resort to strict loss supervision, AnoOnly suspends it for normal data and introduces a form of weak supervision. This weak supervision is instantiated through the utilization of batch normalization, which implicitly performs cluster learning on normal data. When integrated into existing SSAD methods, the proposed AnoOnly demonstrates remarkable performance enhancements across various models and datasets, achieving new state-of-the-art performance. Additionally, our AnoOnly is natively robust to label noise when suffering from data contamination. Our code is publicly available at https://github.com/cool-xuan/AnoOnly.
Microstructure, room-temperature mechanical properties, and oxidation resistance of Mo-12Si-8B-xTi (x = 10, 20, 30, 40, at%) alloys fabricated via hot pressing sintering are investigated. 10Ti and 20Ti alloys comprise alpha-Mo + Mo3Si + Mo5SiB2, whereas 30Ti and 40Ti alloys comprise alpha-Mo + Mo5SiB2 + Ti5Si3. CALPHAD thermodynamic modeling and EDS indicate a significant increase in Ti solubility in alpha-Mo and Mo5SiB2 phases with higher Ti content, reaching levels of 30-50 %. EBSD analysis reveals evident grain coarsening, the average grain size improves from 1.43 to 2.97 mu m. HRTEM characterization demonstrates alpha-Mo/Mo5SiB2 and alpha-Mo/Ti5Si3 interfaces are both incoherent. 10Ti alloy exhibits the highest flexural strength of 454 MPa and fracture toughness of 9.9 MPa m1/2, while minor differences in fracture toughness are observed among 20-40Ti alloys. The combination of adequate alpha-Mo matrix, appropriate grain size, and brittle intermetallic phases collectively contribute to determining the strength and toughness. Simultaneous thermal analysis and cyclic oxidation show 40Ti alloy offers the best oxidation resistance at 800-1200 degrees C, because of the protective TiO2 SiO2 layer generated through selective oxidation of Ti5Si3 and Ti/Si. Using Pilling-Bedworth ratio as the sole criterion to assess the oxidation resistance is inadequate, as temperature and phase constitution must be considered.
Weakly supervised video anomaly detection (WVAD) aims to detect where abnormal events occur in videos using only video-level labels during training. Many existing WVAD methods concentrate on learning comprehensive representations for each frame, making them susceptible to interference from irrelevant backgrounds or scenes. In this paper, we address this challenge by introducing disentangled representation learning to WVAD, presenting a new modular component named the Event-Centric Disentangler (ECD). Specifically, our ECD incorporates an Event Focus Attention (EFA) module, estimating channel-wise attention to focus on event representations while filtering out irrelevant information from backgrounds or scenes. Alongside EFA, a Frame Importance Allocator (FIA) learns frame-wise weighting factors to aggregate frame-level predictions for the generation of video-level predictions. Furthermore, we introduce a video-level contrastive loss to provide disentanglement supervision for ECD training, within the constraints of weakly supervised settings. Integrating ECD into existing WVAD methods, we achieve state-of-the-art performance on two benchmark datasets, with 87.53 https://github.com/quyiii/ECD .
The direct utilization of low-light images hinders downstream visual tasks. Traditional low-light image enhancement (LLIE) methods, such as Retinex-based networks, require image pairs. A spiking-coding methodology called intensity-to-latency has been used to gradually acquire the structural characteristics of an image. convLSTM has been used to connect the features. This study introduces a simplified DCENet to achieve unsupervised LLIE as well as the spiking coding mode of a spiking neural network. It also applies the comprehensive coding features of convLSTM to improve the subjective and objective effects of LLIE. In the ablation experiment for the proposed structure, the convLSTM structure was replaced by a convolutional neural network, and the classical CBAM attention was introduced for comparison. Five objective evaluation metrics were compared with nine LLIE methods that currently exhibit strong comprehensive performance, with PSNR, SSIM, MSE, UQI, and VIFP exceeding the second place at 4.4% (0.8%), 3.9% (17.2%), 0% (15%), 0.1% (0.2%), and 4.3% (0.9%) on the LOL and SCIE datasets. Further experiments of the user study in five non-reference datasets were conducted to subjectively evaluate the effects depicted in the images. These experiments verified the remarkable performance of the proposed method.
In weakly supervised video anomaly detection (WVAD), where only video-level labels indicating the presence or absence of abnormal events are available, the primary challenge arises from the inherent ambiguity in temporal annotations of abnormal occurrences. Inspired by the statistical insight that temporal features of abnormal events often exhibit outlier characteristics, we propose a novel method, BN-WVAD, which incorporates BatchNorm into WVAD. In the proposed BN-WVAD, we leverage the Divergence of Feature from the Mean vector (DFM) of BatchNorm as a reliable abnormality criterion to discern potential abnormal snippets in abnormal videos. The proposed DFM criterion is also discriminative for anomaly recognition and more resilient to label noise, serving as the additional anomaly score to amend the prediction of the anomaly classifier that is susceptible to noisy labels. Moreover, a batch-level selection strategy is devised to filter more abnormal snippets in videos where more abnormal events occur. The proposed BN-WVAD model demonstrates state-of-the-art performance on UCF-Crime with an AUC of 87.24%, and XD-Violence, where AP reaches up to 84.93%. Our code implementation is accessible at https://github.com/cool-xuan/BN-WVAD.
The microstructure, mechanical properties, and oxidation behavior of Mo-10Si-8B-xTiC (at.%, x = 0,5,10,15) alloys fabricated via oscillatory pressure sintering are studied. The phase constitution of three TiC-added alloys is alpha-Mo + Mo3Si + Mo5SiB2 + Mo2C + TiC; in 15TiC alloy alpha-Mo phase no longer maintains its contiguity as in the other three alloys. With further TiC addition, the microhardness improves by 60-90 % (up to 1500 HV), Young's elastic modulus increases from 328 to 368 GPa, and the fracture toughness decreases from 14.23 to 10.8 MPa m(1/2), accounting for the less-ductile phase and newly introduced brittle phases. Thermogravimetric analysis indicates that the addition of TiC reduces the oxidation resistance. For isothermal oxidation at 800, 1000, and 1200 degrees C, the serious oxidation occurs for all TiC-added alloys except for 15TiC alloy at 1200 degrees C because the oxidation sequence changes to TiC -> Mo5SiB2 -> Mo3Si -> alpha-Mo. The anomalous phenomenon for 15TiC alloy at 1200 degrees C could be attributed to the rapid formation and cover of protective SiO2-TiO2 layer on alloy surface after the transient MoO3 volatilization.
To further enhance the recognition accuracy of automatic modulation recognition, improve communication efficiency, strengthen security, and optimize resource management, this paper designs a high-precision hybrid deep learning model featuring early-stage layer fusion. This model combines with Convolutional Neural Networks (CNN), Transformers, and Deep Neural Networks (DNN) to enhance the model’s feature extraction capabilities, thereby improving modulation recognition accuracy. Experiments are performed on RadioML2016.10a and RadioML2018.01a, and the results show that this architecture can effectively combine the advantages of different types of models, making the overall performance more robust and suitable for complex automatic modulation recognition problems.
The fast and accurate detection of steel surface defects has become an important goal of research in various fields. As one of the most important and effective methods of detecting steel surface defects, the successive generations of YOLO algorithms have been widely used in these areas; however, for the detection of tiny targets, it still encounters difficulties. To solve this problem, the first modified PP-YOLOE algorithm for small targets is proposed. By introducing Coordinate Attention into the Backbone structure, we encode channel relationships and long-range dependencies using accurate positional information. This improves the performance and overall accuracy of small target detection while maintaining the model parameters. Additionally, simplifying the traditional PAN+FPN components into an optimized FPN feature pyramid structure allows the model to skip computationally expensive but less relevant processes for the steel surface defect dataset, effectively reducing the computational complexity of the model. The experimental results show that the overall average accuracy (mAP) of the improved PP-YOLOE algorithm is increased by 4.1%, the detection speed is increased by 2.06 FPS, and the accuracy of smaller targets (with a pixel area less than 322) that are more difficult to detect is significantly improved by 13.3% on average, as compared to the original algorithm. The detection performance is also higher than that of the mainstream target detection algorithms, such as SSD, YOLOv3, YOLOv4, and YOLOv5, and has a high application value in industrial detection.
Mo-Si-B alloys are a crucial focus for the development of the next generation of ultra-high-temperature structural materials. They have garnered significant attention over the past few decades due to their high melting point and superior strength and oxidation resistance compared to other refractory metal alloys. However, their low fracture toughness at room temperature and poor oxidation resistance at medium temperature are significant barriers limiting the processing and application of Mo-Si-B alloys. Therefore, this review was carried out to compare the effectiveness of doped metallic elements and second-phase particles in solving these problems in detail, in order to provide clear approaches to future research work on Mo-Si-B alloys. It was found that metal doping can enhance the properties of the alloys in several ways. However, their impact on oxidation resistance and fracture toughness at room temperature is limited. Apart from B-rich particles, which significantly improve the high-temperature oxidation resistance of the alloy, the doping of second-phase particles primarily enhances the mechanical properties of the alloys. Additionally, the application of additive manufacturing to Mo-Si-B alloys was discussed, with the observation of high crack density in the alloys prepared using this method. As a result, we suggest a future research direction and the preparation process of oscillatory sintering, which is expected to reduce the porosity of Mo-Si-B alloys, thereby addressing the noted issues.
Class imbalance is a common challenge in real-world recognition tasks, where the majority of classes have few samples, also known as tail classes. We address this challenge with the perspective of generalization and empirically find that the promising Sharpness-Aware Minimization (SAM) fails to address generalization issues under the class-imbalanced setting. Through investigating this specific type of task, we identify that its generalization bottleneck primarily lies in the severe overfitting for tail classes with limited training data. To overcome this bottleneck, we leverage class priors to restrict the generalization scope of the class-agnostic SAM and propose a class-aware smoothness optimization algorithm named Imbalanced-SAM (ImbSAM). With the guidance of class priors, our ImbSAM specifically improves generalization targeting tail classes. We also verify the efficacy of ImbSAM on two prototypical applications of class-imbalanced recognition: long-tailed classification and semi-supervised anomaly detection, where our ImbSAM demonstrates remarkable performance improvements for tail classes and anomaly. Our code implementation is available at https://github.com/cool-xuan/Imbalanced_SAM.
The aircraft engine is a core component of an airplane, and its critical components work in harsh environments, making it susceptible to a variety of surface defects. To achieve efficient and accurate defect detection, this paper establishes a dataset of surface defects on aircraft engine components and proposes an optimized object detection algorithm based on YOLOv5 according to the features of these defects. By adding a dual-path routing attention mechanism in the Biformer model, the detection accuracy is improved; by replacing the C3 module with C3-Faster based on the FasterNet network, robustness is enhanced, accuracy is maintained, and lightweight modeling is achieved. The NWD detection metric is introduced, and the normalized Gaussian Wasserstein distance is used to enhance the detection accuracy of small targets. The lightweight upsampling operator CARAFE is added to expand the model's receptive field, reorganize local information features, and enhance content awareness performance. The experimental results show that, compared with the original YOLOv5 model, the improved YOLOv5 model's overall average precision on the aircraft engine component surface defect dataset is improved by 10.6%, the parameter quantity is reduced by 11.7%, and the weight volume is reduced by 11.3%. The detection performance is higher than mainstream object detection algorithms such as SSD, RetinaNet, FCOS, YOLOv3, YOLOv4, and YOLOv7. Moreover, the detection performance on the public dataset (NEU-DET) has also been improved, providing a new method for the rapid defect detection of aircraft engines and having high application value in various practical detection scenarios.
Semi-supervised anomaly detection (SSAD) methods have demonstrated their effectiveness in enhancing unsupervised anomaly detection (UAD) by leveraging few-shot but instructive abnormal instances. However, the dominance of homogeneous normal data over anomalies biases the SSAD models against effectively perceiving anomalies. To address this issue and achieve balanced supervision between heavily imbalanced normal and abnormal data, we develop a novel framework called AnoOnly (Anomaly Only). Unlike existing SSAD methods that resort to strict loss supervision, AnoOnly suspends it and introduces a form of weak supervision for normal data. This weak supervision is instantiated through the utilization of batch normalization, which implicitly performs cluster learning on normal data. When integrated into existing SSAD methods, the proposed AnoOnly demonstrates remarkable performance enhancements across various models and datasets, achieving new state-of-the-art performance. Additionally, our AnoOnly is natively robust to label noise when suffering from data contamination. Our code is publicly available at https://github.com/cool-xuan/AnoOnly.
Semi-supervised anomaly detection (SSAD) methods have demonstrated their effectiveness in enhancing unsupervised anomaly detection (UAD) by leveraging few-shot but instructive abnormal instances. However, the dominance of homogeneous normal data over anomalies biases the SSAD models against effectively perceiving anomalies. To address this issue and achieve balanced supervision between heavily imbalanced normal and abnormal data, we develop a novel framework called AnoOnly ( Ano maly Only ). Unlike existing SSAD methods that resort to strict loss supervision, AnoOnly suspends it and introduces a form of weak supervision for normal data. This weak supervision is instantiated through the utilization of batch normalization, which implicitly performs cluster learning on normal data. When integrated into existing SSAD methods, the proposed AnoOnly demonstrates remarkable performance enhancements across various models and datasets, achieving new state-of-the-art performance. Additionally, our AnoOnly is natively robust to label noise when suffering from data contamination. Our code is publicly available at https://github.com/cool-xuan/AnoOnly .
快速、准确地检测材料表面缺陷已成为各领域研究的重要目标,为增加检测效率,实现设备轻量化,提出了一种基于YOLOv5的目标检测优化算法,添加DyHead检测头,融合多个注意力机制,增强模型的检测精度;更换aLRPLoss损失函数,减少超参数调节工作,优化训练过程;基于FasterNet提出C3-Faster,代替网络中的C3模块,以PConv的思想提升模型检测性能,减少模型体积;最后添加轻量级上采样算子CA-RAFE,扩大模型感受野,提升对不同大小目标的检测效果.实验结果表明,改进后的YOLOv5模型相比于原版模型,在钢材表面缺陷数据集上总体平均精度提高了4.174%,参数量减少了11.25%,计算复杂度减少了13.75%,权重体积减少了10.72%,检测性能高于SSD、RetinaNet、FCOS、YOLOv3、YOLOv4等主流目标检测算法,在工业检测中具有较高的应用价值.
The integration of artificial intelligence with steel manufacturing operations holds great potential for enhancing factory efficiency. Object detection algorithms, as a category within the field of artificial intelligence, have been widely adopted for steel defect detection purposes. However, mainstream object detection algorithms often exhibit a low detection accuracy and high false-negative rates when it comes to detecting small and subtle defects in steel materials. In order to enhance the production efficiency of steel factories, one approach could be the development of a novel object detection algorithm to improve the accuracy and speed of defect detection in these facilities. This paper proposes an improved algorithm based on the YOLOv5s-7.0 version, called YOLOv5s-7.0-FCC. YOLOv5s-7.0-FCC integrates the basic operator C3-Faster (C3F) into the C3 module. Its special T-shaped structure reduces the redundant calculation of channel features, increases the attention weight on the central content, and improves the algorithm’s computational speed and feature extraction capability. Furthermore, the spatial pyramid pooling-fast (SPPF) structure is replaced by the Content Augmentation Module (CAM), which enriches the image feature content with different convolution rates to simulate the way humans observe things, resulting in enhanced feature information transfer during the process. Lastly, the upsampling operator Content-Aware ReAssembly of Features (CARAFE) replaces the “nearest” method, transforming the receptive field size based on the difference in feature information. The three modules that act on feature information are distributed reasonably in YOLOv5s-7.0, reducing the loss of feature information during the convolution process. The results show that compared to the original YOLOv5 model, YOLOv5s-7.0-FCC increases the mean average precision (mAP) from 73.1% to 79.5%, achieving a 6.4% improvement. The detection speed also increased from 101.1 f/s to 109.4 f/s, an improvement of 8.3 f/s, further meeting the accuracy requirements for steel defect detection.
The Mo-12Si-8.5B alloy was surface-remelted by laser and electron beam, and the microstructure of its melt pool and substrate regions were analyzed by scanning electron microscopy (SEM), X-ray diffraction (XRD), and energy spectrometry (EDS) techniques. It was found that the composition of the surface phases in the Mo-12Si-8.5B alloy did not change by the high-energy beam surface remelting process, but the microstructure of the molten pool region was significantly different from that of the substrate region, and its phase distribution was more uniform. Dendrites appeared on the surface of the material under the action of both processes, and the Si- and B-rich phases were mainly gathered in the interdendritic region. In the melt pool of the laser-remelted specimens, the α-Mo phase was continuously distributed with an average dendrite length of 70 µm, while the α-Mo phase distribution in the melt pool of the electron beam remelted specimens were relatively concentrated, with a larger dendrite size and an average dendrite length of 120 µm. The dendrite size in the melt pool of the laser remelted material was smaller, and the distribution of the elements was relatively uniform. Using a laser beam as the heat source was more favorable for the next step of the additive manufacturing of the core parts of hypersonic vehicles.
Diffusion aluminum coating is crucial to protect aero-engine turbine blades from high-temperature oxidation. Slurry aluminizing, as a commonly-used coating preparation technology, has variations in the process parameters that directly affect the quality of the coating. Therefore, this paper investigates the effect of slurry thickness on coating quality. Different forms of aluminized coatings were obtained by coating nine DZ22B nickel-based superalloy plates of the same size with different slurry thicknesses while keeping other parameters constant. These aluminized coatings were characterized using a scanning electron microscope (SEM) with an energy dispersive spectrometer (EDS), an X-ray diffractometer (XRD), and a surface gauge. The results show that the AlNi phase dominates the matrix of the aluminized coating, and the outer layer of the coating has white dotted precipitates of Cr. As the slurry thickness increases, the coating thickness increases, and the proportion of the outer layer in the overall coating increases. In contrast, the thickness of the interdiffusion layer does not change significantly. The thicker the slurry, the higher the Al content of the coating surface. A medium-thickness slurry can form a smooth aluminizing coating with a roughness Ra < 4.5 μm surface. The combined results show that a medium-thick slurry can produce a high-quality coating.
An aluminized coating can improve the high-temperature oxidation resistance of turbine blades, but the inter-diffusion of elements renders the coating's thickness difficult to achieve in non-destructive testing. As a typical method for coating thickness inspection, X-ray fluorescence mainly includes the fundamental parameter method and the empirical coefficient method. The fundamental parameter method has low accuracy for such complex coatings, while it is difficult to provide sufficient reference samples for the empirical coefficient method. To achieve accurate non-destructive testing of aluminized coating thickness, we analyzed the coating system of aluminized blades, simulated the spectra of reference samples using the open-source software XMI-MSIM, established the mapping between elemental spectral intensity and coating thickness based on partial least squares and back-propagation neural networks, and validated the model with actual samples. The experimental results show that the model's prediction error based on the back-propagation neural network is 4.45% for the Al-rich layer and 16.89% for the Al-poor layer. Therefore, the model is more suitable for predicting aluminized coating thickness. Furthermore, the Monte Carlo simulation method can provide a new way of thinking for materials that have difficulty in fabricating reference samples.
In recent years, the application of object detection in the military field has become more and more extensive, and the detection of aircraft objects in remote sensing images can provide data support for accurate object strikes. In this paper, we propose an real-time aircraft object detection method in remote sensing images based on YOLOv4 object detection algorithm. We improved the YOLOv4 object detection algorithm by replacing the traditional convolution in Res_unit with a depthwise separable convolution, replacing the Mish activation function in the backbone with the ELU activation function, and adding an SE module in each CSP_unit. The obtained algorithm was named Aircraft-YOLOv4 finally. The mAP and fps of Aircraft-YOLOv4 when detecting aircraft objects in remote sensing images can reach 86.92% and 29.62, respectively, realizing real-time detection, which is 2.82% and 7.01 higher than YOLOv4. And Aircraft-YOLOv4 has improved performance in all aspects when model is tested on UCAS-AOD, a dataset similar to the RSOD-Dateset used for training. The experimental results show that Aircraft-YOLOv4 has good generalization and is more suitable for aircraft object detection tasks in remote sensing images in the military field than YOLOv4.