Timely detection and repair of defective power components in transmission lines are critical to the safe operation of power grids. With the widespread adoption of unmanned aerial vehicles (UAVs) for the inspection of transmission lines, effectively balancing the model’s resource consumption and detection accuracy remains a challenge. To address this issue, this paper proposes an enhanced YOLOv8n-based model for small-object and multi-target detection, achieving both lightweight design and improved detection performance. First, the lightweight MobileViT is integrated into the backbone's C2f module to create the C2f-ViT module. This design fuses multi-region features, mitigates computational consumption, and boosts feature representation. Second, the Efficient Multi-Scale Attention (EMA) mechanism is embedded into the neck’s C2f module. It reinforces the model's focus on vital regions amid complex interference and improves its target detection accuracy. Subsequently, the Dynamic Head (DyHead) module is integrated into the detection head, boosting multi-scale feature fusion capability through adaptive cross-level feature integration. Finally, model compression is achieved via a layer-adaptive magnitude-based pruning (LAMP) method, which reduces parameter count and memory usage while maximizing accuracy retention. Experimental results indicate that, compared with the standard YOLOv8n, the parameters of our proposed model are reduced by 74
To address the inherent limitations of current detection methods, namely weak feature representation for small objects, insufficient mining of critical details, and poor anti-interference performance, we propose a discriminative feature network for small object detection. The introduced discriminative feature perception, selection, and enhancement modules enable in-depth perception, effective filtering, and targeted enhancement of edge details and semantic information pertaining to small objects, thereby significantly strengthening the model's feature representation capability. Extensive experimental results demonstrate that the proposed method outperforms other comparative algorithms on three authoritative benchmarks: AI-TOD, AI-TODv2 and VisDrone2019. Furthermore, in the context of a ship water gauge reading intelligent recognition project, we establish a novel Small-Character dataset comprising 5051 images and 60878 instances, where our approach achieves robust detection performance in practical real-world scenarios.
This survey paper provides a comprehensive analysis of deep learning advancements in humanoid robot vision, focusing on key areas such as real-time object detection, 3D scene understanding, and multimodal sensor fusion. We systematically review foundational architectures (CNNs, Transformers) and emerging paradigms (self-supervised learning, embodied AI) that enable robust visual perception in dynamic environments. The paper highlights critical algorithmic innovations, including YOLO variants for efficient detection, spatial-aware transformers for 3D reasoning, and fusion techniques for heterogeneous sensor integration. We identify persistent challenges in computational efficiency, generalization under environmental dynamics, and ethical deployment considerations. The survey synthesizes interdisciplinary research directions, emphasizing neuro-symbolic integration, open-vocabulary perception, and energy-efficient architectures as pivotal for advancing autonomous human-robot interaction. By evaluating 60+ works, this study establishes a roadmap for overcoming current limitations while leveraging breakthroughs in visual representation learning and embodied cognition to push the boundaries of robotic vision systems.
Efficient and accurate detection of foreign object intrusion on transmission lines based on UAV aerial images and deep learning algorithms enables rapid identification of potential safety hazards, thereby improving the lean operation and maintenance capability of power grids. The research progress of deep learning-based foreign object intrusion detection methods for transmission lines using UAV aerial images was reviewed. First, deep learning-based foreign object detection methods for UAV aerial images were systematically reviewed from four dimensions: Faster R-CNN, SSD, the YOLO series algorithms, and other derivative algorithms. Second, an improved model for UAV aerial image-based foreign object intrusion detection in transmission lines was elaborated. It illustrated the application strategies and performance optimization effects of core technologies in detection tasks, including attention mechanisms, multi-scale feature fusion, and lightweight networks. Subsequently, the DETR-based foreign object detection models built on the vanilla Transformer architecture were investigated, as well as large Transformer models tailored for the power industry and their engineering applications. Finally, in view of the current challenges in UAV aerial image-based intelligent foreign object detection for transmission lines, future research prospects were proposed.
Cylindrical lithium-ion batteries (CLBs), as essential components in electric vehicles and portable electronic devices, have attracted considerable attention due to concerns about their safety and reliability. Accurate and efficient detection of various surface defects on these batteries continues to pose a significant challenge. To address this issue, this paper proposes a novel network architecture, which integrates Transformer and Mamba mechanisms. First, the model introduces a newly designed Mamba-based backbone network, whose core feature extraction module combines gated nested architectures with residual connections, effectively captures local dependencies, thereby enhancing the local modeling capability of State Space Models (SSM). Second, when dealing with high-level feature layers, conventional multi-head attention is divided into high-frequency and low-frequency branches, thereby fully leveraging the distinct feature extraction capabilities of each attention head. Furthermore, a novel neck fusion framework is proposed to incorporate shallow features from the backbone as auxiliary pathways into deeper layers, while removing the large object detection head. This architectural design alleviates the mismatch problem between the target scale distribution and the receptive field of detection heads. Experimental results indicate that proposed method achieves a mean average precision (mAP) across an intersection over union (IoU) threshold from 0.5 to 0.95 (mAP50-95) of 0.633 and a mAP at an IoU threshold of 0.5 (mAP50) of 0.944, surpassing state-of-the-art models in performance. Finally, through comparative experiments with publicly available printed circuit board (PCB) defect dataset and hypothesis testing under real industrial scenarios, the generalization ability and robustness of the proposed model are further verified.
Composite cBN- Y2O3 ceramics with a 30 vol% cBN content were successfully fabricated at sintering temperature ranging from 800 to 1600 °C under a high pressure of 7.2 GPa. The high sintering pressure was observed to prevent phase transformation of cubic BN phase, while simultaneously inducing the transformation from cubic Y2O3 to monoclinic Y2O3 phase. No significant change in the grain size of cBN within the composites was observed, attributed to the low sintering temperatures and short holding time. In contrast, the grain size of Y2O3 within the cBN- Y2O3 composite was reduced to less than 100 nm after sintering at 800 and 1000 °C. This observation indicates that the high pressure-induced phase transformation of Y2O3contributed to its grain refinement. As a result of the retained cubic BN phase and refined Y2O3 grains, the cBN- Y2O3 composite ceramics sintered for 60 s achieved a relative density of 94 % and a hardness of approximately 11 GPa. These results demonstrated that the high pressure and high temperature (HPHT) technique is a promising route in achieving fully dense cBN-Y2O3 composite ceramics with enhanced performance.
The recognition and positioning of characters on the water gauge are important components of artificial intelligence system for reading ship draft weighing. Meanwhile, the manual annotation of oriented objects has a large workload and low accuracy. To address these issues, this paper proposes a novel oriented detection framework. It introduces three additional parameters to refine the horizontally detected bounding boxes from the model and employs two symmetric functions to confine the angles and scales within specified ranges. In the training and testing of the model, only the horizontal annotation information of the object is needed, reducing the workload of annotation of the target in the dataset. A residual network has been added to the backbone of object detector(YOLO) to enhance the feature extraction capability of deep modules and improve the sensitivity to small objects. For more accurate oriented estimation, we utilize improved Intersection over Union(IoU) to the bounding box regression loss.This method is particularly effective for objects with a dominant orientation and approximately rectangular shapes, such as the characters, vehicles, and buildings commonly found in aerial imagery. A simple implementation of our method has achieved state-of-the-art performances on aerial objects datasets, with a negligible reduction to detection speed. After applying the novel oriented detection method to the intelligent system, real-time testing was conducted on 120 drone videos, resulting in a 35.3% improvement in the accuracy of the system’s water gauge readings.
Accurate waterline detection is critical for automated ship draft monitoring but remains challenging due to weak textures, low contrast, and dynamic maritime interferences. This paper presents TGNet, a task-guided framework that jointly optimizes character recognition and waterline keypoint localization. TGNet introduces a triple attention network (TAnet) with channel, spatial, and texture attention modules to enhance discriminative feature extraction. Crucially, a task-to-task guidance mechanism leverages detected draft characters to spatially constrain and crop feature maps, focusing the keypoint detection head on the most relevant waterline region. Extensive experiments on three large-scale aerial datasets show that TAnet consistently improves baseline detectors by an average of 2.8% recall and 3.6% mean average precision (mAP). On our self-built UAV ship draft dataset, TGNet achieves 92.3% recall and 93.8% mAP75 at 56 FPS, outperforming state-of-the-art keypoint and segmentation methods while maintaining real-time efficiency. The proposed approach demonstrates a robust and efficient solution for practical autonomous draft reading.
Regular detection of pavement cracks is essential for infrastructure maintenance. However, existing methods often ignore the challenges such as the continuous evolution of crack features between video frames and the difficulty of defect quantification. To this end, this paper proposes an integrated framework for pavement crack detection, segmentation, tracking and counting based on Transformer. Firstly, we design the VitSeg-Det network, which is an integrated detection and segmentation network that can accurately locate and segment tiny cracks in complex scenes. Second, the TransTra-Count system is developed to automatically count the number of defects by combining defect tracking with width estimation. Finally, we conduct experimental verification on three datasets. The results show that the proposed method is superior to the existing deep learning methods in detection accuracy. In addition, the actual scene video test shows that the framework can accurately label the defect location and output the number of defects in real time.
ABSTRACT With the increase of semiconductor integration density, in order to cope with the increase of wafer defect complexity and types, especially the low recognition accuracy of overlapping mixed defects and unknown wafer defects, this study proposes a lightweight model for wafer defect detection called LightWMNet. First, using a hierarchical attention Encoder‐Decoder architecture, the features of wafer defect pattern (WDP) are channel recalibrated to generate high‐resolution fine‐grained features and low‐resolution coarse‐grained features. Secondly, the backbone network incorporates two novel attention modules—feedforward spatial attention (FFSa) and feedforward channel attention (FFCa)—to amplify responses in critical defect regions and suppress noise from stochastic discrete pixels. These mechanisms synergistically enhance feature discriminability without introducing significant parametric overhead. Finally, the Dice loss function and the cross entropy loss function are combined to jointly evaluate the segmentation and classification accuracy of the model. Experimental results on the public mixed wafer defect dataset MixedWM38 show that the pixel accuracy (PA), intersection over union (IoU) and Dice coefficient of the proposed network reach 98.26%, 94.83% and 97.22%, respectively. Without significantly increasing the computational complexity and size of the model, compared with the existing state‐of‐the‐art (SOTA) model, the classification accuracy of lightWMNet in single defect, three mixed defects and four mixed defects is improved by 0.5%, 0.25% and 0.89% respectively. Furthermore, we used transfer learning for the first time to evaluate the model's generalisation ability for unseen defect categories. The results showed that LightWMNet still has a certain recognition ability even in untrained wafer defects.
This paper presents a systematic survey of machine vision-based surface defect detection technologies, focusing on five core challenges in the field: interference from complex backgrounds, small object detection, class imbalance, dynamic scene modeling, and cross-scenario generalization. It reviews key technical approaches corresponding to these challenges over the past five years. Furthermore, a dataset characterization analysis framework is established around these challenges, summarizing and comparing the characteristics of over 40 publicly available datasets across more than ten scenarios, including PCB, photovoltaic, metal, and pavement surfaces. Quantitative selection metrics (such as the small target coefficient and texture complexity) are proposed for challenges like small target detection and complex backgrounds, offering a methodological guide for aligning research questions with benchmark data. Finally, the paper summarizes current limitations and provides an outlook on new paradigms driven by large-scale models and the construction of high-quality benchmark datasets, aiming to offer valuable references for both research and engineering practices in this field.
The fault diagnosis technology of photovoltaic (PV) components is very important to ensure the stable operation of PV power station. The application of intelligent fault detection method can effectively improve the accuracy and efficiency of fault detection. In this paper, the latest progress in the field of PV module fault diagnosis in recent years is reviewed, with emphasis on fault detection methods based on electrical characteristic parameters and image processing technology. Firstly, this paper introduces the types, causes and traditional diagnosis methods of the common faults of PV modules. Then, the fault detection technology based on electrical characteristic parameters was discussed, including the method based on I-V characteristic curve analysis and mathematical model. Then, the paper reviews the fault detection technology based on image processing, especially the application of unmanned aerial vehicle (UAV)-assisted image detection technology in fault location and identification of PV modules, and focuses on the development and challenges of machine vision technology in fault detection. Finally, the existing data sets and performance evaluation indicators are summarized, and the development trend of intelligent fault diagnosis technology for PV modules in the future is prospected. This paper aims to provide reference for researchers in related fields and promote the innovation and development of PV module fault diagnosis technology.
In the present study, cubic boron nitride (cBN) and yttrium oxide (Y2O3) composite ceramics were densified via spark plasma sintering of mixed cBN-Y2O3 composite powders at temperatures ranging from 1000 to 1300 degrees C. Four distinct phases, including cBN, hexagonal BN (hBN), Y2O3 and Y3BO6,were observed within sintered cBN-Y2O3 composite ceramics, attributed to cBN phase transformation and chemical reactions during the sintering process. The phase transformed hBN was found to degrade the properties of ceramic samples, resulting in reduced relative density and hardness. The incorporation of micron-sized cBN particles could improve the hardness of composite ceramics due to enhanced densification and inhibited phase transformation. In contrast, nano-sized cBN led to reduced densification and hardness of the composite ceramics due to the accelerated phase transformation.
In this study, Y2O3-MgO nanocomposite ceramics exhibiting high transparency in the visible spectrum, composed of a 50:50 vol ratio, was fabricated through the combination of sol-gel and low-temperature high-pressure (LTHP) sintering in the absence of exogenous dopants. The Y2O3-MgO nanocomposite precursor nanopowders was synthesized using citric acid monohydrate-nitrate combustion route. The effects of sintering parameters on the phase composition, microstructure, and optical property of the samples were investigated. The results indicated highly visible-spectrum transparent Y2O3-MgO nanocomposite ceramics with a grain size below 10 nm were obtained at an applied pressure of 5 GPa at 300 degrees C. The optical transmittance of 2 mm thick Y2O3-MgO ceramics reached a maximum of 78 % at a wavelength of 1.2 mu m, significantly surpassing the existing opaque Y2O3-MgO nanocomposite transparent ceramics in the near-IR wavelength region. The reported findings could inspire further development of transparent Y2O3-MgO composite ceramics for a wider range of optical applications.
Small object detection is a challenging issue in computer vision. Most of the existing methods measure the similarity between two bounding boxes using the intersection over union. However, intersection over union is very sensitive to small shifts between two bounding boxes, which is not conducive to small object detection. In common decoupled-head structure, the design that classification and localization enjoy the shared input cannot solve the conflict between localization and classification well, resulting in an imperfect balance between these two tasks. To overcome the above limitations, we first propose the center deviation degree. By integrating center deviation degree with the intersection over union variants, we establish a novel hybrid evaluation metric that substantially improves detection accuracy across multiple object scales, while specifically addressing the critical challenges associated with small object detection. Moreover, we generate semantic context-rich features and detailed spatial features as the inputs for classification and localization branch, respectively. Decouple the two tasks by providing them with independent input features for more accurate object detection. Extensive experiments show that, when equipped with hybrid evaluation metric and classification-localization context decoupling, our method achieves gains of 1.8, 1.7, 1.7 and 1.9 points on very tiny objects, tiny objects, small objects and medium objects respectively, compared with the baseline.
The image data acquired through unmanned aerial vehicle (UAV)-based infrared imaging typically feature complex backgrounds. Moreover, photovoltaic (PV) faults appear as small, low-contrast anomalies, making them difficult to detect accurately using conventional methods. In this article, we design a dual-path multiscale and reparameterized real-time detection transformer (DMR-RTDETR) network, which mainly consists of multipath cross-scale attention feature fusion (MPCAFF) module, RepshuffleRoss module, and optimized prediction head. First, the MPCAFF module recalibrates and fuses global and local feature maps in parallel to enhance the detection of small faults. Second, the RePshuffleRoss reparameterization feature shuffling module uses multipath convolutions and channel shuffling to boost feature diversity, while reparameterization during inference reduces computational complexity. Third, an optimized RepConv-based detection head is designed to streamline the decoder and enhance accuracy by selectively integrating $40\times 40$ and $80\times 80$ scale feature maps. In addition, we also propose the PV fault detection (PVFD) dataset, which is the first target detection database covering multiple PV fault types. Experimental results show that the DMR-RTDETR model has a mean average precision (mAP) of 0.871 and an average precision (AP) of 0.448 on the PVFD dataset, confirming its effectiveness in detecting small target faults, typically defined as occupying less than $32\times 32$ pixels. Finally, through comparative experiments on the public VisDrone2019 benchmark dataset, the generalization ability of the proposed model is further verified.
Object detection has greatly improved over the past decade, thanks to advances in deep learning and large-scale datasets. However, detecting objects reflected on surfaces remains an underexplored area. Reflective surfaces are ubiquitous in daily life, appearing in homes, offices, public spaces, and natural environments. Accurate detection and interpretation of reflected objects are essential for various applications. This paper addresses this gap by introducing an extensive benchmark specifically designed for Reflected Object Detection. Our Reflected Object Detection (ROD) dataset features a diverse collection of images showcasing reflected objects in various contexts, providing standard annotations for both real and reflected objects. This distinguishes it from traditional object detection benchmarks. The ROD dataset encompasses 10 categories and 6 reflective surfaces, including 23,520 images of real and reflected objects on different backgrounds, complete with standard bounding-box annotations and the classification of objects as real or reflected. In addition, we present baseline results by adapting five state-of-the-art object detection models to address this challenging task. The experimental results underscore the limitations of existing methods when applied to reflected object detection, highlighting the need for specialized approaches. By releasing the ROD dataset, we aim to support and advance future research on detecting reflected objects. The dataset and code are available at: https://github.com/jirouvan/ROD .
The phase stability of cubic boron nitride (cBN) at elevated temperatures is an important region for both pure BN and cBN-based composite ceramics due to the significant property differences between distinct BN polymorphs. This study aims to explore the phase transformation behavior of cBN via the spark plasma sintering (SPS) process. It is revealed that the phase transformation from cBN to hBN favors an indirect phase transformation route with an intermediate phase of rBN. In addition, the phase transformed bulk BN through SPS exhibits high densification (similar to 96.7% of theoretical density), exceptional nanomechanical properties and good thermal conductivity. The fundamental study on the phase transformation of cBN, along with the excellent properties of dense BN bulks via the SPS process, provides a profound insight into BN-related materials.