Potholes are a critical infrastructure distress that damages vehicles and creates safety risks, motivating timely detection for effective maintenance. Automated detection via computer vision is a scalable solution, but its progress is hindered by the scarcity of high-quality, large-scale public datasets. Here we present HRP4K, a high-resolution, perspective-view road image dataset for developing and evaluating pothole detectors. Images are captured by vehicle-mounted full-frame mirrorless cameras across 1,100 km of urban and rural roads in China. The dataset comprises 6,003 images, including 4,003 positive images containing at least one pothole and 2,000 negative images of pothole-absent road surfaces. In total, HRP4K provides 7,217 pothole instances annotated with bounding boxes and is released in both YOLO and COCO formats. A human-in-the-loop pipeline involving algorithmic preprocessing, privacy anonymization, and iterative model-assisted annotation ensured consistent, high-fidelity labels. The data exhibit realistic long-tailed distributions of small and ultra-small potholes in visually complex scenes, providing a challenging benchmark for object detection. We report baseline results using six modern detectors to facilitate standardized comparison and benchmarking in automated infrastructure monitoring.
Accurate and efficient pixel-level crack detection is essential for infrastructure health monitoring. Although recent encoder-decoder architectures have achieved notable results, their reliance on aggressive downsampling often compromises the measurement-grade spatial precision required for fine crack detection. To address this gap, we present HACNetV2, an enhanced full-resolution architecture that builds upon our previous HACNet. Through key architectural refinements and novel components, HACNetV2 achieves substantially improved efficiency and effectiveness compared to its predecessor. HACNetV2 employs a dual-branch structure that integrates full-resolution and high-resolution processing to maintain critical spatial detail without imposing a heavy computational burden. Additionally, we propose two novel components: the hybrid atrous spatial pyramid attention module (HybridASPA), which expands the receptive field while reducing computational complexity, and the crack-aware attention module (CrackAM), which enhances sensitivity to fine and complex crack features. Together, these components enable HACNetV2 to enhance feature representation effectively. The experimental results on the public BCL and CHCrack5K benchmark datasets demonstrate that HACNetV2 outperforms recent models, achieving superior mIoU and F1 scores while maintaining a lightweight structure with only 0.52 M parameters. Furthermore, HACNetV2 supports real-time inference capabilities, processing 480 x 480 images at 47 FPS on an RTX 3090 GPU.
BackgroundEarly diagnosis in oral cancer is essential to reduce both morbidity and mortality. This study explores the use of uncertainty estimation in deep learning for early oral cancer diagnosis.MethodsWe develop a Bayesian deep learning model termed 'Probabilistic HRNet', which utilizes the ensemble MC dropout method on HRNet. Additionally, two oral lesion datasets with distinct distributions are created. We conduct a retrospective study to assess the predictive performance and uncertainty of Probabilistic HRNet across these datasets.ResultsProbabilistic HRNet performs optimally on the In-domain test set, achieving an F1 score of 95.3% and an AUC of 96.9% by excluding the top 30% high-uncertainty samples. For evaluations on the Domain-shift test set, the results show an F1 score of 64.9% and an AUC of 80.3%. After excluding 30% of the high-uncertainty samples, these metrics improve to an F1 score of 74.4% and an AUC of 85.6%.ConclusionRedirecting samples with high uncertainty to experts for subsequent diagnosis significantly decreases the rates of misdiagnosis, which highlights that uncertainty estimation is vital to ensure safe decision making for computer-aided early oral cancer diagnosis.
道路破损对于道路安全有着巨大威胁,准确检测道路破损对道路养护修缮具有重要意义.针对人工检测方法检测效率低、成本高,国内外公司开发的商用路面检测系统成本高、性价比低的等问题,文章设计并实现一种基于YOLOv7模型与边缘检测设备Jetson Orin的路面破损检测系统,通过USB摄像头或工业相机采集道路图像,边缘计算设备Jetson Orin进行图像处理,利用YOLOv7模型进行路面破损检测,可检测出裂缝、井盖、坑槽、修补四种类型的缺陷.系统体积小,成本低、性能强、算法准确率高,可部署在普通的家用汽车进行道路缺陷实时检测.
Aiming at low efficiency, high risk, and difficulty in quantifying the degree of loosening of high-strength bolts in steel bridges, a batch bolt loosening detection method based on key point identification is proposed. Firstly,the key points of bolts in the image are located by the convolutional neural network model, then the K-means clustering algorithm is used to envelope the key points of the bolt, and the initial angle and loosening angle of the bolt are calculated. The performance of the algorithm is verified by collecting bolt images in the laboratory. The results show that the root mean square error of the loosening test in the laboratory is 0.63°~1.83°; the maximum error is 2°, and the average detection time is only 51 ms. The loosening angle identified by this method is very consistent with the actual loosening angle, and the test error meets the requirements of engineering application.
针对钢桥高强度螺栓人工批量巡检效率低、接触传感设备成本高昂的问题,提出了一种基于图像识别的高强度螺栓松动检测方法.利用螺栓角点位置识别算法对螺栓图像样本进行白色掩膜构建、掩膜小型噪点剔除和感兴趣区域分割等处理,确定螺栓角点位置坐标,结合相机成像相似映射原理推算螺栓松动角度,进而根据螺栓松动角度评估预紧力损失.对不同型号螺栓在不同水平视距下旋转10°、20°、30°采集样本,并将其导入算法进行试验验证.结果表明,基于该方法的螺栓松动角度检测准确率达90%以上,满足工程检测要求,并能够有效评估高强度螺栓预紧力损失.
Automated pixel-level crack detection is one of the essential tasks in the field of defect inspection. Deep convolutional neural networks, typically using encoder–decoder architectures, have been successfully applied to many crack detection scenes in recent works. However, encoder–decoder networks commonly rely on downsampling and upsampling operations and have a large number of parameters, which may influence the accuracy of crack prediction due to the cracks usually have long, narrow sizes, and the labeled training set is always limited. To address these issues, we propose a simple and effective hybrid atrous convolutional network (HACNet). HACNet maintains the same spatial resolution throughout the whole architecture. It can retain more spatial precision in prediction. HACNet uses atrous convolutions with the proper dilation rates to enlarge the receptive field and a hybrid approach connecting these convolutions to aggregate multiscale features. The resulting architecture can achieve accurate segmentation with relatively few parameters. Evaluations on the public CFD data set, CrackTree206 data set, Deepcrack data set (DCD), and Yang et al. Crack data set (YCD) demonstrate that our method can obtain promising results, compared with other recent approaches. Evaluation on self-collected images and SDNET2018 data set illustrates the good potential of HACNet for practical applications.
SIGNIFICANCE:Oral cancer is a quite common global health issue. Early diagnosis of cancerous and potentially malignant disorders in the oral cavity would significantly increase the survival rate of oral cancer. Previously reported smartphone-based images detection methods for oral cancer mainly focus on demonstrating the effectiveness of their methodology, yet it still lacks systematic study on how to improve the diagnosis accuracy on oral disease using hand-held smartphone photographic images.AIM:We present an effective smartphone-based imaging diagnosis method, powered by a deep learning algorithm, to address the challenges of automatic detection of oral diseases.APPROACH:We conducted a retrospective study. First, a simple yet effective centered rule image-capturing approach was proposed for collecting oral cavity images. Then, based on this method, a medium-sized oral dataset with five categories of diseases was created, and a resampling method was presented to alleviate the effect of image variability from hand-held smartphone cameras. Finally, a recent deep learning network (HRNet) was introduced to evaluate the performance of our method for oral cancer detection.RESULTS:The performance of the proposed method achieved a sensitivity of 83.0%, specificity of 96.6%, precision of 84.3%, and F1 of 83.6% on 455 test images. The proposed "center positioning" method was about 8% higher than that of a simulated "random positioning" method in terms of F1 score, the resampling method had additional 6% of performance improvement, and the introduced HRNet achieved slightly better performance than VGG16, ResNet50, and DenseNet169, with respect to the metrics of sensitivity, specificity, precision, and F1.CONCLUSIONS:Capturing oral images centered on the lesion, resampling the cases in training set, and using the HRNet can effectively improve the performance of deep learning algorithm on oral cancer detection. The smartphone-based imaging with deep learning method has good potential for primary oral cancer diagnosis.
Cracks are one of the most common types of surface defects that occur on various engineering infrastructures. Visual-based crack detection is a challenging step due to the variation of size, shape, and appearance of cracks. Existing convolutional neural network (CNN)-based crack detection networks, typically using encoder-decoder architectures, may suffer from loss of spatial resolution in the high-to-low and low-to-high resolution processes, affecting the accuracy of prediction. Therefore, we propose H R N e t e , an enhanced version of a high-resolution network (HRNet), by removing the downsampling operation in the initial stage, reducing the number of high-resolution representation layers, using dilated convolution, and introducing hierarchical feature integration. Experiments show that the proposed H R N e t e with relatively few parameters can achieve more accuracy and robust performance than other recent approaches.
针对传统项目化教学法实施过程中学生积极性低、学习效果差和教学管理难等问题,对迭代式教学法在高职计算机类项目化课程中的实践展开研究,并以"JavaScript程序设计"课程为例阐述迭代式教学法在计算机类项目化课程中的具体应用.教学实践表明,该方法能激发学生的热情和兴趣,发挥学生的创新能力,提高学生的实践设计水平.
针对传统机器人学习离线、任务特定、无法在学习过程中扩展智能, 且实时性、适应性差等问题, 借鉴认知机器人、神经生物学、认知科学等思想, 提出一种仿人脑多巴胺调控机制的机器人视觉自发育学习算法, 模拟人脑海马与前额叶神经回路, 用多巴胺调控学习进程, 实现仿人类的视觉学习能力.实验结果表明: 该算法能有效地实现视觉图像的自发育学习, 能完成非特定任务, 且识别率高, 实时性好, 适应性强.
Poor road conditions, such as potholes, are a nuisance to society, which would annoy passengers, damage vehicles, and even cause accidents. Thus, detecting potholes is an important step toward pavement maintenance and rehabilitation to improve road conditions. Potholes have different shapes, scales, shadows, and illumination effects, and highly complicated backgrounds can be involved. Therefore, detection of potholes in road images is still a challenging task. In this study, we focus on pothole detection in 2D vision and present a new method to detect potholes based on location-aware convolutional neural networks, which focuses on the discriminative regions in the road instead of the global context. It consists of two main subnetworks: the first localization subnetwork employs a high recall network model to find as many candidate regions as possible, and the second part-based subnetwork performs classification on the candidates on which the network is expected to focus. The experiments using the public pothole dataset show that the proposed method could achieve high precision (95.2%), recall (92.0%) simultaneously, and outperform the most existing methods. The results also demonstrate that accurate part localization considerably increases classification performance while maintains high computational efficiency. The source code is available at https://github.com/hanshenchen/pothole-detection.
Crack detection is one of the most important works in the system of pavement management. Cracks do not have a certain shape and the appearance of cracks usually changes drastically in different lighting conditions, making it hard to be detected by the algorithm with imagery analytics. To address these issues, we propose an effective U-shaped fully convolutional neural network called UCrackNet. First, a dropout layer is added into the skip connection to achieve better generalization. Second, pooling indices is used to reduce the shift and distortion during the up-sampling process. Third, four atrous convolutions with different dilation rates are densely connected in the bridge block, so that the receptive field of the network could cover each pixel of the whole image. In addition, multi-level fusion is introduced in the output stage to achieve better performance. Evaluations on the two public CrackTree206 and AIMCrack datasets demonstrate that the proposed method achieves high accuracy results and good generalization ability.
Cracks are one of the most common categories of pavement distress that may potentially threaten road and highway safety. Thus, a reliable and efficient pixel-level method of crack detection is necessary for real-time measurement of the crack. However, many existing encoder-decoder architectures for crack detection are time-consuming because the part of decoder module always has lots of convolutional layers and feature channels that lead to performance that highly relies on computing resources, which is a handicap in scenarios with limited resources. In this study, we propose a simple and effective method to boost the algorithmic efficiency based on encoder-decoder architecture for crack detection. We develop a switch module, called SWM, to predict whether the image is positive or negative and then skip the decoder module to save computation time when it is negative. This method uses the encoder module as the fixed feature extractor and only needs to place a light-weight classifier head on the end of the encoder module to output the final class probability. We choose the classical UNet and DeepCrack as examples of the encoder-decoder architectures to show how SWM is integrated into the architectures to reduce computation complexity. Evaluations on the public CrackTree206 and AIMCrack datasets demonstrate that our method can significantly boost the efficiency of the encoder-decoder architectures in all tasks, while without affecting the performance. The SWM can also be easily embedded into other encoder-decoder architectures for further improvement. The source code is available at https://github.com/hanshenchen/crack-detection.
车道检测是辅助驾驶和自动驾驶的重要研究内容.针对现有车道检测算法的鲁棒性和复杂度较难均衡等问题,提出一种基于多帧叠加和窗口搜索的快速车道检测算法.首先,通过逆透视变换(IPM)把指定的感兴趣区域(ROI)转换成乌瞰图,结合多帧叠加的方法把RGB图像转化成二值图.其次,根据近视场中的像素密度分布,计算当前帧的车道线起始点,并采用滑动窗口搜索的方法提取整个车道线.最后,根据车道线的特征,选择不同的车道模型,使用最小二乘法(LSE)拟合得到模型参数.大量的实际道路行驶测试结果表明,该算法能快速地检测车道线,并具有一定的鲁棒性和准确性.
In this paper,the design of compatibility dual-display based on embedded system is presented.A monochrome STN-LCD,a TFT-LCD and the LPC2478 microprocessor are introduced as examples.The two LCD hardware interface circuit based on LPC2478 are completely designed.The mainframe of the graphical user interface software design and some details of programming are introduced in detail.Finally,the results of the testing and analysis are presented.The experimental results show that all information can be integrity displayed on the LCD screen by the purposed system,and the design has a good repartition.