Robust pavement crack detection in complex scenes remains a significant challenge. This stems not merely from the scarcity of annotated data, but more critically, from the severe lack of pattern diversity within existing datasets. Key variations in morphology, scale, background texture, and imaging conditions are often underrepresented, which fundamentally impedes the generalization capability of recognition models. While generative approaches (e.g., GANs and diffusion models) offer a potential path to augment this diversity synthetically, they commonly suffer from poor background realism, entangled structural-appearance representations, and a lack of precise control over generated defects. This paper presents the Crack Diffusion Generator (DiffCrack), a diffusion-based framework designed for semantic-structural controllability in crack image synthesis. DiffCrack decouples crack geometry and visual appearance through two conditioning inputs: a binary mask to anchor spatial layout, and a Hierarchical Prompt Attention (HPA) module to independently modulate attributes such as width, depth, color, and texture. This design enables the controllable and targeted generation of diverse, photorealistic crack patterns that are often missing from real-world datasets. Extensive experiments on real datasets demonstrate that training with DiffCrack-generated images enhances the F1-score of segmentation models by up to 23% on average under complex scene conditions. This result validates that DiffCrack is a scalable, pattern-aware data augmentation tool. By addressing the critical bottleneck of data diversity, our framework offers a principled pathway to improving the robustness and generalization of infrastructure inspection models.
更多
查看译文
关键词
Image generation,Pavement cracks,Diffusion model,Data enhancement,Multimodal data