Zhenjiang Research Institute of Advanced Equipment
被引用0|浏览0
摘要
Noisy labels are ubiquitous in real-world datasets, causing deep neural networks to overfit and suffer from poor generalization. Despite label corruption, prior works observe that during the “early learning” phase, learned representations of images from the same category still congregate together. Digging deeper into this phenomenon, we propose a novel framework to mitigate noisy supervision by creating synthetic samples on Riemannian manifolds. Unlike Euclidean methods that ignore intrinsic data geometry, our approach synthesizes features by aggregating original samples with their top-K neighbors along the geodesic paths, where the aggregation weights are determined by modeling the loss distribution. These synthetic samples serve as denoised proxies that effectively smooth the decision boundary and prevent the memorization of erroneous labels. Additionally, these synthesized representations estimate soft targets, progressively correcting noisy labels and yielding more separated, clearly bounded clusters. Experiments on different datasets demonstrate that our method outperforms state-of-the-art approaches, enhancing the robustness of learned representations.