Data leakage, where semantically or visually similar samples exist across training and test splits, continues to threaten the reliability of object detection benchmarks. The D-LeDe method was recently proposed as a statistical technique for detecting data leakage, showing promise in initial applications. This paper extends the D-LeDe method to evaluate its robustness and generalizability through a focused experimental study. We revisit the KITTI dataset which was previously identified as leakage-prone and introduce a new application of D-LeDe on SODA10M. The core investigation centers on whether visually similar images, measured using perceptual hashing, are the primary cause of data leakage indications captured by D-LeDe. To this end, we progressively remove visually similar image pairs from the test sets of both datasets and observe changes in the Relative Increase Rate, the key decision metric of D-LeDe. Results show that even after removing about 50
更多
查看译文
关键词
Data leakage detection,Object detection,YOLOv7,Kitti,Soda10m,Automotive perception systems