AcetylBench: a Deduplicated Benchmark Dataset Reveals Systematic Performance Inflation and Cross-Database Generalisation Failure in Lysine Acetylation Site Prediction | AMiner
AcetylBench: a Deduplicated Benchmark Dataset Reveals Systematic Performance Inflation and Cross-Database Generalisation Failure in Lysine Acetylation Site Prediction
Muhammad Uzair Ashraf,Shaista Khan,Mohammad Sarfraz,Rizwan Hasan Khan
Lysine acetylation is a pervasive post-translational modification with critical regulatory roles, yet computational prediction of acetylation sites remains hampered by unreliable benchmarking practices including dataset redundancy and non-independent test set evaluation. We report a systematic investigation demonstrating that three independent sources of metric inflation, including near-ubiquitous sequence redundancy, non-independent test sets, and distribution-locked dimensionality reduction collectively reduced an apparent accuracy of 99.38
更多
查看译文
关键词
Lysine acetylation,Post-translational modification,Protein language models,Benchmark dataset,Cross-database evaluation,Sequence redundancy,Machine learning,AcetylBench