In machine learning for QSAR/QSPR, the choice of train–test splitting algorithm alters both a model’s realized performance and the accuracy with which that performance is estimated from the held-out test set. Prior work has established that structure-aware splits such as Kennard–Stone and SPXY produce optimistically biased internal estimates, but these characterizations typically examine one method family at a time. Here we quantify how fifteen splitting strategies from four method families differ across fifteen drug-discovery datasets on two criteria (realized external benchmark performance and performance-estimation bias), and further assess whether splitting strategy affects model-selection quality when multiple model classes are optimized simultaneously. A single, consistent pattern holds across all fifteen datasets and all eleven evaluation measures (RMSE, MAE, MedAE, R^2 , Spearman ρ , Pearson r, Kendall τ , enrichment factor at 5/10/20