Machine learning models for toxicity prediction are routinely evaluated using random train/test splits, allowing structurally similar compounds to appear in both sets and inflating reported performance metrics. A rigorous benchmark quantifying this overestimation alongside calibration, uncertainty, and applicability domain analyses is needed to guide practitioners in selecting and trusting toxicity prediction models. We present ToxBench, a leakage-audited benchmark comprising three toxicology datasets (Tox21, 7,538 compounds, 12 tasks; ClinTox, 1,379 compounds, 2 tasks; SIDER, 1,350 compounds, 27 tasks) processed through a transparent standardization pipeline with explicit reporting of conflicting-label removals. Four model classes were evaluated: Random Forest, XGBoost, MLP, and Graph Neural Network (GNN), each trained under random and Bemis-Murcko scaffold-based splits across five independent seeds (120 experimental conditions). Analyses included post-hoc probability calibration, ensemble-based uncertainty quantification, nearest-neighbor applicability domain analysis, and scaffold-level error analysis. Scaffold splitting consistently reduced AUROC by 0.057–0.079 points across all model classes on Tox21 (mean drop: 0.070) and by 0.031–0.035 points on SIDER for three of four models, demonstrating systematic performance overestimation under random splitting. ClinTox showed reversed performance ordering due to small dataset size and extreme class imbalance, with high seed-to-seed variance (± 0.085–0.160) confirming results are dominated by sampling noise. Post-hoc calibration reduced ECE by 67–68