Reply To: Limited Benchmarks Constrain the Conclusions of a General-Purpose Versus Clinical AI Comparison | AMiner