Artificial intelligence (AI) systems have quietly crossed an uncomfortable threshold: on a growing range of diagnostic tasks, they often match, and in some well-defined tasks even exceed, human performance. Evidence from large language models such as Med-PaLM and GPT-4, and deep learning systems in pathology, radiology, and hematology, shows high accuracy and reproducibility, challenging traditional human benchmarks, although still being subject to substantial heterogeneity in the studies cited. Yet most evaluation frameworks continue to treat human consensus on curated, retrospectively labeled datasets as ground truth, despite well-documented record of cognitive bias, interobserver variability, and diagnostic errors in clinical practice. This creates a striking asymmetry: AI is scrutinized against a putative gold standard whose gold content has never been formally tested—a benchmark illusion. Acknowledging AI's own distinctive failure modes and generalizability problems, this Viewpoint argues that such limitations do not justify clinging to an untested reference. It calls for explicit modeling of ground truth uncertainty, outcome-based validation, and, where appropriate, using AI as an additional reference layer for human performance. The relevant benchmark is no longer clinicians alone, but the human–AI ensemble evaluated within real diagnostic ecosystems and against patient-relevant outcomes. Funding None.
更多
查看译文
关键词
Artificial intelligence,Ground truth uncertainty,Diagnostic benchmarking,Human–AI ensemble