High accuracy in medical AI often reflects leaky data rather than model sophistication. Simpler, interpretable models can match complex ones while being faster and more auditable for fairness and safety issues.
This paper audits cardiovascular screening models trained on health survey data, revealing that their reported high accuracy (AUROC ~0.89) comes from data leakage rather than genuine learning. By systematically removing leaky features and testing multiple model types—from simple classifiers to advanced foundation models—the authors show all models collapse to similar performance.