When a model's training data resembles its test scenarios, inflating performance metrics and obscuring true capabilities.