An evaluation metric that compares model output to expected output without validating actual execution results.