Targeted fine-tuning where an expert rewrites only the failing steps in a weaker model's rollout, rather than providing full trajectory demonstrations.