The difference between a reasoning-tuned teacher model and its base pre-trained version, used as a training target.