Positional encodings like RoPE and ALiBi don't automatically help transformers generalize to unseen token distances—data diversity and task structure matter more than the encoding scheme itself.
This paper investigates how transformers generalize to different token distances between training and inference, using synthetic copy tasks. It compares positional encoding schemes (RoPE, ALiBi, no encoding) and finds that understanding distance generalization requires rethinking how we use positional information.