An alternative to the attention mechanism that processes sequences more efficiently by scaling sub-quadratically, meaning it uses fewer computational resources as sequences get longer.
Performance retention over long documents and conversations