Computing gradients separately for chunks of a sequence rather than the entire sequence at once, reducing peak memory.