Multi-stream architectures like mHC provide more capacity than models actually use—most layers rely on just two streams, and late-layer stream mixing can be removed with minimal performance loss, suggesting opportunities for efficiency improvements.
This paper investigates how DeepSeek-V4-Flash uses its multi-stream residual pathways (mHC architecture). Researchers found that despite having four parallel streams, the model typically uses only two streams per layer, with minimal mixing between streams in later layers. Interventions show that early-layer mixing is critical for performance, while late-layer mixing is largely redundant.