Mixture of Contexts for Long Video Generation

“Minute-long context memory with short-video cost”
Anonymous Authors
Learnable Sparse Attention Routing

A non-parametric yet learnable router selects informative history chunks to calculate attention. The routing is implicitly differentiable and optimized end-to-end from large-scale long video data.

Multi-shot Long Video Generation

For 8-shot, minute-long videos (~180k tokens), MoC prunes ≈ 85 % of token pairs and cuts FLOPs by 7× while delivering minute-long context coherence.

Some prompts are from Twitter artworks.

Single-shot Short Video Generation

On 8-second, 320×192 clips (~6.5k tokens), MoC routing prunes ≈ 83 % of token pairs, while maintaining or even improving video quality.

Qualitative Comparisons

Side-by-side comparisons between dense attention and our Mixture of Contexts.

Multi-shot Comparisons
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Single-shot Comparisons
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)
Dense Attention
Mixture of Contexts (ours)