Investigates techniques for long-term memory and context management in video generation models, enabling coherent and extended video sequences.
This topic focuses on methods for maintaining long-term memory and context within video generation models. It encompasses techniques such as compressed memory tokens, KV cache consolidation, state-space models, and hybrid attention, all aimed at enabling coherent, long-horizon video generation. The scope includes memory mechanisms in video world models but specifically excludes reinforcement learning-focused world models and theoretical certification works not directly addressing video generation memory.