生成式模型的进步提高了视频保真度,使得长时域生成、交互式世界建模以及不断演变的视觉环境成为可能。自回归(AR)视频生成通过因果展开来扩展视觉序列。然而,一个根本性的瓶颈随之出现:随着生成序列的扩展,实际模型必须在严格受限的上下文窗口、存储和计算能力下运行。因此,关键的历史信息,例如实体身份、动态状态以及由干预引发的因果变化,往往在它们的相关性降低之前很久就脱离了活跃上下文。克服这一限制并维持时间持久性构成了一个根本性的记忆问题。我们系统地综述了AR视频生成中的记忆机制。我们将记忆操作性地定义为跨外层AR步骤维持的持久历史信息,即使原始证据不再在局部可访问,它仍能影响未来的生成。基于这一统一框架,我们通过五个互补的视角组织文献:(I)形式,即历史的表征载体;(II)功能,即需要保留的具体语义和物理信息;(III)操作,即记忆的写入、读取、更新、管理和整合的生命周期;(IV)学习,即在闭环展开下对记忆行为的优化;以及(V)评估,即诊断真正记忆能力的范式。最后,我们综合了开放挑战,包括可组合且资源感知的记忆架构、可信的状态更新、自展开学习以及标准化评估。通过连接表征、机制和学习范式,本文建立了开发可靠的条件记忆视频生成系统的结构化基础。
Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, practical models must operate under strictly bounded context windows, storage, and computational limits. Consequently, critical historical information, e.g., entity identities, dynamic states, and intervention-induced causal changes, often leaves the active context long before its relevance diminishes. Overcoming this limitation and maintaining temporal persistence constitutes a fundamental memory problem. We present a systematic and comprehensive review of memory mechanisms in AR video generation. We formulate memory operationally as persistent historical information maintained across outer AR steps, capable of influencing future generation even after the originating evidence is no longer locally accessible. Building upon this unified framework, we organize the literature through five complementary perspectives: (I) Forms, the representational carriers of history; (II) Functions, the specific semantic and physical information requiring preservation; (III) Operations, the lifecycle of writing, reading, updating, managing, and integrating memory; (IV) Learning, the optimization of memory behaviors under closed-loop rollouts; and (V) Evaluation, the paradigms for diagnosing genuine memory capabilities. We conclude by synthesizing open challenges, including composable and resource-aware memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation. By bridging representations, mechanisms, and learning paradigms, this paper establishes a structured foundation for developing reliable, memory-conditioned video generation systems.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-27 | 9.77 | 16 | 入选 |
| 2026-09-26 | 10.22 | 23 | 未入选 |