近期视频扩散技术的进步展示了令人瞩目的高保真生成能力,模型能够在数秒内渲染出逼真的场景。然而,尽管扩散模型能够生成高保真的视频片段,但将它们转化为连贯的长篇叙事引擎仍然充满挑战。
现有的大多数智能体管道通过链式模块自动化这一过程,但由于独立的手写提示(prompting),它们往往受到语义漂移(镜头间角色服装或景致的细微变化)和级联故障(例如,上游资产瑕疵破坏下游视频合成)的影响。由于早期错误会传播并破坏长视域的一致性,该过程通常需要详尽的人工干预。从结构角度来看,这反映了经典的信用分配问题,因为终端故障难以追溯至特定的提示词。此外,现有方法还存在特征漂移的问题,即实体和环境逐渐发生非预期的变化,或者内容崩溃,即叙事无法有意义地推进。
今天,我们介绍了关于AI视频联合导演(co-director)的研究,这是一个统一的多智能体框架,明确规划多镜头叙事中的视觉连续性。该框架构建在Gemini和Veo之上作为编排层,原生继承了如SynthID水印等安全机制。通过将长格式生成视为全局优化和世界状态跟踪问题,我们开发了一系列框架——Co-Director(将发表于COLM 2026)、CANVAS(将发表于EMNLP 2026)、A²RD和VQQA——它们通过自动化重复的编排任务(从多模型提示到镜头链式连接,再到闭环视觉优化),将高层的人类创意规范转化为执行。
我们设计这些框架旨在作为响应式的创意伙伴,抽象掉维持视觉连续性的负担,让用户能够专注于讲故事的艺术。这种架构通过将质量建模为测试时目标(test-time objective),将创意合成与一致性解耦。在全面的评估中,我们的框架在多镜头叙事一致性和角色持久性方面取得了显著的提升,成功生成了长达数分钟的视频,同时减轻了视觉漂移和管道错误传播的问题。
Recent advancements in video diffusion demonstrate remarkable high-fidelity generation with models that can render realistic scenes in seconds. However, while diffusion models generate high-fidelity video clips, transforming them into coherent long storytelling engines remains challenging.
Most existing agentic pipelines automate this process via chained modules but suffer from semantic drift (subtle shifts in character attire or scenery across shots) and cascading failures (e.g., an upstream asset artifact corrupting downstream video synthesis) due to independent, handcrafted prompting. Because early errors propagate and break long-horizon consistency, the process often requires exhaustive manual intervention. From a structural perspective, this reflects the classical credit assignment problem, as terminal failures are difficult to trace back to specific prompts. Furthermore, existing methods suffer from feature drift , where entities and environments gradually change unintentionally, or content collapse , where narratives fail to progress meaningfully.
Today, we introduce our research on an AI video co-director, a unified, multi-agent framework that explicitly plans visual continuity in multi-shot narratives. Built as an orchestration layer on top of Gemini and Veo , this framework natively inherits safety mechanisms like SynthID watermarking . By treating long-form generation as a global optimization and world-state tracking problem, we have developed a suite of frameworks — Co-Director (to appear at COLM 2026 ), CANVAS (to appear at EMNLP 2026 ), A²RD , and VQQA —that translate high-level human creative specification into execution by automating repetitive orchestration tasks, from multi-model prompting and shot chaining to closed-loop visual refinement.
We designed these frameworks to act as responsive creative partners that abstract away the burdens of maintaining visual continuity, freeing users to concentrate on the art of storytelling. This architecture decouples creative synthesis from consistency by modeling quality as a test-time objective. Across comprehensive evaluations, our framework demonstrates substantial gains in multi-shot narrative consistency and character persistence, successfully generating minutes-long videos while mitigating visual drift and pipeline error propagation.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-05 | 8.4 | 20 | 入选 |
| 2026-10-04 | 8.4 | 33 | 未入选 |
| 2026-10-03 | 8.4 | 44 | 未入选 |
| 2026-09-25 | 10.84 | 16 | 入选 |