近期视频扩散技术的进展展示了令人瞩目的超高保真生成能力,相关模型能够在数秒内渲染出逼真的场景。然而,尽管扩散模型能够生成高保真的视频片段,但将其转化为连贯的长篇叙事引擎仍然充满挑战。
现有的大多数智能体流水线通过链式模块自动化这一过程,但由于采用独立的、手工设计的提示词,它们往往存在语义漂移(即镜头间角色服装或景物的细微变化)和级联故障(例如上游资产伪影破坏下游视频合成)的问题。由于早期错误会传播并破坏长周期的连贯性,该过程通常需要详尽的人工干预。从结构角度来看,这反映了经典的信用分配问题,因为终端故障难以追溯至特定的提示词。此外,现有方法还存在特征漂移问题,即实体和环境逐渐发生非预期的变化,或者内容坍塌问题,即叙事无法有意义地推进。
今天,我们介绍了关于AI视频联合导演的一项研究,这是一个统一的、多智能体框架,明确规划多镜头叙事中的视觉连续性。该框架构建在Gemini和Veo之上作为编排层,原生继承了如SynthID水印等安全机制。通过将长格式生成视为全局优化和世界状态跟踪问题,我们开发了一系列框架——Co-Director(将发表于COLM 2026)、CANVAS(将发表于EMNLP 2026)、A²RD和VQQA——它们通过自动化重复的编排任务,从多模型提示词和镜头链接到闭环视觉优化,将高层的人类创意规范转化为执行。
我们设计这些框架旨在作为响应式的创意伙伴,抽象出维持视觉连续性的负担,使用户能够专注于讲故事的艺术。这种架构通过将质量建模为测试时目标,将创意合成与一致性解耦。在全面的评估中,我们的框架在多镜头叙事一致性和角色持久性方面取得了显著的提升,成功生成了长达数分钟的视频,同时减轻了视觉漂移和流水线错误传播的问题。
Recent advancements in video diffusion demonstrate remarkable high-fidelity generation with models that can render realistic scenes in seconds. However, while diffusion models generate high-fidelity video clips, transforming them into coherent long storytelling engines remains challenging.
Most existing agentic pipelines automate this process via chained modules but suffer from semantic drift (subtle shifts in character attire or scenery across shots) and cascading failures (e.g., an upstream asset artifact corrupting downstream video synthesis) due to independent, handcrafted prompting. Because early errors propagate and break long-horizon consistency, the process often requires exhaustive manual intervention. From a structural perspective, this reflects the classical credit assignment problem, as terminal failures are difficult to trace back to specific prompts. Furthermore, existing methods suffer from feature drift , where entities and environments gradually change unintentionally, or content collapse , where narratives fail to progress meaningfully.
Today, we introduce our research on an AI video co-director, a unified, multi-agent framework that explicitly plans visual continuity in multi-shot narratives. Built as an orchestration layer on top of Gemini and Veo , this framework natively inherits safety mechanisms like SynthID watermarking . By treating long-form generation as a global optimization and world-state tracking problem, we have developed a suite of frameworks — Co-Director (to appear at COLM 2026 ), CANVAS (to appear at EMNLP 2026 ), A²RD , and VQQA —that translate high-level human creative specification into execution by automating repetitive orchestration tasks, from multi-model prompting and shot chaining to closed-loop visual refinement.
We designed these frameworks to act as responsive creative partners that abstract away the burdens of maintaining visual continuity, freeing users to concentrate on the art of storytelling. This architecture decouples creative synthesis from consistency by modeling quality as a test-time objective. Across comprehensive evaluations, our framework demonstrates substantial gains in multi-shot narrative consistency and character persistence, successfully generating minutes-long videos while mitigating visual drift and pipeline error propagation.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-05 | 8.4 | 20 | 入选 |
| 2026-10-04 | 8.4 | 33 | 未入选 |
| 2026-10-03 | 8.4 | 44 | 未入选 |
| 2026-09-25 | 10.84 | 16 | 入选 |