Nano Banana、Stable Diffusion 和 Flux 等文本生成图像 AI 模型的快速发展,从根本上改变了创意设计的格局,使得任何人都能够根据文本描述合成逼真且高保真的图像。然而,引导这些庞大的模型以满足精确的用户意图、下游目标或严格的视觉约束,仍然是一项微妙且不可预测的平衡艺术。例如,想象一下向模型提示“一只戴着太阳镜的蜥蜴”。模型可能会生成一只逼真的蜥蜴,但它并没有戴太阳镜。或者,强行让模型包含太阳镜可能会导致蜥蜴的面部变形,从而破坏图像质量。
现有的引导或微调图像生成的方法彼此之间非常脱节。一方面,开发人员使用推理时技术(例如无分类器扩散引导)来调整文本提示的影响,并实时引导图像生成过程。另一方面,他们依赖于使用参数高效适配器(如 LoRA)、奖励加权回归或策略梯度等进行重度微调,以改变模型的行为。
由于这些工具在历史上一直被视为独立且不相关的修补方案,该领域缺乏一种统一的、原则性的数学语言来规范、分析和优化我们对生成模型的控制方式。这种碎片化的方法通常迫使工程师在平衡用户偏好对齐与图像质量时依赖猜测。
为了解决这一平衡难题,我们提出了 Diffusion Controller 框架。Diffusion Controller 不再将图像生成视为一系列孤立的固定步骤,而是将整个去噪过程重构为一个平滑、连续的控制问题。我们的结果表明,Diffusion Controller 的轻量级附加网络在匹配人类偏好方面优于行业标准。此外,其完全解锁版本(即具有“白盒”或不受限制访问权限以更改内部模型权重的微调模型)相对于基线模型取得了 90% 的胜利率。
The rapid advancement of text-to-image AI models, such as Nano Banana , Stable Diffusion and Flux , has fundamentally transformed creative design, allowing anyone to synthesize photorealistic, high-fidelity images from textual descriptions. However, steering these massive models to meet precise user intent, downstream goals, or strict visual constraints remains a delicate and unpredictable balancing act. For example, imagine prompting a model for "a lizard wearing sunglasses". The model might generate a realistic lizard that's not wearing sunglasses. Alternatively, forcing the model to include the sunglasses might distort the lizard's face, ruining the image quality.
Existing methodologies that guide or fine-tune image generation are very disconnected. On the one hand, developers use inference-time techniques (e.g., classifier-free diffusion guidance ) to adjust the text prompt’s influence and guide the image generation process on the fly. On the other hand, they rely on heavy fine-tuning using parameter-efficient adapters like LoRA , reward-weighted regressions , or policy gradients to alter a model's behavior.
Because these tools have historically been treated as distinct and unrelated fixes, the field has lacked a single, principled mathematical language to unify, analyze, and optimize how we control generative models. This fragmented approach often forces engineers to rely on guesswork when balancing user preference alignment against image quality.
To solve this balancing act, we present the Diffusion Controller framework. Instead of treating image generation as a rigid sequence of isolated steps, Diffusion Controller reframes the entire denoising process as a smooth, continuous control problem. Our results show that Diffusion Controller’s lightweight add-on network outperformed the industry standard for matching human preferences. Moreover, its fully unlocked version (i.e., the fine-tuned model with "white-box" or unrestricted access to alter internal model weights) achieved a 90% win rate over the baseline model.
首次收录 · 2026-09-30 · 10.79 分