Holo4 是我们全新的智能体模型系列。它提供两种规格:27B 密集型和 35B-A3B 混合专家(Mixture of Experts)模型。两者均可通过 H Models API 获取。同时,我们也在发布 Holotron 3 的更新版本:Holotron4 Nano。
Holo4 基于我们之前的模型构建,能够通过任何可用接口与软件交互:包括图形用户界面(GUI)、代码、MCP 和 API。它在学术基准测试中表现优异,但我们是为真实的业务流程而打造它。它通过在大量环境和任务上进行监督学习和强化学习进行训练,其中包括由我们的 Agentic Task Factory 生成的那些任务。
立即开始使用:
为每种接口打造的模型
Holo4 可以在屏幕上点击和打字,编写并运行自己的代码,并调用 MCP 或 API 工具。它会根据任务需求选择最合适的方式。大多数智能体模型仅针对单一接口进行训练:专注于 GUI 的模型在没有屏幕的情况下会“失明”,而偏好工具调用的模型则会被困在没有 API 的应用程序面前。实际工作并非如此孤立,单个业务任务可能需要结合这些不同的方法。
Holo4 可在桌面端、网页端、Android 设备、代码沙箱以及针对企业 API 的环境中运行。在每种情况下,它都是同一个模型,调用方式也相同。您无需为每个平台选择不同的模型。
Holo4 模型相比其 Qwen 基础模型有了显著提升。在长流程任务中,Holo4 仅落后于最强的闭源模型:在 OSWorld 2.0 上,Holo4 27B 得分为 61.7%,而 Opus 5.5 为 81.8%;Holo4 35B-A3B 达到 30.9%。然而,它是以数量级更少的参数和更低成本实现的。我们开源了我们在公共基准测试中得分背后的每一条轨迹:您可以在 trajectories.hcompany.ai 回放每一步,或从 Hugging Face 下载它们。
与前沿模型竞争,但成本仅为其中一小部分
在桌面控制(OSWorld 2.0)和 API 使用(AutomationBench)最难的学术基准测试中,Holo4 以远低于每项任务的成本与前沿模型竞争。
关于成本性能图表的说明
OSWorld 2.0。成本是根据每次智能体运行的输入和输出 token 估算的。Holo4 按 H Models API 费率定价(单次运行)。Qwen3.8 27B:模型卡片得分,成本来自我们在阿里云列表价格下的运行 token。Qwen3.6 35B-A3B:在我们框架中的单次运行,使用阿里云列表价格,其中缓存命中率为输入价格的 20%。OpenAI 的发布数据提供了 GPT 和 Opus 的努力范围;其他闭源和开源权重点使用官方 OSWorld 2.0 排行榜 。发布版本、框架和任务子集各不相同。该线连接了闭源模型中非支配的得分和成本对;Holo4 被排除在外。
AutomationBench。Holo4、Qwen3.8 27B 和 Qwen3.6 35B-A3B:使用 AutomationBench v1.0.6,得分和成本在我们的内部框架中测量。其他模型:来自 AutomationBench README 的公开集合得分,每项任务的成本来自官方排行榜,该排行榜在私有集合上运行。一旦 Holo4 在私有集合上得到评估,我们将报告其结果。
真正工作的 AI
Holo4 模型在来自我们 Agentic Task Factory 的环境和任务上进行训练,在专业软件方面表现出色。以下示例展示了 Holo4 27B 与其基础模型 Qwen3.8 27B 的对比。两个模型使用相同的提示词和框架。
3D 建模 · 埃菲尔铁塔
在 FreeCAD 中构建埃菲尔铁塔的 3D 模型,比例为 1 毫米至 1 米,以原点为中心并与 X 轴和 Y 轴对齐。按照此设计进行工作。
该塔在任何高度上的平面形状均为正方形,绝非圆形。从中轴线到角落测量的半宽,在地面层为 62.5 毫米,在高度 57 处为 32.5 毫米,在高度 115 处为 17.5 毫米,在高度 276 处为 9.35 毫米。在这些高度之间,半宽遵循一条平滑曲线,在地面附近陡峭下降,而在较高处平缓下降,绝非直线。
四条完全相同的腿,每条位于一个象限,均为方形柱体,其外侧角沿该轮廓延伸。每条腿在地面处的宽度为14毫米,在高度276处收窄至4毫米。这些腿从地面一直延伸至第一层平台,彼此分离,随后随着上升而逐渐汇聚。它们之间的空间没有任何填充物:塔身是开放的,你可以从任何一侧直接看穿它。
三层平台,每层均为以轴线为中心的实心方形板:在高度57处,宽度72毫米,厚度4毫米;在高度115处,宽度40毫米,厚度3毫米;在高度276处,宽度22毫米,厚度3毫米。
从高度276至324的桅杆,呈方形,底部宽度8毫米,尖端收窄至2毫米。
每个部件必须是一个具有非零体积的封闭实体,且任何部件不得填充腿部之间的空间。
Holo4 27B(84次调用,130万令牌)
Qwen3.8 27B(60次调用,100万令牌)
3D建模 · H标志
在FreeCAD中构建H公司标志的3D模型:一个实心圆盘旁边是一个块状无衬线大写字母H,两者挤出相同的厚度,两个形状高度相似且彼此分开以避免重叠,圆盘的圆心与H的中间对齐。将两个形状都着色为黑色。
Holo4 27B(94次调用,150万令牌)
Qwen3.8 27B(118次调用,190万令牌)
游戏设计 · Pac-Man
在Godot中构建一个Pac-Man风格的游戏并让其持续运行。
一个由墙壁组成的矩形迷宫,布局在网格上,每个开放走廊都填满了豆状物。玩家标记沿走廊连续移动,吃掉经过的每个豆状物并获得一分。三个幽灵在同一走廊中移动并追逐玩家。如果幽灵抓住玩家,玩家失去一条生命,所有内容重置为起始位置。分数和生命值绘制在屏幕上。
没人会玩这个游戏。玩家由一个简单的启发式算法驱动:在每个岔路口,它朝向最近的豆状物前进,除非附近有幽灵,此时它会远离幽灵。游戏必须无人值守且无限期运行,完全不需要键盘输入。
当游戏正常运行时,启动游戏并让它继续播放。
Holo4 27B(68次调用,240万令牌,268行)
Qwen3.8 27B(197次调用,1140万令牌,327行)
我们如何构建Holo4
代理任务工厂
我们内部的代理管道集仅从文档(如真实网站截图或开源软件)中构建交互式环境和可验证的任务。迄今为止,它已在Web应用、MCP服务器和桌面环境中生成了约10,000个任务,包括通过GUI和MCP暴露相同状态的混合环境。
训练
Harness
在训练过程中,我们重构建了Harness,即执行模型动作并在数百个步骤中管理其上下文的循环,利用了来自OSWorld 2.0上代理性能的反馈。代理标记了每个任务失败的原因,工程师审查了他们的修复方案。最大的改动是赋予代理可靠的记忆能力以跟踪数百个步骤,以及在桌面机器本身上的Shell。
Opus 5(70.2%)和GPT-5.6 Sol(66.2%)使用OpenAI发布图表中v2026.08.08离线集的最大努力部分奖励,如成本性能图所示。其他参考分数来自模型卡片和官方排行榜。任务发布、子集和Harness各不相同。
Holotron4 Nano
我们的后训练堆栈旨在适应新的基础模型,并生成跨接口和环境通用的代理。作为NVIDIA Nemotron联盟的成员,我们将最新的堆栈应用于Nemotron 3 Nano Omni模型,作为Holotron 3的后续跟进。
相同的配方将Nemotron 3 Nano Omni转化为Holotron4 Nano,这是一个通用代理模型,在GUI工作流程以及暴露MCP、API或代码沙箱的环境中,显著优于基础模型。
收益是相对于Nemotron 3 Nano Omni的绝对百分点提升。
这些成果表明,我们的方法具有良好的可迁移性,能够将通用模型转化为智能体专家。该方法并不依赖于特定的模型规模。
自行运行
这两种规模的模型今日均已上线 H Models API。权重文件已发布在 Hugging Face 上,提供 BF16、FP8、NVFP4 以及 4-bit GGUF 格式,与我们的小型模型 Holotron4 Nano 一同可供下载。
我们将在未来几天内发布经过优化的 DSpark drafter 检查点,以进一步加速推理过程。
Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts. Both are available on the H Models API. We are also releasing an updated version of Holotron 3: Holotron4 Nano.
Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs. It scores well on academic benchmarks, but we built it for real business workflows. It was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by our Agentic Task Factory.
Get started now:
Models built for every interface
Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools. It uses whichever fits the task. Most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API. Real work is not siloed that way, and a single business task can require combining these different approaches.
Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs. It is the same model in each case and it is called the same way. You do not need to select a different model for each platform.
Holo4 models improve significantly over their Qwen base. Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%. However, it does so with orders of magnitude fewer parameters and at a much lower cost. We open-source every trajectory behind our scores on public benchmarks: replay each step at trajectories.hcompany.ai or download them from Hugging Face .
Competitive with the frontier, at a fraction of the cost
On the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench), Holo4 competes with frontier models at a much lower cost per task.
Notes on the cost-performance charts
OSWorld 2.0. Costs are estimated from the input and output tokens of each agentic run. Holo4 is priced at H Models API rates (single run). Qwen3.8 27B: model card score, cost from the tokens of our run at Alibaba Cloud list prices. Qwen3.6 35B-A3B: single run in our harness, at Alibaba Cloud list prices with cache hits at 20% of the input price. OpenAI launch data supplies the GPT and Opus effort sweeps; other closed and open-weight points use the official OSWorld 2.0 leaderboard . Releases, harnesses and task subsets differ. The line connects non-dominated score and cost pairs among the closed models; Holo4 is excluded.
AutomationBench. Holo4, Qwen3.8 27B and Qwen3.6 35B-A3B: AutomationBench v1.0.6, scores and costs measured in our internal harness. Other models: public-set scores from the AutomationBench README , cost per task from the official leaderboard , which runs on the private set. We will report Holo4 on the private set once it is evaluated.
AI that does work
Trained on environments and tasks from our Agentic Task Factory, Holo4 models excel on professional software. The examples below show Holo4 27B alongside Qwen3.8 27B, its base model. Same prompt and harness for both models.
3D modeling · Eiffel tower
Build a 3D model of the Eiffel Tower in FreeCAD, at a scale of 1 mm to 1 metre, centred on the origin and aligned to the X and Y axes. Work to this design.
The tower is square in plan at every height, never round. Its half-width, measured from the central axis out to the corner, is 62.5 mm at ground level, 32.5 mm at height 57, 17.5 mm at height 115, and 9.35 mm at height 276. Between those heights the half-width follows a smooth curve that falls steeply near the ground and gently higher up, never a straight line.
Four identical legs, one per quadrant, each a square column whose outer corner follows that profile. Each leg is 14 mm across at the ground and tapers to 4 mm at height 276. The legs stand apart from the ground up to the first platform, and converge as they rise. Nothing fills the space between them: the tower is open, and you can see straight through it from every side.
Three platforms, each a solid square slab centred on the axis: 72 mm across and 4 mm thick at height 57; 40 mm across and 3 mm thick at height 115; 22 mm across and 3 mm thick at height 276.
A mast from height 276 to 324, square, 8 mm across at its base tapering to 2 mm at the tip.
Every part must be a closed solid with non-zero volume, and no part may fill the space between the legs.
Holo4 27B (84 calls, 1.3M tokens)
Qwen3.8 27B (60 calls, 1.0M tokens)
3D modeling · H logo
Build a 3D model in FreeCAD of the H company logo: a solid filled disc beside a blocky sans-serif capital letter H, both extruded to the same thickness, the two shapes of similar height and set apart so they do not overlap, with the centre of the disc level with the middle of the H. Colour both shapes black.
Holo4 27B (94 calls, 1.5M tokens)
Qwen3.8 27B (118 calls, 1.9M tokens)
Game design · Pac-Man
Build a Pac-Man-style game in Godot and leave it running.
A rectangular maze of walls laid out on a grid, with pellets filling every open corridor. A player marker moves continuously along the corridors, eating each pellet it passes over and scoring a point for it. Three ghosts move through the same corridors and chase the player. If a ghost catches the player, the player loses a life and everything resets to its starting position. Score and lives are drawn on screen.
No one is going to play this. The player drives itself with a simple heuristic: at each junction it heads toward the nearest pellet, unless a ghost is close, in which case it moves away from the ghost. The game must run unattended and indefinitely, with no keyboard input at all.
When it works, start the game and leave it playing.
Holo4 27B (68 calls, 2.4M tokens, 268 lines)
Qwen3.8 27B (197 calls, 11.4M tokens, 327 lines)
How we built Holo4
Agentic task factory
Our internal set of agentic pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.
Training
Harness
Alongside training, we rebuilt our harness, the loop that executes the model's actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0. Agents tagged why each task failed and engineers reviewed their fixes. The largest changes were giving the agent a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself.
Opus 5 (70.2%) and GPT-5.6 Sol (66.2%) use max-effort partial rewards on the v2026.08.08 offline set from OpenAI's launch chart , as in the cost-performance plot. Other reference scores come from model cards and the official leaderboard. Task releases, subsets and harnesses vary.
Holotron4 Nano
Our post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments. As a member of the NVIDIA Nemotron Coalition , we applied our latest stack to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.
The same recipe turns Nemotron 3 Nano Omni into Holotron4 Nano, a generalist agentic model that significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes.
Gains are absolute percentage-point improvements over Nemotron 3 Nano Omni.
These gains show that our recipe transfers well and can turn a generalist model into an agentic expert. Nothing in it is size-specific.
Run it yourself
Both sizes are available today on the H Models API . Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to our small model, Holotron4 Nano .
We will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.
首次收录 · 2026-09-29 · 10.35 分