亚马逊云科技(Amazon Web Services)发布了一个开源决策模型,其灵感来源于 TypeSafe 的 Jev。随着人工智能开发者越来越寻求比前沿大语言模型更适用于计算机自动化的智能能力,这一趋势日益明显。
在 OpenAI 宣布推出类似产品的一周后,亚马逊发布了 Strands Decider 2B。这是一种高速、低成本的方式,用于在预先决定的选项之间进行排序,并提供对其选择置信度的衡量。该模型完全开源,现已可用,且体积足够小,可以在本地运行。
亚马逊杰出工程师 Marc Brooker 在看到 Jev 并尝试构建自己的类似模型后提出了该项目。这个自制项目取得了足够的成功——它曾短暂地在同类大小的模型 Jevbench 排名中位居榜首——以至于亚马逊工程师对其进行了整理,并将其作为其 Strands Labs(一个致力于开发用于部署 AI 代理的新工具和协议的组织)的一项产品发布。
Brooker 表示,这种工具的需求源于与 AWS 客户的对话,这些客户的智能体工作流程并不总是需要全功能大语言模型的能力或成本。
“最初让我对这个类别的模型产生兴趣的是,它们成为工作流步骤的完美决策者——‘基于我目前的位置,我接下来该做什么?’”Brooker 告诉 TechCrunch。他表示,这为客户提供了“一种可以通过置信度分数和答案的封闭领域以更具可靠性的方式结构化的工作流步骤,[并且]延迟更低,潜在成本也更低。”
与其他决策模型一样,Strands Decider 建立在大型语言模型的“躯干”之上,在本例中是 Qwen3.5-2B,但它不生成文本,而是提供校准后的选择。TypeSafe 将其模型命名为 Jev,以纪念经济学家威廉·斯坦利·杰文斯(William Stanley Jevons),希望唤起他的理论,即某物成本的下降——比如计算机智能的成本——实际上会增加其需求。
自 TypeSafe 推出其理念以来,研究人员已经推出了数十个类似的模型,这一事实表明了广泛的兴趣,但也引发了关于它们价值几何的问题。Brooker 表示,挑战将在于优化模型的快速决策能力,同时不损害其智能。
“需要找到一个非常微妙的平衡点,既要推动其在准确性以及这类任务的校准方面的性能,又不能降低其在理解不同语言、拥有现有知识方面的表现,而这些正是使其成为通用、有趣且有用的原因。”他告诉 TechCrunch。
尽管如此,他并不一定期望前沿实验室会主导这一领域,特别是因为随着市场规模较小,构建有趣产品的成本仅在数百或数千美元之间。
至于 TypeSafe 的高管们则表示,他们正埋头苦干,改进未来的模型。
“我理解人们认为这是一场淘金热,但他们可能低估了让模型真正变聪明的难度。”首席执行官兼创始人 Diogo Almeida 告诉 TechCrunch,并表示目前他尚未看到针对其公司的真正竞争出现。
“目前的这批产品看起来更像是机器学习人员想要实现一种酷炫的架构,而不是一个致力于使智能变得有用的团队。”
Amazon Web Services released an open source decision model inspired by TypeSafe’s Jev , with AI developers increasingly seeking intelligence that is more suited to computer automation than frontier LLMs.
Amazon’s Strands Decider 2B, released the same week OpenAI announced a similar offering , is a high-speed, low-cost way to sort between pre-decided options and deliver a measure of how confident it is in its choice. The model is fully open sourced, available now, and small enough to run locally.
Amazon distinguished engineer Marc Brooker came up with the project after seeing Jev and trying to build his own take on such a model. The homebrew project was successful enough — it briefly reached the top spot on the Jevbench ranking for models of its size — that Amazon engineers cleaned it up and released it as an offering from their Strands Labs , an organization developing new tools and protocols for deploying AI agents.
Brooker says the need for a tool like this emerged in conversations with AWS customers, whose agentic workflows didn’t always require the capability or cost of a fully featured LLM all the time.
“What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step — ‘what is the next thing for me to do here, based on where I am?’” Brooker told TechCrunch. He said it offers customers “a workflow step that can be structured in a way that is more reliable, thanks to the confidence scores, thanks to the closed domain of answers, [and is] lower latency, potentially lower cost.”
Like other decision models, Strands Decider is built on the “torso” of an LLM, in this case Qwen3.5-2B, but instead of generating text, it delivers calibrated choices. TypeSafe named their model Jev after the economist William Stanley Jevons, with hopes of invoking his theory that the falling cost of something — like computer intelligence — can, in fact, increase its demand.
The fact that dozens of similar models have been produced by researchers since TypeSafe debuted its idea shows the wide interest, but also raises the question of how valuable they can be. Brooker suggests that the challenge will be in optimizing the model’s speedy decision-making without compromising its intelligence.
“There is a very careful balance to be found where you want to push its performance on accuracy and calibration on these kinds of tasks, without degrading its performance on understanding different languages, on having the kind of knowledge it has, which is what makes it general purpose and interesting and useful,” he told TechCrunch.
Still, he doesn’t necessarily expect the frontier labs to dominate the space, especially since, with smaller markets, the cost to build something interesting is in the hundreds or thousands of dollars.
For their part, TypeSafe executives say they are keeping their heads down and improving future models.
“I get that people think it’s a gold rush, but they might be underestimating the difficulty of making the models actually smart,” CEO and founder Diogo Almeida told TechCrunch, saying that for now, he didn’t see real competition for his company emerging yet.
“The current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful.”
首次收录 · 2026-10-02 · 12.69 分