在您的工作流中,您需要不止一个模型。一个较低成本的模型处理常规调用,而一个更强的模型处理复杂任务。LangChain 和 CrewAI 等智能体框架是构建此类系统的一种方式。两者都提供了编排层,但组织方式不同。LangChain 的运行时 LangGraph 为您提供显式的状态和控制流。CrewAI 则将工作组织为基于角色的智能体和事件驱动的流程。
当您正在构建一个能够规划、保持状态、调用工具或委派工作的智能体时,这两个框架都能提供帮助。如果您只需要为每个步骤选择一个模型,并在某个模型失败时进行回退,那么完整的编排框架会引入您不需要的部分。
本文根据每项工作所承担的职责,对比了 LangChain 和 LangGraph、CrewAI 以及 OpenRouter 原生路由。它展示了针对 OpenRouter 直接编写和通过 LangChain 编写的相同两步管道,描述了我们的 Agent SDK 如何介于两者之间,并展示了当您需要同时具备编排和路由功能时,如何将 OpenRouter 置于任一框架之下。
简而言之
多模型编排包含三个层次。工作流编排涉及规划、状态、记忆和委派。模型路由是为调用选择模型,并在返回错误时进行回退。提供商路由是选择由哪个提供商端点来服务您已选择的模型。
LangGraph 和 CrewAI 执行工作流编排。OpenRouter 执行模型路由和提供商路由。两者互不替代。
OpenRouter 的 models 参数是一个有序的回退列表。当第一个模型返回错误时,我们会尝试下一个。回退是基于错误的,并不评判答案质量。
我们的 Agent SDK 涵盖了中间情况,即带有验证、流式传输和停止条件的有界多轮工具循环,无需持久化图或基于角色的团队。
LangChain 拥有专用的 ChatOpenRouter 集成,而 CrewAI 通过其 LLM 类将 OpenRouter 记录为提供商。您可以保留任一框架的编排功能,并将 OpenRouter 用作底层的模型层。
三个层次被统称为一个名字
通过订阅,您同意接收 OpenRouter 通讯:每月约一封邮件,包含模型使用数据、产品更新和研究报告。您可以通过每封电子邮件中的链接随时取消订阅。请参阅我们的隐私政策。
“多模型编排”这一短语涵盖了三个不同的决策。将它们分开可以使框架对比更简洁。
工作流编排涉及规划、状态、记忆和委派。您将任务分解为步骤,在轮次之间和运行之间保持状态,暂停以进行人工审查,并将子任务交给其他智能体。LangGraph 和 CrewAI 就是为这一层而构建的。
模型路由是选择哪个模型处理给定的调用,以及当该模型返回错误时会发生什么。OpenRouter 的 models 参数在一个请求字段中处理此问题。
提供商路由是选择由哪个提供商端点来服务您已经选定的模型。OpenRouter 上的许多模型由多个提供商提供服务。我们为每个请求在符合条件的提供商中进行选择,对于包含工具调用的请求,Auto Exacto 会根据工具调用性能重新排列这些提供商。提供商路由永远不会改变您请求的模型。
一个工作流可能需要这三个层次。规划逻辑决定下一步做什么,模型路由决定哪个模型执行该操作,提供商路由决定哪个端点服务该模型。您不需要第一层来获取后两层。使用 models 列表和一些 if 语句即可在模型之间进行路由,而无需智能体框架。
LangChain 和 LangGraph
LangChain 当前的文档描述了具有不同角色的两个产品。LangChain 是智能体框架,提供用于模型、工具和智能体循环的抽象和集成。LangGraph 是其底层的低级编排运行时,专注于持久化执行、流式传输、人在回路和持久性。您可以不使用 LangChain 而单独使用 LangGraph,且 LangChain 的预构建智能体运行在 LangGraph 之上。
LangGraph 将工作流建模为节点的图。您可以在同一个图中混合使用确定性的、手工编码的步骤与由模型驱动的步骤。持久化功能来自两个组件。检查点保存器(checkpointer)会为线程保存图的状态,从而为您提供对话连续性、容错性、时间旅行能力以及人工审查的基础。存储层(store)则在图状态之外持久化应用程序数据,以实现长期的、跨线程的记忆。interrupt() 函数会在节点的任意位置暂停运行,通过检查点保存器保存状态,并等待您使用 Command 恢复它,以便在运行继续之前,由人工批准、编辑或拒绝某个步骤。
这种控制权的代价是您需要自行描述控制流。每个节点、边、状态字段、检查点保存器和中断都是您需要编写和维护的代码。如果您需要一个多步骤管道,其中每个节点都可检查且可恢复,LangGraph 为您提供了结构支持。而如果您只需要一个带有回退机制的模型调用,那就超出了任务所需。
CrewAI
CrewAI 的文档描述了两个构建块。Crews(团队)是由智能体组成的团队,每个智能体都定义了角色、目标和背景故事,它们通过分配的任务开展工作。Crew 可以按顺序执行任务,其中每个任务的输出成为下一个任务的上下文;也可以按分层过程执行,由管理器模型或管理器智能体分配和协调任务。当启用 allow_delegation(允许委托)时,智能体可以相互委托任务,每个智能体都有一个 max_iter(最大迭代次数)限制,默认为 20,以及一个可选的 max_execution_time(最大执行时间)。
Flows(流程)是围绕 Crews 构建的结构化、事件驱动层。Flow 定义了步骤、在它们之间移动的状态以及控制流,包括条件逻辑、循环和分支。每个 Flow 实例都携带一个具有唯一 ID 的状态对象,该对象在运行期间持久存在。CrewAI 的介绍将 Flows 定位为应用程序的骨干,而 Crews 则是 Flow 内的工作单元。
CrewAI 的模式更接近于任务描述而非图定义。您将更多的精力投入到角色、目标和任务字符串上,而较少的精力用于连接节点。这与 LangGraph 是不同的权衡取舍,而不是它的一个较小版本。Crews 留给智能体决定如何完成任务的空间,而 Flows 则是您重新夺回控制权的地方。
比较
LangChain 和 LangGraph CrewAI OpenRouter direct
基于图的编排,具有显式状态、持久化和人工审查 事件驱动流程内的角色型智能体团队 每次调用选择一个模型、错误驱动的 fallback(回退)以及提供商路由
多模型支持 是,每个节点或智能体一个模型对象 是,每个智能体、Crew 或管理器一个 LLM 是,每个请求一个 models 列表
规划、记忆和委托 是,在图、检查点保存器和存储中显式定义 是,通过智能体、流程和 Flow 状态实现 否,仅路由
流式传输 是 是,在 Crew 级别使用 stream=True 是,每个请求
人工审查 interrupt() 配合检查点保存器 您编写的 Flow 逻辑 未提供
您需要编写的内容 节点、边、状态模式、持久化配置 智能体、任务、Crew 和 Flow 定义 请求体
OpenRouter 原生路由
这里的“原生”意味着没有框架。您将请求发送到 https://openrouter.ai/api/v1/chat/completions ,在 models 参数中以优先级顺序列出您的模型,其余的由我们处理。如果第一个模型返回错误,我们会尝试列表中的下一个模型。默认情况下,任何错误都可能触发回退,包括上下文长度验证错误、过滤模型的审核标记、速率限制和停机时间。如果回退模型也返回错误,我们将返回该错误。我们根据最终提供服务的模型对请求进行定价,response model 字段会告诉您具体是哪一个。
Fallback(回退)是对错误的反应。它不会评估第一个模型的答案是否良好。如果您希望更强的模型审查较弱模型的输出,那是您代码中的第二步,而不是 models 列表为您做的事情。
考虑一个两步流水线:先用一个模型起草,再用另一个模型审查草稿,如果第一个模型返回错误,则在每一步都进行回退。直接针对 OpenRouter 编写代码时,这涉及一个函数和两个模型列表。
import os
import requests
def route (models: list[ str ], prompt: str ) -> tuple[ str , str ]:
response = requests.post(
"https://openrouter.ai/api/v1/chat/completions" ,
headers = { "Authorization" : f "Bearer { os.environ[ 'OPENROUTER_API_KEY' ] } " },
json = { "models" : models, "messages" : [{ "role" : "user" , "content" : prompt}]},
timeout = 120 ,
)
response.raise_for_status()
body = response.json()
return body[ "choices" ][ 0 ][ "message" ][ "content" ], body[ "model" ]
draft, draft_model = route(
[ "anthropic/claude-sonnet-5" , "openai/gpt-5.6-sol" ],
"Draft a one-paragraph summary of what a model fallback list does." ,
)
review, review_model = route(
[ "openai/gpt-5.6-sol" , "anthropic/claude-sonnet-5" ],
f "Review this draft for accuracy and suggest one improvement: \n\n{ draft } " ,
)
print ( f "draft by { draft_model } , review by { review_model } " )
print (review)
要将步骤发送到不同的模型,只需更改列表即可。由于我们的 API 与 OpenAI 兼容,相同的聊天补全请求结构适用于目录中的每个聊天模型,而模型变更仅是字符串的变更。服务于其他端点(如嵌入、视频、文本转语音或语音转文本)的模型则使用各自独立的请求结构。
通过 LangChain 实现的相同流水线使用了专用的 ChatOpenRouter 模型类。你需要为每个步骤创建一个模型对象,使用 LangChain 的 with_fallbacks 包装每个对象,以便在调用失败时重试下一个模型对象,并通过 invoke 调用结果。
from langchain_openrouter import ChatOpenRouter
drafter = ChatOpenRouter( model = "anthropic/claude-sonnet-5" ).with_fallbacks(
[ChatOpenRouter( model = "openai/gpt-5.6-sol" )]
)
reviewer = ChatOpenRouter( model = "openai/gpt-5.6-sol" ).with_fallbacks(
[ChatOpenRouter( model = "anthropic/claude-sonnet-5" )]
)
draft = drafter.invoke( "Draft a one-paragraph summary of what a model fallback list does." )
review = reviewer.invoke(
f "Review this draft for accuracy and suggest one improvement: \n\n{ draft.content } "
)
print (review.content)
这两个版本都在相同的两个模型之间路由相同的两个步骤,且具有相同的回退顺序。区别在于回退运行的位置。在直接版本中,我们在单个请求的服务器端运行它。在 LangChain 版本中,框架在你的进程中捕获失败的调用并发送第二个请求。对于使用 LangChain 的 create_agent 构建的智能体,框架还提供了 ModelFallbackMiddleware,当主模型失败时尝试备用模型。框架版本为你提供模型对象和共享的 invoke 接口,这是在存在图、检查点或一组需要管理的工具时你所需的结构。如果没有这些,它就是一种开销。
无论哪种版本,通过我们进行路由都会带来三件事。每个响应都包含一个 usage 对象,其中带有令牌计数和以积分计算的成本,无需额外参数,因此你可以在决定是否值得采用分层设置之前看到每次调用的成本。对支持模型的提示缓存可降低重复上下文的成本。当请求包含工具时,Auto Exacto 默认运行。它根据吞吐量、工具调用成功率和基准数据重新排序你所选模型的服务提供商,使工具调用落在具有强大工具调用记录的服务提供商上,而无需你在侧进行配置。Auto Exacto 改变的是服务提供商顺序,而非模型本身。
直接路由不规划、不在多轮对话中保持状态,也不决定哪个子任务分配给哪个智能体。那是工作流编排层,而模型列表并不提供它。下一步不一定是完整的框架。
OpenRouter Agent SDK
在直接路由和完整框架之间,是我们的 Agent SDK(即 @openrouter/agent 包)。聊天补全(chat completion)是无状态的:你发送消息并得到一个回复。将其转变为智能体(agent)意味着运行一个循环:模型请求调用工具,你的代码验证参数并执行该工具,结果返回给模型,然后循环重复直到工作完成。Agent SDK 将这一循环封装为一个 callModel 函数。
你使用 tool() 辅助函数和 Zod schema 定义工具,SDK 则处理跨轮次的验证、执行和对话状态。停止条件(如 stepCountIs 和 maxCost)限制了循环的范围。每个条件在步骤完成后进行检查,因此 maxCost 会在达到阈值的那个步骤之后停止循环,而不是阻止该步骤;默认情况下,SDK 随后会再进行一次模型调用以生成最终答案。请将 maxCost 视为一条停止规则,而非支出上限。流式传输(Streaming)已内置支持,你可以将远程 MCP 服务器作为工具源接入。
import { OpenRouter, tool, stepCountIs, maxCost } from "@openrouter/agent" ;
import { z } from "zod" ;
const client = new OpenRouter ({ apiKey: process.env. OPENROUTER_API_KEY });
const result = client. callModel ({
model: "anthropic/claude-sonnet-5" ,
input: "What time is it in Tokyo?" ,
tools: [
tool ({
name: "get_time" ,
description: "Get the current time in a timezone" ,
inputSchema: z. object ({ timezone: z. string () }),
execute : async ({ timezone }) => ({
time: new Date (). toLocaleString ( "en-US" , { timeZone: timezone }),
}),
}),
],
stopWhen: [ stepCountIs ( 5 ), maxCost ( 0.5 )],
});
const text = await result. getText ();
console. log (text);
该 SDK 使用 TypeScript 编写,Python 和 Go 端口保持同步。它为你提供带有验证、流式传输和停止条件的有界工具循环。但它不提供持久化的图结构、检查点或基于角色的团队功能。
对于单个请求内的委托,openrouter:subagent 服务器工具允许模型在生成过程中将一个自包含的任务交给工作模型(worker model)。工作模型可以是 OpenRouter 上的任何模型。每个任务都是独立的。工作模型仅能看到任务描述,且在任务之间不保留任何记忆。服务器工具目前处于测试阶段,API 和行为可能会发生变化。
该 SDK 还涵盖了循环内的人工审批和持久化的对话状态。工具可以设置 requireApproval 以在运行前暂停,而 StateAccessor 则在 callModel 调用之间持久化消息、审批和工具结果。它不提供带有显式转换和检查点的持久化图结构,也不提供跨智能体团队的协调功能。当你需要这些功能时,就是升级到 LangGraph 或 CrewAI 的时候了。
在 LangChain 或 CrewAI 底层使用 OpenRouter
你不必在框架和网关之间做出选择。你可以保留框架的计划、状态和委托功能,并将 OpenRouter 用作其底层的模型层。
对于 LangChain,我们维护着一个专门的集成。langchain-openrouter(Python)和 @langchain/openrouter(JavaScript)包为你提供 ChatOpenRouter 模型,你可以将智能体和图指向该模型。LangChain 的文档目前将 Python 集成标记为测试版。
from langchain_openrouter import ChatOpenRouter
model = ChatOpenRouter(
model = "anthropic/claude-sonnet-5" ,
temperature = 0 ,
model_kwargs = { "models" : [ "anthropic/claude-sonnet-5" , "openai/gpt-5.6-sol" ]},
)
ChatOpenRouter 对象发送一个 model 值。要在 LangChain 下使用我们的服务端回退功能,请通过 model_kwargs 传递 models 列表,该包会将其展开到请求体中。如果没有这样做,回退将是框架的职责,如上文 with_fallbacks 示例所示。
对于 CrewAI,你使用我们的端点配置其 LLM 类。CrewAI 的 LLM 文档将 OpenRouter 列为使用 LiteLLM 的提供商之一,因此你需要安装 crewai[litellm] 扩展包,在模型 slug 前加上 openrouter/ 前缀,并传入我们的基础 URL 和你的密钥。
import os
from crewai import LLM
llm = LLM(
model = "openrouter/anthropic/claude-sonnet-5" ,
base_url = "https://openrouter.ai/api/v1" ,
api_key = os.env
You need more than one model in a workflow. One lower-cost model handles routine calls and a stronger model handles the difficult ones. Agent frameworks such as LangChain and CrewAI are one way to build that. Both provide an orchestration layer, and they organize it differently. LangChain’s runtime, LangGraph, gives you explicit state and control flow. CrewAI organizes work as role-based agents and event-driven flows.
Both frameworks help when you’re building an agent that plans, keeps state, calls tools, or delegates work. If you only need to choose a model for each step and fall back when one fails, a full orchestration framework adds parts you won’t use.
This article compares LangChain and LangGraph, CrewAI, and OpenRouter-native routing by the job each one does. It shows the same two-step pipeline written against OpenRouter directly and through LangChain, describes where our Agent SDK sits between the two, and shows how to put OpenRouter underneath either framework when you need both orchestration and routing.
Tl;dr
Multi-model orchestration is three layers. Workflow orchestration is planning, state, memory, and delegation. Model routing is choosing a model for a call and falling back when it returns an error. Provider routing is choosing which provider endpoint serves the model you chose.
LangGraph and CrewAI do workflow orchestration. OpenRouter does model routing and provider routing. Neither replaces the other.
The OpenRouter models parameter is an ordered fallback list. When the first model returns an error, we try the next one. Fallback is error-driven and does not judge answer quality.
Our Agent SDK covers the middle case, a bounded multi-turn tool loop with validation, streaming, and stop conditions, without a durable graph or a role-based crew.
LangChain has a dedicated ChatOpenRouter integration, and CrewAI documents OpenRouter as a provider through its LLM class. You can keep either framework’s orchestration and use OpenRouter as the model layer underneath.
Three layers that get called one name
By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy .
The phrase “multi-model orchestration” covers three different decisions. Separating them makes the framework comparison shorter.
Workflow orchestration is planning, state, memory, and delegation. You break a task into steps, keep state across turns and across runs, pause for human review, and hand subtasks to other agents. LangGraph and CrewAI are built for this layer.
Model routing is choosing which model handles a given call, and what happens when that model returns an error. The OpenRouter models parameter handles this in one request field.
Provider routing is choosing which provider endpoint serves the model you already chose. Many models on OpenRouter are served by more than one provider. We select among the eligible providers for each request, and on requests that include tools, Auto Exacto reorders those providers by tool-calling performance. Provider routing never changes which model you asked for.
A workflow can need all three. Planning logic decides what to do next, model routing decides which model does it, and provider routing decides which endpoint serves that model. You don’t need the first layer to get the second two. A models list and a few if statements route between models without an agent framework.
LangChain and LangGraph
LangChain’s current documentation describes two products with different roles. LangChain is the agent framework, with abstractions and integrations for models, tools, and agent loops. LangGraph is the low-level orchestration runtime underneath it, focused on durable execution, streaming, human-in-the-loop, and persistence. You can use LangGraph without LangChain, and LangChain’s prebuilt agents run on LangGraph.
LangGraph models a workflow as a graph of nodes. You can mix deterministic, hand-coded steps with model-driven steps in the same graph. Persistence comes from two components. A checkpointer saves the graph state for a thread, which gives you conversation continuity, fault tolerance, time travel, and the foundation for human review. A store persists application data outside the graph state for long-term, cross-thread memory. The interrupt() function pauses a run at any point in a node, saves state through the checkpointer, and waits for you to resume it with a Command , so a person can approve, edit, or reject a step before the run continues.
The cost of that control is that you describe the control flow yourself. Each node, edge, state field, checkpointer, and interrupt is code you write and maintain. If you need a multi-step pipeline where every node is inspectable and resumable, LangGraph gives you the structure. If you need a model call with a fallback, it’s more than the job requires.
CrewAI
CrewAI’s documentation describes two building blocks. Crews are teams of agents, each defined with a role, a goal, and a backstory, that work through assigned tasks. A crew runs tasks in a sequential process, where each task’s output becomes context for the next, or in a hierarchical process, where a manager model or manager agent assigns and coordinates the tasks. Agents can delegate to each other when allow_delegation is enabled, and each agent has a max_iter limit, which defaults to 20, and an optional max_execution_time .
Flows are the structured, event-driven layer around crews. A flow defines the steps, the state that moves between them, and the control flow, including conditional logic, loops, and branching. Each flow instance carries a state object with a unique ID that persists for the run. CrewAI’s introduction positions flows as the backbone of an application and crews as the units of work inside a flow.
CrewAI’s model is closer to a task description than a graph definition. You spend more of your effort on role, goal, and task strings and less on connecting nodes. That’s a different tradeoff from LangGraph rather than a smaller version of it. Crews leave the agents room to decide how to complete a task, and flows are where you take that control back.
Comparison
LangChain and LangGraph CrewAI OpenRouter direct
Built for Graph-based orchestration with explicit state, persistence, and human review Role-based agent teams inside event-driven flows Choosing a model per call, error-driven fallback, and provider routing
Multi-model support Yes, one model object per node or agent Yes, one LLM per agent, crew, or manager Yes, a models list per request
Planning, memory, and delegation Yes, explicit in the graph, checkpointer, and store Yes, through agents, processes, and flow state No, routing only
Streaming Yes Yes, at the crew level with stream=True Yes, per request
Human review interrupt() with a checkpointer Flow logic you write Not provided
What you write Nodes, edges, state schema, persistence config Agent, task, crew, and flow definitions A request body
OpenRouter-native routing
“Native” here means no framework. You send a request to https://openrouter.ai/api/v1/chat/completions , list your models in priority order in the models parameter, and we handle the rest. If the first model returns an error, we try the next model in the list. By default any error can trigger a fallback, including context length validation errors, moderation flags for filtered models, rate limiting, and downtime. If the fallback model also returns an error, we return that error. We price the request using the model that ultimately served it, and the response model field tells you which one that was.
Fallback reacts to errors. It doesn’t evaluate whether the first model’s answer was good. If you want a stronger model to review a weaker model’s output, that’s a second step in your code, not something the models list does for you.
Take a two-step pipeline that drafts with one model and then reviews the draft with a different one, falling back at each step if the first model returns an error. Written against OpenRouter directly, that’s one function and two models lists.
import os
import requests
def route (models: list[ str ], prompt: str ) -> tuple[ str , str ]:
response = requests.post(
"https://openrouter.ai/api/v1/chat/completions" ,
headers = { "Authorization" : f "Bearer { os.environ[ 'OPENROUTER_API_KEY' ] } " },
json = { "models" : models, "messages" : [{ "role" : "user" , "content" : prompt}]},
timeout = 120 ,
)
response.raise_for_status()
body = response.json()
return body[ "choices" ][ 0 ][ "message" ][ "content" ], body[ "model" ]
draft, draft_model = route(
[ "anthropic/claude-sonnet-5" , "openai/gpt-5.6-sol" ],
"Draft a one-paragraph summary of what a model fallback list does." ,
)
review, review_model = route(
[ "openai/gpt-5.6-sol" , "anthropic/claude-sonnet-5" ],
f "Review this draft for accuracy and suggest one improvement: \n\n{ draft } " ,
)
print ( f "draft by { draft_model } , review by { review_model } " )
print (review)
To send a step to a different model, you change the list. Because our API is OpenAI-compatible, the same chat completions request shape works for every chat model in the catalog, and a model change is a change to one string. Models that serve other endpoints, such as embeddings, video, text to speech, or speech to text, use their own request shapes.
The same pipeline through LangChain uses the dedicated ChatOpenRouter model class. You create a model object per step, wrap each one with LangChain’s with_fallbacks so a failed call retries on the next model object, and call the result through invoke .
from langchain_openrouter import ChatOpenRouter
drafter = ChatOpenRouter( model = "anthropic/claude-sonnet-5" ).with_fallbacks(
[ChatOpenRouter( model = "openai/gpt-5.6-sol" )]
)
reviewer = ChatOpenRouter( model = "openai/gpt-5.6-sol" ).with_fallbacks(
[ChatOpenRouter( model = "anthropic/claude-sonnet-5" )]
)
draft = drafter.invoke( "Draft a one-paragraph summary of what a model fallback list does." )
review = reviewer.invoke(
f "Review this draft for accuracy and suggest one improvement: \n\n{ draft.content } "
)
print (review.content)
Both versions route the same two steps across the same two models with the same fallback order. The difference is where the fallback runs. In the direct version, we run it server-side within one request. In the LangChain version, the framework catches the failed call in your process and sends a second request. For agents built with LangChain’s create_agent , the framework also offers ModelFallbackMiddleware , which tries alternative models when the primary model fails. The framework version gives you model objects and a shared invoke interface, which is the structure you want once there’s a graph, a checkpointer, or a set of tools to manage. When there isn’t, it’s overhead.
Three things come with routing through us in either version. Every response includes a usage object with token counts and the cost in credits, with no extra parameter required, so you can see what each call cost before deciding whether a tiered setup is worth it. Prompt caching on supported models reduces the cost of repeated context. And when a request includes tools, Auto Exacto runs by default. It reorders the providers for your chosen model using throughput, tool-calling success rate, and benchmark data, so tool calls land on providers with strong tool-calling records without configuration on your side. Auto Exacto changes the provider order, not the model.
Direct routing doesn’t plan, keep state across turns, or decide which subtask goes to which agent. That’s the workflow orchestration layer, and a models list doesn’t provide it. The next step up isn’t always a full framework.
The OpenRouter Agent SDK
Between direct routing and a full framework sits our Agent SDK , the @openrouter/agent package. A chat completion is stateless. You send messages and get one response. Turning that into an agent means running a loop where the model requests a tool call, your code validates the arguments and runs the tool, the result goes back to the model, and the loop repeats until the work is done. The Agent SDK packages that loop into one callModel function.
You define tools with the tool() helper and a Zod schema, and the SDK handles validation, execution, and conversation state across turns. Stop conditions such as stepCountIs and maxCost bound the loop. Each condition is checked after a step completes, so maxCost stops the loop after the step that reaches the threshold rather than preventing that step, and by default the SDK then makes one more model turn to produce a final answer. Treat maxCost as a stopping rule, not as a spending ceiling. Streaming is built in, and you can plug in a remote MCP server as a tool source.
import { OpenRouter, tool, stepCountIs, maxCost } from "@openrouter/agent" ;
import { z } from "zod" ;
const client = new OpenRouter ({ apiKey: process.env. OPENROUTER_API_KEY });
const result = client. callModel ({
model: "anthropic/claude-sonnet-5" ,
input: "What time is it in Tokyo?" ,
tools: [
tool ({
name: "get_time" ,
description: "Get the current time in a timezone" ,
inputSchema: z. object ({ timezone: z. string () }),
execute : async ({ timezone }) => ({
time: new Date (). toLocaleString ( "en-US" , { timeZone: timezone }),
}),
}),
],
stopWhen: [ stepCountIs ( 5 ), maxCost ( 0.5 )],
});
const text = await result. getText ();
console. log (text);
The SDK is written in TypeScript, with Python and Go ports kept in sync. It gives you a bounded tool loop with validation, streaming, and stop conditions. It does not give you a durable graph, a checkpointer, or a role-based crew.
For delegation within a single request, the openrouter:subagent server tool lets a model hand a self-contained task to a worker model mid-generation. The worker can be any model on OpenRouter. Each task is independent. The worker sees only the task description and keeps no memory between tasks. Server tools are in beta, and the API and behavior may change.
The SDK also covers human approval and persisted conversation state within that loop. A tool can set requireApproval to pause before it runs, and a StateAccessor persists messages, approvals, and tool results between callModel invocations. What it doesn’t give you is a durable graph with explicit transitions and checkpoints, or coordination across a team of agents. When you need those, that’s the point to move up to LangGraph or CrewAI.
Using OpenRouter underneath LangChain or CrewAI
You don’t have to choose between a framework and a gateway. You keep the framework’s planning, state, and delegation, and use OpenRouter as the model layer underneath it.
With LangChain, we maintain a dedicated integration . The langchain-openrouter package for Python and the @langchain/openrouter package for JavaScript give you a ChatOpenRouter model that you point your agents and graphs at. LangChain’s documentation currently marks the Python integration as beta.
from langchain_openrouter import ChatOpenRouter
model = ChatOpenRouter(
model = "anthropic/claude-sonnet-5" ,
temperature = 0 ,
model_kwargs = { "models" : [ "anthropic/claude-sonnet-5" , "openai/gpt-5.6-sol" ]},
)
A ChatOpenRouter object sends one model value. To use our server-side fallback under LangChain, pass the models list through model_kwargs , which the package spreads into the request body. Without it, fallback is the framework’s job, as in the with_fallbacks example above.
With CrewAI, you configure its LLM class with our endpoint. CrewAI’s LLM documentation lists OpenRouter as a provider that uses LiteLLM, so you install the crewai[litellm] extra, prefix the model slug with openrouter/ , and pass our base URL and your key.
import os
from crewai import LLM
llm = LLM(
model = "openrouter/anthropic/claude-sonnet-5" ,
base_url = "https://openrouter.ai/api/v1" ,
api_key = os.env
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-04 | 9.22 | 22 | 入选 |
| 2026-10-03 | 9.95 | 28 | 未入选 |