嵌入模型决定了你的检索系统能够找到什么。它将每个输入转化为向量,并将相关的输入放置在彼此附近,从而使你的应用程序能够通过语义而非精确措辞进行检索。
最佳选择取决于你需要搜索的材料类型。英文知识库、多语言支持档案、源代码仓库以及图文集合各有不同的需求。向量维度、上下文长度、公开权重以及价格也会影响决策。
我们针对英文 RAG、多语言检索、代码搜索、图文检索以及低成本索引筛选出了候选模型,然后通过我们的嵌入端点向每个模型发送了实时请求。
最后验证日期:2026 年 9 月 11 日。我们的嵌入模型目录在该日期返回了 37 条记录,其中包括部分模型的批量版本和预览版本。
提供商可以添加或移除路由,且提示价格可能会发生变化。在启动大规模索引或重新索引任务之前,请查看当前的模型页面。如果模型页面与本指南不符,请以模型页面为准。
TL;DR(摘要)
对于英文 RAG,请从 openai/text-embedding-3-small 开始。
当你的输入超过 8,192 个 token,或者你希望在 Voyage 4 不同层级之间迁移而无需重建索引时,请测试 voyageai/voyage-4-large。如果你希望获得具有公开权重的替代方案,qwen/qwen3-embedding-8b 拥有更长的上下文窗口且提示价格更低。
对于具有公开权重的多语言检索,请使用 qwen/qwen3-embedding-8b;对于代码搜索,请使用 voyageai/voyage-code-4;对于图文检索,请使用 google/gemini-embedding-2 或 voyageai/voyage-multimodal-3.5。
在我们筛选的付费文本模型中,perplexity/pplx-embed-v1-0.6b 拥有最低的提示价格。nvidia/nemotron-3-embed-1b:free 是一个免费的文本嵌入路由,提供 32,768 个 token 的上下文窗口。
我们的 API 检查确认的是请求和响应行为,而非检索质量。在构建或重建完整索引之前,请使用来自你应用程序的标记查询和文档对至少两个候选者进行比较。
按用例分类的最佳嵌入模型
通过订阅,你同意接收 OpenRouter 通讯:包括模型使用数据、产品更新和研究报告,大约每周一封电子邮件。随时通过每封电子邮件中的链接退订。请参阅我们的隐私政策。
| 用例 | 推荐模型 | 入选理由 |
|---|---|---|
| 默认英文 RAG | openai/text-embedding-3-small | 低提示价格、8,192-token 上下文以及可调整的输出维度 |
| 超过 8,192 token 的输入 | voyageai/voyage-4-large | 32,000-token 上下文、四种可选维度以及与其它 Voyage 4 层级兼容的嵌入向量 |
| 具有公开权重的多语言检索 | qwen/qwen3-embedding-8b | 支持超过 100 种语言、32,768-token 上下文窗口以及公开权重 |
| 代码搜索 | voyageai/voyage-code-4 | 专为检索代码及相关技术内容而构建 |
| 图文检索 | google/gemini-embedding-2 | 将文本和图像置于同一嵌入空间。voyageai/voyage-multimodal-3.5 是我们验证的第二个选项 |
| 最低价格的付费文本选项 | perplexity/pplx-embed-v1-0.6b | 每百万输入 token 提示价格为 $0.004,拥有 32,000-token 上下文窗口 |
| 免费文本嵌入 | nvidia/nemotron-3-embed-1b:free | 一个免费路由,提供 32,768-token 上下文窗口和多语言模型卡片 |
按输入类型选择模型
从你需要搜索的材料开始。该图表首先按输入类型排序,其次按是否需要公开权重排序,最后按你的主要约束条件排序。
该图表为你提供了一个起点。在构建或重建大型索引之前,请使用来自你应用程序的查询和文档对至少两个候选者进行比较,并跟踪每个模型检索相关片段的情况。
我们如何选择这些模型
我们使用实时目录确认了可用性、上下文窗口、输入类型和提示价格。我们查阅了模型提供方的文档以获取能力说明和基准测试结果,随后通过我们的嵌入端点发送了实时请求,以确认请求和响应行为。
我们向 19 个模型发送了两字符串的批量请求,并总共运行了 28 项检查,涵盖批量输入、可配置维度、图像输入、图文输入以及错误处理。这 16 个付费模型为每个输入返回一个向量。这些响应证实了下表中列出的默认维度,并表明 dimensions 参数适用于 OpenAI Text Embedding 3 Small、Gemini Embedding 2 和 Voyage 4 Large。针对不存在的模型发出的请求返回了 400 错误,消息内容为“Model openai/does-not-exist does not exist”(模型 openai/does-not-exist 不存在)。
三个免费路由从我们的测试账户返回了 404 错误,因为其隐私设置不允许将流量路由至可能会在免费模型提示上进行训练的提供方。我们将在下文“免费选项”部分描述该设置。我们为这些模型列出的默认维度来自其模型卡片,而非我们的响应数据。
这些请求证实了 API 兼容性。它们并未衡量检索质量。我们排除了响应时间,因为每个模型在不同服务条件下仅接收了一个小型请求。我们将每个用例与文档化的能力相匹配,然后利用语言覆盖范围、上下文长度、输出维度、公开权重和提示价格来缩小候选范围。已发布的评估结果(包括 MTEB 和 CoIR)帮助我们确定了需要测试的模型。您标注的检索结果应决定最终部署哪一个模型。
比较入围的嵌入模型
以下价格是我们目录中于 2026 年 9 月 11 日列出的提示价格,Gemini Embedding 2 的图像价格则是其于 2026 年 9 月 18 日的目录价值。默认维度来自我们的实时响应数据,免费 NVIDIA 模型的默认维度除外,其数据来自其模型卡片。
模型 | 输入 | 上下文 | 默认维度 | 每百万 token 的提示价格 | 权重
openai/text-embedding-3-small | 文本 | 8,192 | 1,536 | $0.02 | 闭源
openai/text-embedding-3-large | 文本 | 8,192 | 3,072 | $0.13 | 闭源
voyageai/voyage-4-lite | 文本 | 32,000 | 1,024 | $0.02 | 闭源
voyageai/voyage-4 | 文本 | 32,000 | 1,024 | $0.06 | 闭源
voyageai/voyage-4-large | 文本 | 32,000 | 1,024 | $0.12 | 闭源
voyageai/voyage-code-4 | 文本,针对代码优化 | 32,000 | 1,024 | $0.12 | 闭源
voyageai/voyage-multimodal-3.5 | 文本和图像 | 32,000 | 1,024 | $0.12 | 闭源
qwen/qwen3-embedding-8b | 文本 | 32,768 | 4,096 | $0.01 | 开源
qwen/qwen3-embedding-4b | 文本 | 32,768 | 2,560 | $0.02 | 开源
perplexity/pplx-embed-v1-0.6b | 文本 | 32,000 | 1,024 | $0.004 | 开源
perplexity/pplx-embed-v1-4b | 文本 | 32,000 | 2,560 | $0.03 | 开源
google/gemini-embedding-2 | 文本和图像 | 8,192 | 3,072 | 文本 $0.20,图像 token $0.45 | 闭源
nvidia/nemotron-3-embed-1b:free | 文本 | 32,768 | 2,048 | 免费 | 开源
维度列影响原始索引大小。请参阅下方的“计算向量存储”以获取公式和示例。
我们的目录包含的模型多于表格中所示。baai/bge-m3 提供了另一种开源多语言选项,mistralai/mistral-embed-2312 提供了一般文本替代方案,而 mistralai/codestral-embed-2505 则针对代码检索。在我们的文本请求检查中,这三者均返回了向量。目录还列出了一组来自 BAAI、E5、GTE 和 Sentence Transformers 的 512-token 开源模型,价格为每百万 token $0.005 至 $0.01,我们未在本指南中对此进行测试。
英文 RAG 的最佳默认选择
当您需要用于英文知识库的管理型模型时,请从 openai/text-embedding-3-small 开始。其默认向量包含 1,536 个值,仅为 openai/text-embedding-3-large 默认值的一半。在相同的数值格式下,Small 版本所需的原始向量存储量也减半。建议在首次评估时保持 1,536 值的默认设置,仅当默认设置已满足检索目标但存储或搜索成本仍构成限制时,才考虑降低该维度。
你可以使用 dimensions 参数来缩短向量的维度。我们的请求中设置了 "dimensions": 256,返回了一个包含 256 个值的向量。较小的向量占用更少的数据库存储空间,并减少相似度搜索所需的工作量,但也可能降低检索质量。在更改现有索引之前,请先测试较小的尺寸。
OpenAI 将 text-embedding-3-large 描述为其用于英语和非英语任务的最强大的嵌入模型。当 Small 模型未能返回相关结果时,请对其进行测试,特别是在用户跨语言搜索时。切换到 Large 模型会改变向量大小和提示价格,因此在重新嵌入语料库之前,请在保留的问题集上比较这两个模型。
如果你的输入超过了 Small 模型的 8,192 个 token 限制,请测试 Voyage 4 Large。对于较短的输入,请使用相同的文本块、查询和相关性标签来比较这两个模型。
适用于更长输入的 Voyage 4 系列
当你需要一个可调节维度且兼容所有 Voyage 4 层级的托管 32,000-token 模型时,请测试 voyageai/voyage-4-large。它支持最多 32,000 个 token,并支持 256、512、1,024 和 2,048 种维度。在我们的 API 检查中,将 dimensions 设置为 512 返回了一个包含 512 个值的向量。
该系列还包括 voyageai/voyage-4 和 voyageai/voyage-4-lite,它们使用相同的上下文窗口。Voyage 表示,使用 4 系列创建的所有嵌入彼此兼容。这允许你将另一个 Voyage 4 层级与现有索引进行比较。在部署之前,请在具有代表性的样本上验证层级变更。兼容的向量空间消除了重建的需求,但并不保证检索结果完全相同。
将 Large 模型和至少一个较低层级的 Voyage 4 模型针对同一评估集运行。如果它们在相同的截止点检索到相同的相关段落,则较低层级的模型可能满足你的需求。
Voyage 模型是托管服务。如果你需要用于自托管或部署控制的公开权重,Qwen3 Embedding 8B 是这份短名单中的开源权重替代方案。
多语言检索的最佳开源权重模型
qwen/qwen3-embedding-8b 是我们选择的开源权重多语言模型。它支持超过 100 种语言,你可以从 Hugging Face 下载权重,或通过我们的 API 调用托管模型。我们默认的 API 请求返回了 4,096 个值。在相同的数值格式下,包含 4,096 个值的向量使用的原始存储空间是包含 1,024 个值的向量的四倍。在估算索引大小时请纳入这一差异。我们的 API 检查并未确定 Qwen3 Embedding 8B 的检索效果是否优于 nvidia/nemotron-3-embed-1b:free。当路由成本和向量大小影响决策时,请在做出承诺之前在同一组带标签的数据上测试这两个模型。
Qwen 模型卡报告称,截至 2025 年 6 月 5 日,该模型在大规模文本嵌入基准测试(MTEB)上的多语言得分为 70.58。这是供应商报告的公开基准测试结果。你可以用它来筛选 Qwen,然后测试你的应用程序支持的所有语言和领域,因为你的语料库可能会产生不同的排名。
模型页面列出了 32,768-token 的上下文窗口,但该限制是针对每个端点设置的。在 2026 年 9 月 11 日,Qwen3 Embedding 8B 的 DeepInfra 和 SiliconFlow 端点列出了 32,768 个 token,而 Nebius 端点列出了 32,000 个。如果你发送超过 32,000 个 token 的输入,请检查端点列表并锁定具有更大限制的提供商,或者将输入保持在 32,000 个 token 或以下。
4B 版本使用相同的模型级上下文窗口,在我们的检查中返回了 2,560 个值。它运行所需的资源更少。托管价格取决于可用的提供商,因此在做出选择之前,请比较 Qwen3 Embedding 8B 和 4B 的当前模型页面。
baai/bge-m3 为您提供第二个开源多语言模型,以便在您自己的数据集上与 Qwen3 进行比较。其模型说明涵盖了 100 多种语言,支持高达 8,192 个 token 的输入,而我们的请求返回了一个包含 1,024 个值的向量。上游模型可以返回多种表示类型,包括密集向量和稀疏向量。我们的标准嵌入响应返回的是密集向量。如果您的检索设计依赖于其其他表示类型,请使用上游实现。
代码搜索的最佳嵌入模型
当查询需要检索函数、文件或代码文档片段时,请使用 voyageai/voyage-code-4。Voyage 为该模型专为代码检索和编码代理工作负载而构建。其 32,000 token 的上下文窗口可以接受长文件或长查询,尽管索引块仍应代表在检索时有用的代码单元。
mistralai/codestral-embed-2505 是我们目录中的主要特定于代码的替代方案。在我们的文本请求检查中,它返回了 1,536 个值。其上下文窗口为 8,192 token,而 Voyage Code 4 为 32,000 token。建议先测试 Voyage Code 4,并将 Codestral Embed 用作对比模型。
公开评估可以帮助您比较候选模型。代码信息检索基准(CoIR)包含十个数据集,涵盖七个领域的八项代码检索任务。在使用自己仓库的工作对最强候选模型进行测试之前,请使用它来比较模型。一项有用的内部评估是将问题描述或开发者问题映射到解决它们所需的文件和块。
文本和图像检索的最佳模型
当同一索引需要同时搜索文本和图像时,请使用 google/gemini-embedding-2。它将两种输入类型置于同一个嵌入空间中,因此文本查询可以检索具有相关含义的图像,而图像查询可以检索相关的文本。
我们通过嵌入端点验证了文本、base64 编码的 PNG 图像以及组合的文本和图像请求。每个请求返回一个包含 3,072 个值的向量。我们还发送了一个带有 "dimensions": 768 的文本请求,并收到了一个包含 768 个值的向量。在我们的检查中,一个 64x64 像素的 PNG 计为 258 个提示 token,费用为 $0.000128。我们的目录对该模型的图像输入单独定价,截至 2026 年 9 月 18 日,模型页面上的价格为每百万图像 token $0.45,因此在索引大型图像集合之前请查看模型页面。本指南将推荐限制在文本和图像输入上,因为这是我们验证的请求类型。
voyageai/voyage-multimodal-3.5 是我们验证的第二个文本和图像模型。其图像请求和组合的文本与图像请求均返回一个包含 1,024 个值的向量,同样的 64x64 像素 PNG 计为 89 个提示 token,费用为 $0.00003。它具有 32,000 token 的上下文窗口和每百万 token $0.12 的提示价格。Gemini Embedding 2 和 Voyage Multimodal 3.5 使用不同的向量空间,因此在索引之前请选择其中一个。
google/gemini-embedding-001 是一个单独的仅文本模型,具有 20,000 token 的上下文窗口。这两个 Gemini 模型使用不同的向量空间,因此在这两者之间切换需要您重新嵌入索引。
nvidia/llama-nemotron-embed-vl-1b-v2:free 是我们目录中一个免费的文本和图像选项,具有 131,072 token 的上下文窗口。由于下文所述的隐私设置原因,我们的测试账户无法调用它,因此我们将其排除在推荐之外。
最佳低成本和免费选项
perplexity/pplx-embed-v1-0.6b 在我们的短名单中拥有最低的付费文本价格,为每百万输入 token $0.004。它接受高达 32,000 token 的输入,并返回一个包含 1,024 个值的向量。较大的 perplexity/pplx-embed-v1-4b 返回了 2,560 个值,提示价格为每百万输入 token $0.03。
Perplexity 将 0.6B 模型定位为轻量级、低延迟的检索方案,而将 4B 模型用于实现更高的检索质量。我们对 API 的检查确认了默认的向量维度,但并未验证哪个 Perplexity 模型的检索效果更佳。应将提供商的定位视为候选名单的信号,并在标注查询集上对这两个模型进行比较。
在评估阶段可使用免费通道,但不建议将其作为大型生产环境索引任务的唯一路径。当需要一个免费的候选模型进行测试时,可从 nvidia/nemotron-3-embed-1b:free 开始。该模型拥有 32,768 token 的上下文窗口,NVIDIA 的模型说明卡报告其输出向量为 2,048 维,支持在 34 种语言上进行评估,并允许通过 L2 重新归一化将向量切片为更小的维度。NVIDIA 以 OpenMDW-1.1 许可证发布该模型的权重。在使用免费通道进行生产环境索引任务前,请查看当前模型页面以确认可用性及速率限制。可用性和速率限制可能在批量处理期间发生变化。
An embedding model decides what your retrieval system can find. It turns each input into a vector and places related inputs near one another, so your application can retrieve by meaning instead of exact wording.
The best choice depends on the material you need to search. An English knowledge base, a multilingual support archive, a source-code repository, and a text-and-image collection have different requirements. Vector size, context length, public weights, and price can also change the decision.
We shortlisted models for English RAG, multilingual retrieval, code search, text-and-image retrieval, and low-cost indexing, then sent live requests to each one through our embeddings endpoint.
Last verified: 11 September 2026. Our embedding model catalog returned 37 entries on that date, including batch and preview variants of some models.
Providers can add or remove routes, and prompt prices can change. Check the current model page before starting a large indexing or re-indexing job. If the model page differs from this guide, use the model page.
TL;DR
Start with openai/text-embedding-3-small for English RAG.
Test voyageai/voyage-4-large when your inputs exceed 8,192 tokens or you want to move between Voyage 4 tiers without rebuilding the index. qwen/qwen3-embedding-8b has a longer context window at a lower prompt price if you want an open-weight alternative.
Use qwen/qwen3-embedding-8b for multilingual retrieval with public weights, voyageai/voyage-code-4 for code search, and google/gemini-embedding-2 or voyageai/voyage-multimodal-3.5 for text-and-image retrieval.
perplexity/pplx-embed-v1-0.6b has the lowest prompt price among the paid text models in our shortlist. nvidia/nemotron-3-embed-1b:free is a free text-embedding route with a 32,768-token context window.
Our API checks confirm request and response behavior, not retrieval quality. Compare at least two candidates on labeled queries and documents from your application before building or rebuilding the full index.
Best embedding models by use case
By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy .
Use case Recommended model Why it is on the shortlist
Default English RAG openai/text-embedding-3-small Low prompt price, 8,192-token context, and adjustable output dimensions
Inputs longer than 8,192 tokens voyageai/voyage-4-large 32,000-token context, four selectable dimensions, and embeddings compatible with the other Voyage 4 tiers
Multilingual retrieval with public weights qwen/qwen3-embedding-8b Support for more than 100 languages, a 32,768-token context window, and public weights
Code search voyageai/voyage-code-4 Built for retrieving code and related technical content
Text-and-image retrieval google/gemini-embedding-2 Places text and images in the same embedding space. voyageai/voyage-multimodal-3.5 is the second option we verified
Lowest-priced paid text option perplexity/pplx-embed-v1-0.6b Prompt price of $0.004 per million input tokens with a 32,000-token context window
Free text embeddings nvidia/nemotron-3-embed-1b:free A free route with a 32,768-token context window and a multilingual model card
Choose a model by input type
Start with the material you need to search. The chart sorts by input type first, then by whether you need public weights, then by your main constraint.
The chart gives you a starting point. Compare at least two candidates with queries and documents from your application, and track how often each model retrieves the relevant chunks before you build or rebuild a large index.
How we chose these models
We used our live catalog to confirm availability, context windows, input types, and prompt prices. We checked the model providers’ documentation for capabilities and benchmark results, then sent live requests through our embeddings endpoint to confirm request and response behavior.
We sent a two-string batch request to 19 models and ran 28 checks in total, covering batch input, configurable dimensions, image input, text-and-image input, and error handling. The 16 paid models returned one vector per input. The responses confirmed the default dimensions in the table below and showed that the dimensions parameter works with OpenAI Text Embedding 3 Small, Gemini Embedding 2, and Voyage 4 Large. A request for a model that does not exist returned a 400 error with the message Model openai/does-not-exist does not exist .
The three free routes returned a 404 error from our test account because its privacy settings do not allow routing to providers that may train on free-model prompts. We describe that setting in the free options section below. The default dimensions we list for those models come from their model cards, not from our responses.
These requests confirmed API compatibility. They did not measure retrieval quality. We excluded response times because each model received one small request under different serving conditions. We matched each use case to a documented capability, then used language coverage, context length, output dimensions, public weights, and prompt price to narrow the candidates. Published evaluations, including MTEB and CoIR, helped us identify models to test. Your labeled retrieval results should decide which one you deploy.
Compare the shortlisted embedding models
The prices below are the prompt prices listed in our catalog on 11 September 2026, and the image price for Gemini Embedding 2 is the catalog value on 18 September 2026. The default dimensions come from our live responses, except for the free NVIDIA model, whose default comes from its model card.
Model Input Context Default dimensions Prompt price per million tokens Weights
openai/text-embedding-3-small Text 8,192 1,536 $0.02 Closed
openai/text-embedding-3-large Text 8,192 3,072 $0.13 Closed
voyageai/voyage-4-lite Text 32,000 1,024 $0.02 Closed
voyageai/voyage-4 Text 32,000 1,024 $0.06 Closed
voyageai/voyage-4-large Text 32,000 1,024 $0.12 Closed
voyageai/voyage-code-4 Text, optimized for code 32,000 1,024 $0.12 Closed
voyageai/voyage-multimodal-3.5 Text and image 32,000 1,024 $0.12 Closed
qwen/qwen3-embedding-8b Text 32,768 4,096 $0.01 Open
qwen/qwen3-embedding-4b Text 32,768 2,560 $0.02 Open
perplexity/pplx-embed-v1-0.6b Text 32,000 1,024 $0.004 Open
perplexity/pplx-embed-v1-4b Text 32,000 2,560 $0.03 Open
google/gemini-embedding-2 Text and image 8,192 3,072 $0.20 for text, $0.45 for image tokens Closed
nvidia/nemotron-3-embed-1b:free Text 32,768 2,048 Free Open
The dimensions column affects raw index size. See “Calculate vector storage” below for the formula and examples.
Our catalog contains more models than the table shows. baai/bge-m3 adds another open multilingual option, mistralai/mistral-embed-2312 provides a general text alternative, and mistralai/codestral-embed-2505 targets code retrieval. All three returned vectors in our text request check. The catalog also lists a set of 512-token open models from BAAI, E5, GTE, and Sentence Transformers at $0.005 to $0.01 per million tokens, which we did not test for this guide.
Best default for English RAG
Start with openai/text-embedding-3-small when you need a managed model for an English knowledge base. Its default vector has 1,536 values, half as many as the default from openai/text-embedding-3-large . With the same numeric format, Small uses half as much raw vector storage. Keep the 1,536-value default for the first evaluation and reduce it only if the default meets the retrieval target but storage or search cost remains a constraint.
You can shorten its vectors with the dimensions parameter. Our request with "dimensions": 256 returned a vector with 256 values. Smaller vectors use less database storage and reduce the work required for similarity search, but they can also reduce retrieval quality. Test the smaller size before changing an existing index.
OpenAI describes text-embedding-3-large as its most capable embedding model for English and non-English tasks. Test it when Small misses relevant results, particularly when users search across languages. Moving to Large changes the vector size and prompt price, so compare both models on held-out questions before re-embedding the corpus.
If your inputs exceed Small’s 8,192-token limit, test Voyage 4 Large. For shorter inputs, compare both models using the same chunks, queries, and relevance labels.
Voyage 4 family for longer inputs
Test voyageai/voyage-4-large when you need a managed 32,000-token model with adjustable dimensions and compatibility across Voyage 4 tiers. It accepts up to 32,000 tokens and supports 256, 512, 1,024, and 2,048 dimensions. In our API check, setting dimensions to 512 returned a vector with 512 values.
The family includes voyageai/voyage-4 and voyageai/voyage-4-lite , which use the same context window. Voyage states that all embeddings created with the 4 series are compatible with each other. This lets you test another Voyage 4 tier against an existing index. Validate a tier change on a representative sample before deployment. A compatible vector space removes the rebuild requirement, but it does not guarantee identical retrieval results.
Run Large and at least one lower Voyage 4 tier against the same evaluation set. If they retrieve the same relevant passages at the same cutoff, the lower tier may meet your requirements.
The Voyage models are managed services. If you need public weights for self-hosting or deployment control, Qwen3 Embedding 8B is the open-weight alternative in this shortlist.
Best open-weight model for multilingual retrieval
qwen/qwen3-embedding-8b is our open-weight multilingual choice. It supports more than 100 languages, and you can download the weights from Hugging Face or call the hosted model through our API. Our default API request returned 4,096 values. With the same numeric format, a 4,096-value vector uses four times the raw storage of a 1,024-value vector. Include that difference when estimating index size. Our API checks do not establish whether Qwen3 Embedding 8B retrieves better than nvidia/nemotron-3-embed-1b:free . When route cost and vector size affect the decision, test both on the same labeled set before you commit.
The Qwen model card reports a multilingual score of 70.58 on the Massive Text Embedding Benchmark, MTEB, as of 5 June 2025. This is a vendor-reported public benchmark result. Use it to shortlist Qwen, then test every language and domain your application supports, because your corpus can produce a different ranking.
The model page lists a 32,768-token context window, but the limit is set per endpoint. On 11 September 2026 the DeepInfra and SiliconFlow endpoints for Qwen3 Embedding 8B listed 32,768 tokens and the Nebius endpoint listed 32,000. If you send inputs longer than 32,000 tokens, check the endpoint list and pin a provider with the larger limit, or keep inputs at or under 32,000 tokens.
The 4B version uses the same model-level context window and returned 2,560 values in our check. It requires fewer resources to run yourself. Hosted prices depend on the available providers, so compare the current model pages for Qwen3 Embedding 8B and 4B before choosing.
baai/bge-m3 gives you a second open multilingual model to compare with Qwen3 on your own dataset. Its model card covers more than 100 languages and inputs of up to 8,192 tokens, and our request returned a 1,024-value vector. The upstream model can return several representation types, including dense vectors and sparse vectors. Our standard embeddings response returns the dense vector. Use the upstream implementation if your retrieval design depends on its other representation types.
Best embedding models for code search
Use voyageai/voyage-code-4 when a query needs to retrieve a function, file, or piece of code documentation. Voyage built the model for code retrieval and coding-agent workloads. Its 32,000-token context window can accept long files or queries, although the index chunks should still represent code units that are useful when retrieved.
mistralai/codestral-embed-2505 is the main code-specific alternative in our catalog. It returned 1,536 values in our text request check. Its context window is 8,192 tokens, against 32,000 for Voyage Code 4. Test Voyage Code 4 first and use Codestral Embed as the comparison model.
Public evaluations can help you compare the candidates. The Code Information Retrieval Benchmark, CoIR , contains ten datasets covering eight code-retrieval tasks across seven domains. Use it to compare models before testing the strongest candidates on work from your own repositories. One useful internal evaluation maps issue descriptions or developer questions to the files and chunks needed to resolve them.
Best models for text-and-image retrieval
Use google/gemini-embedding-2 when the same index needs to search text and images. It places both input types in one embedding space, so a text query can retrieve an image with related meaning and an image query can retrieve related text.
We verified text, base64-encoded PNG image, and combined text-and-image requests through our embeddings endpoint. Each request returned one vector with 3,072 values. We also sent a text request with "dimensions": 768 and received a vector with 768 values. In our check, a 64 by 64 pixel PNG counted as 258 prompt tokens and cost $0.000128. Our catalog prices image input for this model separately, at $0.45 per million image tokens on the model page as of 18 September 2026, so check the model page before indexing a large image collection. This guide limits the recommendation to text and image inputs because those are the request types we verified.
voyageai/voyage-multimodal-3.5 is the second text-and-image model we verified. Its image request and its combined text-and-image request each returned a vector with 1,024 values, and the same 64 by 64 pixel PNG counted as 89 prompt tokens and cost $0.00003. It has a 32,000-token context window and a prompt price of $0.12 per million tokens. Gemini Embedding 2 and Voyage Multimodal 3.5 use different vector spaces, so choose one before you index.
google/gemini-embedding-001 is a separate text-only model with a 20,000-token context window. The two Gemini models use different vector spaces, so changing between them requires you to re-embed the index.
nvidia/llama-nemotron-embed-vl-1b-v2:free is a free text-and-image route with a 131,072-token context window in our catalog. Our test account could not call it for the privacy-setting reason described below, so we left it out of the recommendation.
Best low-cost and free options
perplexity/pplx-embed-v1-0.6b has the lowest paid text price in our shortlist at $0.004 per million input tokens. It accepts up to 32,000 tokens and returned a vector with 1,024 values. The larger perplexity/pplx-embed-v1-4b returned 2,560 values and has a prompt price of $0.03 per million input tokens.
Perplexity positions the 0.6B model for lightweight, low-latency retrieval and the 4B model for higher retrieval quality. Our API checks confirmed the default vector sizes, not which Perplexity model retrieves better. Treat the provider’s positioning as a shortlist signal and compare both models on labeled queries.
Use a free route for evaluation, not as the only path for a large production indexing job. Start with nvidia/nemotron-3-embed-1b:free when you need a free candidate to test. It has a 32,768-token context window, and NVIDIA’s model card reports a 2,048-value output vector, evaluation across 34 languages, and support for slicing the vector to a smaller size with L2 re-normalization. NVIDIA publishes the weights under the OpenMDW-1.1 license. Check the current model page for availability and rate limits before using a free route for a production indexing job. Availability and rate limits can change during a batch
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-27 | 8.49 | 23 | 入选 |
| 2026-09-26 | 8.77 | 36 | 未入选 |
| 2026-09-25 | 9.22 | 37 | 未入选 |
| 2026-09-24 | 9.95 | 34 | 未入选 |