我们正在为在 transformers 中高效运行 GGUF 模型添加支持,因此您可以通过熟悉的 transformers API 使用适合您笔记本电脑内存大小的检查点。从 Hub 选择一个 GGUF,使用 from_pretrained 加载它,然后开始在您的本地机器上进行生成。
在您的笔记本电脑上运行 AI 模型已经变得容易得多,而 llama.cpp 在其中发挥了重要作用。其推理引擎为 Ollama、LM Studio 和 Jan 等本地 AI 工具提供动力。与 MLX 等项目一起,它已使本地推理成为日常使用的可行选择。
本地 AI 体验的一个近期示例:
GGUF 由 llama.cpp 团队开发,是用于本地推理的广泛使用的格式。该团队还在 Hub 上的 ggml-org 下分享量化检查点。Unsloth、LM Studio Community 和 bartowski 等发布者也提供多种量化级别的即用型 GGUF 检查点,因此用户可以选择适合其机器的版本。GGUF 模型已被下载数百万次。
我们也希望让使用 transformers 在本地运行这些模型变得更加容易。兼容性只有在模型易于运行时才有意义。为了将性能提升至接近 llama.cpp 的水平,我们通过 kernels 库重用其底层 ggml 内核,并减少 generate 中的开销。我们的初始重点是 Apple Silicon 上的本地推理,从 Qwen3.5 架构开始。
什么是 GGUF 文件格式?
GGUF 将模型权重和元数据(包括分词器信息和可选的聊天模板)打包在一个文件中。它支持不同的量化级别,让您可以在精度和内存占用之间进行权衡。Q4_K_M 等变体混合了张量精度,主要使用 4 位权重,同时保持敏感张量具有更高的精度。
以下是量化如何改变 Unsloth 的 Qwen3.5-4B 的文件大小:
GGUF 变体
文件大小
权衡
BF16
8.42 GB
未量化的参考
Q6_K
3.53 GB
比更小的变体具有更高的精度
Q5_K_M
3.14 GB
尺寸与精度之间的中间地带
Q4_K_M
2.74 GB
本地推理的实用起点
我们建议从 Q4_K_M 开始,然后在有更多可用内存的情况下尝试 Q5_K_M 或 Q6_K。更激进的量化可以帮助更大的模型适应,但质量权衡取决于模型和任务。在您实际希望模型执行的工作上进行评估。Hub 上的 GGUF 文档描述了可用的量化类型。
使用 transformers 加载 GGUF
要开始使用,您需要:
Apple Silicon Mac。
由已发布的 ggml-量化内核构建支持的 PyTorch 版本,通常是最新的两个 PyTorch 发行版。
transformers 的最新版本(目前为主分支,直到下一个发布)以及兼容版本的 kernels。
pip install -U "git+https://github.com/huggingface/transformers.git" kernels
要加载 GGUF 模型,请将其 Hub model_id 和文件名作为 gguf_file 传递给 from_pretrained。
无需额外配置:当权重保持在 Metal 上打包时,transformers 会自动加载兼容的 ggml/Metal 层内核,并使用 ggml-org/ggml-attn 作为注意力实现。如果无法获取该内核,模型将回退到带有警告的 "sdpa",并且您始终可以通过显式传递 attn_implementation="sdpa" 来强制使用 "sdpa"。有关更多加载选项,请参阅 GGUF 文档。
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(
model_id,
gguf_file=filename
)
这是唯一的 GGUF 特定步骤。之后的所有内容都是标准的 transformers API:
messages = [{ "role" : "user" , "content" : "Explain why the sky is blue in a few sentences." }]
inputs = tokenizer.apply_chat_template(
messages,
tokenize= True ,
add_generation_prompt= True ,
return_dict= True ,
return_tensors= "pt" ,
).to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens= 256 )
print (tokenizer.decode(outputs[ 0 ], skip_special_tokens= True ))
如果没有兼容的量化内核,加载器将回退到反量化模型并使用更多内存。
使用你喜欢的接口提供 GGUF 服务
你还可以使用 transformers serve 来运行相同的检查点,它提供了一个与 OpenAI 兼容的 API:
pip install -U "transformers[serving] @ git+https://github.com/huggingface/transformers.git" kernels
transformers serve "unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf"
模型参数使用
对于支持思考的聊天模板模型,添加 --reasoning off 以跳过它,或添加 --reasoning on 以启用它。默认值 --reasoning auto 遵循聊天模板的默认设置。有关详细信息,请参阅推理选项。
你可以通过添加自定义的 OpenAI 兼容提供程序来连接 Jan 或 Pi 等客户端,设置如下:
设置
值
基础 URL
http://localhost:8000/v1
模型 ID
unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf
transformers 在您的 Mac 上运行模型,而客户端提供对话界面。支持此 API 的其他客户端也可以使用相同的端点。
与 llama.cpp 进行基准测试
我们本地推理性能的参考标准是 llama.cpp。下面的比较专注于三个 GGUF 检查点:一个小规模稠密模型、一个较大规模的稠密模型和一个混合专家(MoE)模型。
llama.cpp 列的数据来自 llama-bench 工具(构建版本 5f55650a7,发布版本 b10200,ggml 0.18.0 的 Metal 后端),运行命令为 llama-bench -m
在 MacBook Pro M2 Max、32 GB 统一内存、macOS 26.6、PyTorch 2.12.1、kernels 0.17.0 环境下测量,并连接电源。
基准测试脚本
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id, filename = "unsloth/Qwen3.5-4B-GGUF" , "Qwen3.5-4B-Q4_K_M.gguf"
model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
inputs = tokenizer( "The capital of France is Paris. The capital of Germany is" , return_tensors= "pt" )
inputs = inputs.to(model.device)
with torch.inference_mode():
model.generate(inputs, max_new_tokens= 8 , min_new_tokens= 8 , do_sample= False )
torch.mps.synchronize()
for _ in range ( 3 ):
time.sleep( 90 )
start = time.perf_counter()
model.generate(inputs, max_new_tokens= 128 , min_new_tokens= 128 , do_sample= False )
torch.mps.synchronize()
print ( f" { 128 / (time.perf_counter() - start): .1 f} tok/s" )
对于另一列数据:
llama-bench -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M -p 0 -n 128 -r 3
在所有三个检查点中,transformers 的性能接近 llama.cpp。图表使用了上述相同的测量方法;它并不意味着基准测试条件完全相同,因为 Transformers 的测量包括预填充时间,而 llama-bench 报告的是仅解码吞吐量。
transformers 和 llama.cpp
当 GGML 和 llama.cpp 加入 Hugging Face 时,我们描述了它们的互补角色:llama.cpp 为本地推理提供基础,而 transformers 为模型定义提供基础。GGUF 的支持使这两者更加紧密地结合在一起。
当你的首要目标是高效的本地推理时,llama.cpp 仍然是我们推荐的引擎。其专用的运行时、内存管理和广泛的硬件支持都是围绕这一目标构建的。此集成为开发者提供了一种便捷的方式,使其能够在 transformers 中使用相同的 GGUF 检查点:
在 Python 和 PyTorch 中实验 GGUF。使用钩子检查中间激活值,修改模型的前向传播过程,或使用熟悉的 PyTorch 工具原型化自定义层。
评估 GGUF 模型。利用现有的 transformers 评估工作流来衡量量化检查点的质量。
验证 GGUF 转换。对于我们开发者而言,在 transformers 中加载原始检查点及其 GGUF 转换版本,使得检查权重是否正确转换变得更加容易,同时考虑到量化误差。
尝试新的解码思路。使用自定义的 logits 处理器和停止条件配合 generate 函数,或在 Python 中编写自己的生成循环。
从 GGUF 检查点进行微调。反量化权重,并继续使用标准的 transformers 训练工作流。
对于最后一种情况,请使用 GgufConfig(dequantize=True):
import torch
from transformers import AutoModelForCausalLM, GgufConfig
model = AutoModelForCausalLM.from_pretrained(
"unsloth/Qwen3.5-4B-GGUF",
gguf_file="Qwen3.5-4B-Q4_K_M.gguf",
quantization_config=GgufConfig(dequantize=True),
dtype=torch.bfloat16,
)
超越 GGUF:用于更多模型的 ggml 内核
更大的机遇在于将 ggml 的性能带给 llama.cpp 尚未支持的模型。
transformers 已经提供了这些架构的 PyTorch 实现。随着 PyTorch 中可用 ggml 内核和量化方案,我们可以致力于加速其支持的操作,而无需先在 llama.cpp 中实现整个模型。这对于新架构、研究模型以及可能永远不会获得专用 llama.cpp 实现的自定义变体尤其有用。
这种机遇不仅限于 GGUF 格式本身。内核操作于张量;它不需要整个模型都来自 GGUF 文件。相同的构建模块可以集成到其他 transformers 模型和加载工作流中。这也为其他模态开辟了道路:计算机视觉模型、音频模型和多模态模型可以重用兼容的注意力、归一化和矩阵乘法内核,而无需先在 llama.cpp 中获得完整的实现。每种架构仍需进行集成和验证;此处最初的 GGUF 示例涵盖文本生成。
使用 Python 和 PyTorch 进行快速本地推理
我们还希望展示在保持模型和生成循环位于 Python 中的情况下,我们能取得多大的进展。通过合适的内核和高效的生成循环,Python 和 PyTorch 能够提供强劲的本地推理性能。内核处理繁重的计算,而生成循环通过避免不必要的同步来保持 GPU 忙碌。
我们的重点是在不需要 torch.compile 的情况下使急切执行(eager execution)变得快速。对于交互式使用,我们希望快速启动并持续输出令牌流,而无需在输入形状变化时出现编译暂停或重新编译。这项工作的两个主要部分是内核和 generate 本身。
重用 ggml 的 Metal 内核
内核是用于在 GPU 上执行操作的小型程序。PyTorch 提供通用实现;专用内核可以减少工作量、合并多个操作,或直接以存储格式读取量化权重。
kernels 库使我们能够在 Hub 上分发兼容的 ggml Metal 内核构建版本,并从 transformers 中调用它们。这将 ggml 的工作带入 PyTorch 模型中,而无需用单独推理运行时替换模型。
内核 | 功能
ggml-quantization | 读取用于矩阵运算的打包量化权重,包括 MoE 模型中选定的专家。它避免了在每次解码操作之前展开整个权重矩阵。
ggml-norm
融合归一化操作,包括 Qwen3.5 和 Qwen3.8 使用的零中心 RMSNorm。
ggml-attn
为提示词处理和令牌解码提供 ggml 的 Metal Flash Attention。
ggml-gated-delta-net
加速 Qwen3.5 和 Qwen3.8 混合架构中线性注意力层所使用的门控 delta 网络。
topk
在 MoE 模型中为每个令牌选择专家,结合 softmax 和 top-k 路由。这是我们要自研的 Metal 实现。
前四个包建立在 ggml 的内核之上;top-k 内核解决了 MoE 路由中的另一个瓶颈。它们共同减少了生成每个令牌所需的 GPU 工作量。
为了展示层内核的贡献,我们比较了带有和不带有这些内核的同一打包 GGUF 检查点。量化内核在两种配置中均保持启用:禁用它也会改变权重的表示方式,并衡量不同的权衡取舍。
保持 CPU 和 GPU 协同工作
如果 GPU 没有工作可做,更快的内核也无济于事。在生成过程中,CPU 负责调度 GPU 操作并控制产生下一个令牌的循环。从 GPU 读取结果可能会迫使 CPU 等待排队操作完成。即使每次令牌都重复微小的等待,也会显著降低吞吐量。
两项更改解决了 generate 中的这一问题,从而为所有 Transformer 模型带来改进(不仅限于运行 GGUF 文件时):
尽早丢弃不必要的注意力掩码 (#48814)。当支持的仅解码器输入没有填充时,其全一填充掩码可以在生成开始时移除。下游注意力代码不再需要反复检查该掩码以确定是否可以跳过。因果注意力仍然得以保留。
推迟停止检查 (#47975)。在支持的路径上,generate 异步复制停止决策并在下一步中消耗它。CPU 可以继续调度工作,而 GPU 正在运行。流式令牌使用相同的方法,并且超出停止条件的额外步骤会从结果中移除。
这些更改改进了模型周围的生成循环,因此其用途超出了 GGUF 的范围。它们与内核工作相辅相成:内核降低了操作的成本,而更少的同步点使 CPU 调度和 GPU 执行能够重叠。
这些测量值保持所有层内核启用;条形图隔离了生成循环中的更改。
当前的限制和后续步骤
初始目标是 Apple Silicon 上的单个交互式对话。有几条边界需要注意:
打包推理路径目前仅限 MPS。通过反量化导入 GGUF 仍然是一个单独的选项;对文件格式的支持并不意味着每种设备上都提供打包内核。
填充和批处理仍需改进。无填充输入受益于上述的掩码优化。有填充的批次无法采用相同的捷径,性能可能较低。我们希望将这项工作扩展到 MPS 上的 generate_batch。
架构覆盖范围有限。打包加载器目前涵盖 Qwen3.5 密集型和 MoE 架构,包括兼容的 Qwen3.8 检查点。添加对其他架构的支持相对简单,我们将逐步扩大覆盖范围。
如果您有希望在 transformers 中使用的 GGUF 模型,请提交包含检查点和您用例的问题。这将帮助我们优先支持人们本地运行的模型。
致谢
我们要感谢 Arthur Zucker 发起这项工作并审阅我的所有 PR,以及 Cyril Vallez 对 generate PR 的贡献。我们感谢 Sayak Paul、llama.cpp 团队和 Bertrand Chevalier 在集成内核方面提供的帮助。我们还要感谢 Aritra Roy Gosthipaty 和 Pedro Cuenca 审阅这篇博客文章,以及 Lysandre Debut 监督该项目。
We're adding support for running GGUF models efficiently in transformers , so you can use checkpoints sized for your laptop's memory through the familiar transformers APIs. Pick a GGUF from the Hub, load it with
from_pretrained , and start generating on your own machine.
Running AI models on your laptop has become much easier, and llama.cpp has been a big part of that. Its inference engine powers local AI tools such as Ollama, LM Studio, and Jan. Alongside projects like MLX , it has helped make local inference a practical option for everyday use.
A recent example of what local AI can feel like:
GGUF , developed by the llama.cpp team, is a widely used format for local inference. The team also shares quantized checkpoints under ggml-org on the Hub . Publishers such as Unsloth , LM Studio Community , and bartowski also provide ready-to-use GGUF checkpoints in a range of quantizations, so users can pick the version that fits their machine. GGUF models have been downloaded millions of times.
We want to make it easier to run these models locally with transformers, too. Compatibility is only useful if the model is pleasant to run. To bring performance close to llama.cpp, we're reusing its underlying ggml kernels through the kernels library, and reducing overhead in generate . Our initial focus is local inference on Apple Silicon, starting with the Qwen3.5 architecture.
What is the GGUF file format?
GGUF packages model weights and metadata, including tokenizer information and an optional chat template, in one file. It supports different quantization levels, letting you trade some precision for a smaller memory footprint. Variants such as Q4_K_M mix tensor precisions, using mostly 4-bit weights while keeping sensitive tensors at higher precision.
Here's how quantization changes the file size of Unsloth's Qwen3.5-4B :
GGUF variant
File size
Tradeoff
BF16
8.42 GB
Unquantized reference
Q6_K
3.53 GB
More precision than the smaller variants
Q5_K_M
3.14 GB
A middle ground between size and precision
Q4_K_M
2.74 GB
A practical starting point for local inference
We suggest starting with Q4_K_M , then trying Q5_K_M or Q6_K if you have more memory available. More aggressive quantization can help larger models fit, but the quality tradeoff depends on the model and the task. Evaluate it on the work you actually want the model to do. The Hub's GGUF documentation describes the available quantization types.
Load GGUF with transformers
To get started, you need:
An Apple Silicon Mac .
A PyTorch version supported by the published ggml-quantization kernel builds , usually the two latest PyTorch releases.
The latest version of transformers (main for now, until the next release) and a compatible version of kernels .
pip install -U "git+https://github.com/huggingface/transformers.git" kernels
To load a GGUF model, pass its Hub model_id and filename as gguf_file to from_pretrained .
No extra configuration is needed: when the weights stay packed on Metal, transformers automatically loads the compatible ggml/Metal layer kernels and uses ggml-org/ggml-attn as the attention implementation. If that kernel cannot be fetched, the model falls back to "sdpa" with a warning, and you can always force "sdpa" by passing attn_implementation="sdpa" explicitly. See the GGUF documentation for more loading options.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(
model_id,
gguf_file=filename
)
That is the only GGUF-specific step. Everything after it is the standard transformers API:
messages = [{ "role" : "user" , "content" : "Explain why the sky is blue in a few sentences." }]
inputs = tokenizer.apply_chat_template(
messages,
tokenize= True ,
add_generation_prompt= True ,
return_dict= True ,
return_tensors= "pt" ,
).to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens= 256 )
print (tokenizer.decode(outputs[ 0 ], skip_special_tokens= True ))
Without a compatible quantization kernel, the loader falls back to dequantizing the model and uses more memory.
Serve GGUF with your preferred interface
You can also use the same checkpoint with transformers serve , which exposes an OpenAI-compatible API:
pip install -U "transformers[serving] @ git+https://github.com/huggingface/transformers.git" kernels
transformers serve "unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf"
The model argument uses
For models whose chat template supports thinking, add --reasoning off to skip it or --reasoning on to enable it. The default, --reasoning auto , follows the chat template’s default. See the reasoning options for details.
You can connect a client such as Jan or Pi by adding a custom OpenAI-compatible provider with these settings:
Setting
Value
Base URL
http://localhost:8000/v1
Model ID
unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf
transformers runs the model on your Mac, while the client provides the conversation interface. The same endpoint can be used by other clients that support this API.
Benchmarking against llama.cpp
Our reference for local inference performance is llama.cpp. The comparison below focuses on three GGUF checkpoints: a small dense model, a larger dense model, and a mixture-of-experts model.
The llama.cpp column comes from the llama-bench tool (build 5f55650a7 , release b10200, Metal backend from ggml 0.18.0), run as llama-bench -m
Measured on a MacBook Pro M2 Max, 32 GB unified memory, macOS 26.6, PyTorch 2.12.1, kernels 0.17.0,
plugged in.
The benchmark script
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id, filename = "unsloth/Qwen3.5-4B-GGUF" , "Qwen3.5-4B-Q4_K_M.gguf"
model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
inputs = tokenizer( "The capital of France is Paris. The capital of Germany is" , return_tensors= "pt" )
inputs = inputs.to(model.device)
with torch.inference_mode():
model.generate(inputs, max_new_tokens= 8 , min_new_tokens= 8 , do_sample= False )
torch.mps.synchronize()
for _ in range ( 3 ):
time.sleep( 90 )
start = time.perf_counter()
model.generate(inputs, max_new_tokens= 128 , min_new_tokens= 128 , do_sample= False )
torch.mps.synchronize()
print ( f" { 128 / (time.perf_counter() - start): .1 f} tok/s" )
For the other column:
llama-bench -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M -p 0 -n 128 -r 3
Transformers is close to llama.cpp across all three checkpoints. The chart uses the same measurements described above; it does not imply identical benchmark conditions, since the Transformers measurement includes prefill while llama-bench reports decode-only throughput.
transformers and llama.cpp
When GGML and llama.cpp joined Hugging Face , we described their complementary roles: llama.cpp provides a foundation for local inference, while transformers provides a foundation for model definition. GGUF support brings those two closer together.
llama.cpp remains our recommended engine when your priority is efficient local inference. Its dedicated runtime, memory management, and broad hardware support are built around that goal. This integration gives developers a convenient way to work with the same GGUF checkpoints inside transformers:
Experiment with GGUF in Python and PyTorch. Inspect intermediate activations with hooks, modify a model's forward pass, or prototype custom layers using familiar PyTorch tools.
Evaluate GGUF models. Use your existing transformers evaluation workflows to measure the quality of quantized checkpoints.
Validate GGUF conversions. For us as developers, loading the original checkpoint and its GGUF conversion in transformers makes it easier to check that the weights were converted correctly, accounting for quantization error.
Try new decoding ideas. Use custom logits processors and stopping criteria with generate , or write your own generation loop in Python.
Fine-tune from a GGUF checkpoint. Dequantize the weights and continue with a standard transformers training workflow.
For that last case, use GgufConfig(dequantize=True) :
import torch
from transformers import AutoModelForCausalLM, GgufConfig
model = AutoModelForCausalLM.from_pretrained(
"unsloth/Qwen3.5-4B-GGUF" ,
gguf_file= "Qwen3.5-4B-Q4_K_M.gguf" ,
quantization_config=GgufConfig(dequantize= True ),
dtype=torch.bfloat16,
)
Beyond GGUF: ggml kernels for more models
The bigger opportunity is bringing ggml's performance to models that llama.cpp does not support.
transformers already provides the PyTorch implementations of these architectures. With ggml kernels and quantization schemes available in PyTorch, we can work toward accelerating their supported operations without first implementing the entire model in llama.cpp. This is especially useful for new architectures, research models, and custom variants that may never receive a dedicated llama.cpp implementation.
That opportunity extends beyond the GGUF format itself. A kernel operates on tensors; it does not require the whole model to come from a GGUF file. The same building blocks can be integrated into other transformers models and loading workflows. This also opens a path to other modalities: computer vision models, audio models, and multimodal models could reuse compatible attention, normalization, and matrix multiplication kernels without first having a full implementation in llama.cpp. Each architecture still needs integration and validation; the initial GGUF examples here cover text generation.
Fast local inference with Python and PyTorch
We also wanted to show how far we can get while keeping the model and generation loop in Python. With the right kernels and an efficient generation loop, Python and PyTorch can deliver strong local inference performance. The kernels handle the heavy computation, while the generation loop keeps the GPU busy by avoiding unnecessary synchronization.
Our focus was to make eager execution fast without requiring torch.compile . For interactive use, we wanted a quick start and a steady stream of tokens, without compilation pauses or recompilation when input shapes change. The two main pieces of that work are the kernels and generate itself.
Reusing ggml's Metal kernels
A kernel is a small program that performs an operation on the GPU. PyTorch supplies general-purpose implementations; a specialized kernel can do less work, combine several operations, or read quantized weights directly in their stored format.
The kernels library lets us distribute compatible builds of ggml's Metal kernels on the Hub and call them from transformers. That brings ggml's work into the PyTorch model without replacing the model with a separate inference runtime.
Kernel
What it does
ggml-quantization
Reads packed quantized weights for matrix operations, including the selected experts in an MoE model. It avoids expanding the whole weight matrix before each decode operation.
ggml-norm
Fuses normalization operations, including the zero-centered RMSNorm used by Qwen3.5 and Qwen3.8.
ggml-attn
Provides ggml's Metal flash attention for prompt processing and token decoding.
ggml-gated-delta-net
Accelerates the gated delta network used in the linear-attention layers of the Qwen3.5 and Qwen3.8 hybrid architectures.
topk
Selects the experts for each token in an MoE model, combining softmax and top-k routing. This is our own Metal implementation.
The first four packages build on ggml's kernels; the top-k kernel addresses a separate bottleneck in MoE routing. Together they reduce the GPU work needed for each generated token.
To show the contribution of the layer kernels, we compare the same packed GGUF checkpoints with and without them. The quantization kernel stays enabled in both configurations: disabling it would also change how weights are represented and would measure a different tradeoff.
Keeping the CPU and GPU working together
Faster kernels only help if the GPU has work to do. During generation, the CPU schedules GPU operations and controls the loop that produces the next token. Reading a result back from the GPU can force the CPU to wait until queued operations finish. Repeating even a small wait for every token can noticeably reduce throughput.
Two changes address this in generate , which results in improvements for all transformers models (not just when running GGUF files):
Drop an unnecessary attention mask early (#48814) . When a supported decoder-only input has no padding, its all-ones padding mask can be removed at the start of generation. Downstream attention code no longer needs to inspect that mask repeatedly to determine whether it can be skipped. Causal attention is still preserved.
Defer the stopping check (#47975) . On supported paths, generate copies the stopping decision asynchronously and consumes it on the following step. The CPU can keep scheduling work while the GPU runs. Streaming tokens use the same approach, and any extra step past the stopping condition is removed from the result.
These changes improve the generation loop around the model, so their usefulness extends beyond GGUF. They complement the kernel work: kernels reduce the cost of an operation, while fewer synchronization points let CPU scheduling and GPU execution overlap.
These measurements keep all layer kernels enabled; the bars isolate the changes to the generation loop.
Current limitations and next steps
The initial target is a single interactive conversation on Apple Silicon. There are a few boundaries to keep in mind:
The packed inference path is MPS-only for now. GGUF import through dequantization remains a separate option; support for the file format does not imply that packed kernels are available on every device.
Padding and batching still need work. Unpadded inputs benefit from the mask optimization described above. Padded batches cannot take the same shortcut and can have lower performance. We want to extend the work to generate_batch on MPS.
Architecture coverage is limited. The packed loader currently covers the Qwen3.5 dense and MoE architectures, including compatible Qwen3.8 checkpoints. Adding support for other architectures is relatively straightforward, and we’ll expand coverage gradually.
If you have a GGUF model you would like to use in transformers, open an issue with the checkpoint and your use case. That will help us prioritize support for the models people are running locally.
Acknowledgments
We would like to thank Arthur Zucker for initiating this work and reviewing all of my PRs, and Cyril Vallez for the generate PRs. We are grateful to Sayak Paul , the llama.cpp team , and Bertrand Chevalier for their help integrating the kernels. We also thank Aritra Roy Gosthipaty and Pedro Cuenca for reviewing this blog post, and Lysandre Debut for overseeing the project.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-28 | 8.4 | 14 | 入选 |
| 2026-09-27 | 8.4 | 25 | 未入选 |
| 2026-09-26 | 8.49 | 40 | 未入选 |
| 2026-09-25 | 8.77 | 43 | 未入选 |
| 2026-09-24 | 9.22 | 46 | 未入选 |
| 2026-09-23 | 9.95 | 32 | 未入选 |