In this article
Transformers now runs llama.cpp quants
Running large language models on a laptop is easier because the Transformers library can now load GGUF files directly. This format, created by the llama.cpp team, packs model weights and metadata into a single file. Users can choose different quantization levels to balance file size against precision.
The integration reuses ggml kernels through the
kernels
library to match the performance of llama.cpp. The initial work focuses on Apple Silicon machines using the Qwen3.5 architecture.
What is the GGUF file format?
GGUF packages model weights, tokenizer information, and optional chat templates into one file. It supports various quantization levels, allowing users to trade precision for a smaller memory footprint.
For example, Unsloth’s Qwen3.5-4B model changes size depending on the variant chosen:
We suggest starting with
Q4_K_M
before trying
Q5_K_M
or
Q6_K
if memory allows. Aggressive quantization helps larger models fit, but the quality loss depends on the model and the task. The Hub’s GGUF documentation describes available quantization types.
Load GGUF with transformers
To get started, you need:
- An Apple Silicon Mac.
- A PyTorch version supported by the published ggml-quantization kernel builds, usually the two latest PyTorch releases.
- The latest version of transformers (main for now, until the next release) and a compatible version of
kernels
.
Install the required packages with:
pip install -U "git+https://github.com/huggingface/transformers.git" kernels
To load a GGUF model, pass its Hub
model_id
and filename as
gguf_file
to
from_pretrained
. No extra configuration is needed: when the weights stay packed on Metal, transformers automatically loads the compatible ggml/Metal layer kernels and uses
ggml-org/ggml-attn
as the attention implementation. If that kernel cannot be fetched, the model falls back to
"sdpa"
with a warning, and you can always force
"sdpa"
by passing
attn_implementation="sdpa"
explicitly.
That is the only GGUF-specific step. Everything after it is the standard transformers API:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(
model_id,
gguf_file=filename
)
messages = [{"role": "user", "content": "Explain why the sky is blue in a few sentences."}]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))Without a compatible quantization kernel, the loader falls back to dequantizing the model and uses more memory.
Serve GGUF with your preferred interface
You can also use the same checkpoint with
transformers serve
, which exposes an OpenAI-compatible API:
pip install -U "transformers[serving] @ git+https://github.com/huggingface/transformers.git" kernels transformers serve "unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf"
The model argument uses
<model_id>:<filename>.gguf
: before the colon is the Hub repository (
unsloth/Qwen3.5-4B-GGUF
), and after it is the file to load (
Qwen3.5-4B-Q4_K_M.gguf
). This selects a specific quantization from a repository that may contain several.
For models whose chat template supports thinking, add
--reasoning off
to skip it or
--reasoning on
to enable it. The default,
--reasoning auto
, follows the chat template’s default.
You can connect a client such as Jan or Pi by adding a custom OpenAI-compatible provider with these settings:
Transformers runs the model on your Mac, while the client provides the conversation interface. The same endpoint can be used by other clients that support this API.
Benchmarking against llama.cpp
Our reference for local inference performance is llama.cpp. The comparison below focuses on three GGUF checkpoints: a small dense model, a larger dense model, and a mixture-of-experts model.
The llama.cpp column comes from the
llama-bench
tool (build
5f55650a7
, release b10200, Metal backend from ggml 0.18.0), run as
llama-bench -m <file> -p 0 -n 128 -r 3
, which reports
tg128
: the token-generation rate over 128 decoded tokens, averaged across three repetitions, with prompt processing excluded. The transformers column is
generate
producing the same 128 tokens from a 12-token prompt, best of three warmed runs, and it includes prefill.
Measured on a MacBook Pro M2 Max, 32 GB unified memory, macOS 26.6, PyTorch 2.12.1, kernels 0.17.0, plugged in.
The benchmark script
To run the Transformers benchmark:
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id, filename = "unsloth/Qwen3.5-4B-GGUF", "Qwen3.5-4B-Q4_K_M.gguf"
model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
inputs = tokenizer("The capital of France is Paris. The capital of Germany is", return_tensors="pt")
inputs = inputs.to(model.device)
with torch.inference_mode():
model.generate(**inputs, max_new_tokens=8, min_new_tokens=8, do_sample=False) # warm up
torch.mps.synchronize()
for _ in range(3):
time.sleep(90) # let the machine cool: back-to-back runs decay by 10% or more
start = time.perf_counter()
model.generate(**inputs, max_new_tokens=128, min_new_tokens=128, do_sample=False)
torch.mps.synchronize()
print(f"{128 / (time.perf_counter() - start):.1f} tok/s")For the other column:
llama-bench -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M -p 0 -n 128 -r 3
Transformers is close to llama.cpp across all three checkpoints. The chart uses the same measurements described above; it does not imply identical benchmark conditions, since the Transformers measurement includes prefill while
llama-bench
reports decode-only throughput.
Transformers and llama.cpp
When GGML and llama.cpp joined Hugging Face, we described their complementary roles: llama.cpp provides a foundation for local inference, while transformers provides a foundation for model definition. GGUF support brings those two closer together.
llama.cpp remains our recommended engine when your priority is efficient local inference. Its dedicated runtime, memory management, and broad hardware support are built around that goal. This integration gives developers a convenient way to work with the same GGUF checkpoints inside transformers:
- Experiment with GGUF in Python and PyTorch. Inspect intermediate activations with hooks, modify a model’s forward pass, or prototype custom layers using familiar PyTorch tools.
- Evaluate GGUF models. Use your existing transformers evaluation workflows to measure the quality of quantized checkpoints.
- Validate GGUF conversions. For us as developers, loading the original checkpoint and its GGUF conversion in transformers makes it easier to check that the weights were converted correctly, accounting for quantization error.
- Try new decoding ideas. Use custom logits processors and stopping criteria with
generate
, or write your own generation loop in Python.
- Fine-tune from a GGUF checkpoint. Dequantize the weights and continue with a standard transformers training workflow.
For that last case, use
GgufConfig(dequantize=True)
:
Source Read original →



