Transformers now runs llama.cpp quants

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 22, 2026 5 min read
Transformers now runs llama.cpp quants


Transformers now runs llama.cpp quants

Running large language models on a laptop is easier because the Transformers library can now load GGUF files directly. This format, created by the llama.cpp team, packs model weights and metadata into a single file. Users can choose different quantization levels to balance file size against precision.

The integration reuses ggml kernels through the

kernels

library to match the performance of llama.cpp. The initial work focuses on Apple Silicon machines using the Qwen3.5 architecture.

What is the GGUF file format?

GGUF packages model weights, tokenizer information, and optional chat templates into one file. It supports various quantization levels, allowing users to trade precision for a smaller memory footprint.

For example, Unsloth’s Qwen3.5-4B model changes size depending on the variant chosen:

GGUF variantFile sizeTradeoff

BF16


8.42 GBUnquantized reference

Q6_K


3.53 GBMore precision than the smaller variants

Q5_K_M


3.14 GBA middle ground between size and precision

Q4_K_M


2.74 GBA practical starting point for local inference

We suggest starting with

Q4_K_M

before trying

Q5_K_M

or

Q6_K

if memory allows. Aggressive quantization helps larger models fit, but the quality loss depends on the model and the task. The Hub’s GGUF documentation describes available quantization types.

Load GGUF with transformers

To get started, you need:

  • An Apple Silicon Mac.
  • A PyTorch version supported by the published ggml-quantization kernel builds, usually the two latest PyTorch releases.
  • The latest version of transformers (main for now, until the next release) and a compatible version of
    kernels

    .

Install the required packages with:

pip install -U "git+https://github.com/huggingface/transformers.git" kernels

To load a GGUF model, pass its Hub

model_id

and filename as

gguf_file

to

from_pretrained

. No extra configuration is needed: when the weights stay packed on Metal, transformers automatically loads the compatible ggml/Metal layer kernels and uses

ggml-org/ggml-attn

as the attention implementation. If that kernel cannot be fetched, the model falls back to

"sdpa"

with a warning, and you can always force

"sdpa"

by passing

attn_implementation="sdpa"

explicitly.

That is the only GGUF-specific step. Everything after it is the standard transformers API:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"

tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    gguf_file=filename
)

messages = [{"role": "user", "content": "Explain why the sky is blue in a few sentences."}]
inputs = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(**inputs, max_new_tokens=256)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Without a compatible quantization kernel, the loader falls back to dequantizing the model and uses more memory.

Serve GGUF with your preferred interface

You can also use the same checkpoint with

transformers serve

, which exposes an OpenAI-compatible API:

pip install -U "transformers[serving] @ git+https://github.com/huggingface/transformers.git" kernels

transformers serve "unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf"

The model argument uses

<model_id>:<filename>.gguf

: before the colon is the Hub repository (

unsloth/Qwen3.5-4B-GGUF

), and after it is the file to load (

Qwen3.5-4B-Q4_K_M.gguf

). This selects a specific quantization from a repository that may contain several.

For models whose chat template supports thinking, add

--reasoning off

to skip it or

--reasoning on

to enable it. The default,

--reasoning auto

, follows the chat template’s default.

You can connect a client such as Jan or Pi by adding a custom OpenAI-compatible provider with these settings:

SettingValue
Base URL

http://localhost:8000/v1



Model ID

unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf



Transformers runs the model on your Mac, while the client provides the conversation interface. The same endpoint can be used by other clients that support this API.

Benchmarking against llama.cpp

Our reference for local inference performance is llama.cpp. The comparison below focuses on three GGUF checkpoints: a small dense model, a larger dense model, and a mixture-of-experts model.

The llama.cpp column comes from the

llama-bench

tool (build

5f55650a7

, release b10200, Metal backend from ggml 0.18.0), run as

llama-bench -m <file> -p 0 -n 128 -r 3

, which reports

tg128

: the token-generation rate over 128 decoded tokens, averaged across three repetitions, with prompt processing excluded. The transformers column is

generate

producing the same 128 tokens from a 12-token prompt, best of three warmed runs, and it includes prefill.

Measured on a MacBook Pro M2 Max, 32 GB unified memory, macOS 26.6, PyTorch 2.12.1, kernels 0.17.0, plugged in.

The benchmark script

To run the Transformers benchmark:

import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id, filename = "unsloth/Qwen3.5-4B-GGUF", "Qwen3.5-4B-Q4_K_M.gguf"

model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
inputs = tokenizer("The capital of France is Paris. The capital of Germany is", return_tensors="pt")
inputs = inputs.to(model.device)

with torch.inference_mode():
    model.generate(**inputs, max_new_tokens=8, min_new_tokens=8, do_sample=False)  # warm up
    torch.mps.synchronize()
    for _ in range(3):
        time.sleep(90)  # let the machine cool: back-to-back runs decay by 10% or more
        start = time.perf_counter()
        model.generate(**inputs, max_new_tokens=128, min_new_tokens=128, do_sample=False)
        torch.mps.synchronize()
        print(f"{128 / (time.perf_counter() - start):.1f} tok/s")

For the other column:

llama-bench -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M -p 0 -n 128 -r 3

Transformers is close to llama.cpp across all three checkpoints. The chart uses the same measurements described above; it does not imply identical benchmark conditions, since the Transformers measurement includes prefill while

llama-bench

reports decode-only throughput.

Transformers and llama.cpp

When GGML and llama.cpp joined Hugging Face, we described their complementary roles: llama.cpp provides a foundation for local inference, while transformers provides a foundation for model definition. GGUF support brings those two closer together.

llama.cpp remains our recommended engine when your priority is efficient local inference. Its dedicated runtime, memory management, and broad hardware support are built around that goal. This integration gives developers a convenient way to work with the same GGUF checkpoints inside transformers:

  • Experiment with GGUF in Python and PyTorch. Inspect intermediate activations with hooks, modify a model’s forward pass, or prototype custom layers using familiar PyTorch tools.
  • Evaluate GGUF models. Use your existing transformers evaluation workflows to measure the quality of quantized checkpoints.
  • Validate GGUF conversions. For us as developers, loading the original checkpoint and its GGUF conversion in transformers makes it easier to check that the weights were converted correctly, accounting for quantization error.
  • Try new decoding ideas. Use custom logits processors and stopping criteria with
    generate

    , or write your own generation loop in Python.

  • Fine-tune from a GGUF checkpoint. Dequantize the weights and continue with a standard transformers training workflow.

For that last case, use

GgufConfig(dequantize=True)

:Source Read original →

Scroll to Top