Product Updates Hugging Face Blog

Transformers now runs llama.cpp quants

GGUFtransformersllama.cpplocal inferencequantization

Transformers is adding support for running GGUF models efficiently, so users can pick checkpoints sized for their laptop's memory and load them through familiar APIs — choose a GGUF from the Hub, load it with from_pretrained, and generate locally. Local inference has become far more practical thanks to llama.cpp, whose engine powers tools like Ollama, LM Studio, and Jan, and alongside projects such as MLX. Hugging Face cites a recent demo of Qwen3.6 27B running inside a Pi coding agent via llama.cpp on a MacBook Pro, which Julien Chaumond described as feeling very close to the latest Opus on non-trivial tasks in Hugging Face codebases.

GGUF, developed by the llama.cpp team (which also publishes quantized checkpoints under ggml-org on the Hub), packages model weights and metadata — including tokenizer information and an optional chat template — into one file, with quantization levels that trade precision for a smaller memory footprint. Variants like Q4_K_M mix tensor precisions, keeping mostly 4-bit weights while holding sensitive tensors at higher precision. Publishers such as Unsloth, LM Studio Community, and bartowski supply ready-to-use GGUF checkpoints, and these models have been downloaded millions of times. For Unsloth's Qwen3.5-4B, file sizes run 8.42 GB for BF16 (unquantized reference), 3.53 GB for Q6_K, 3.14 GB for Q5_K_M, and 2.74 GB for Q4_K_M.

To bring performance close to llama.cpp, the team is reusing its underlying ggml kernels through the kernels library and reducing overhead in generate. The initial focus is local inference on Apple Silicon, starting with the Qwen3.5 architecture. Getting started requires an Apple Silicon Mac, a PyTorch version supported by the published ggml-quantization kernel builds (usually the two latest releases), and the latest version of transformers. Hugging Face suggests starting with Q4_K_M and moving to Q5_K_M or Q6_K with more memory, noting that more aggressive quantization can help larger models fit but the quality tradeoff depends on the model and task — so it should be evaluated on the work you actually want the model to do.

Read original →

← Back to home