{"version":1,"type":"story","url":"https://digestai.news/story/hugging-face-adds-gguf-support-to-transformers-for-local-inference","json":"https://digestai.news/story/hugging-face-adds-gguf-support-to-transformers-for-local-inference.json","markdown":"https://digestai.news/story/hugging-face-adds-gguf-support-to-transformers-for-local-inference.md","slug":"hugging-face-adds-gguf-support-to-transformers-for-local-inference","headline":"Hugging Face adds GGUF support to Transformers for local inference","summary":"Hugging Face announced that the Transformers library can now load GGUF‑format checkpoints directly, letting users run quantized models on a laptop with the familiar from_pretrained API. The feature reuses llama.cpp’s ggml kernels via the new kernels library and initially targets Apple Silicon, starting with the Qwen3.5 architecture. Users can pick a GGUF model from the Hub, specify the filename with the gguffile argument, and generate text without extra configuration.\n\nThe integration aims to match llama.cpp’s performance on a MacBook Pro M2 Max (32 GB unified memory, macOS 26.6) using PyTorch 2.12.1 and kernels 0.17.0. Benchmarks on three checkpoints – a small dense model, a larger dense model, and a mixture‑of‑experts model – show token‑generation rates close to llama.cpp, despite the Transformers measurement including prefill. The update also includes two generation‑loop optimizations: dropping an unnecessary attention mask and deferring the stopping check, which reduce CPU‑GPU synchronization overhead.\n\nGGUF models, which have been downloaded millions of times, support several quantization levels such as Q4KM, Q5KM and Q6K, allowing users to trade precision for a smaller memory footprint. The blog post thanks contributors including Arthur Zucker and Cyril Vallez, and invites the community to request additional model support.","keyPoints":["Transformers now loads GGUF checkpoints via a gguffile argument in from_pretrained.","Initial support focuses on Apple Silicon and Qwen3.5 models using ggml Metal kernels.","Quantization options like Q4KM, Q5KM and Q6K let users fit larger models on laptop memory."],"whyItMatters":"Bringing GGUF support into Transformers simplifies local, offline AI inference, lowering cloud costs and expanding access for developers and hobbyists.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["Hugging Face","llama.cpp","Apple","Unsloth"],"models":["Qwen3.5-4B"],"people":["Arthur Zucker","Cyril Vallez","Sayak Paul"]},"firstPublishedAt":"2026-09-22T00:00:00Z","updatedAt":"2026-09-22T00:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"Hugging Face","title":"Transformers now runs llama.cpp quants","url":"https://huggingface.co/blog/transformers-llama-cpp-quants","publishedAt":"2026-09-22T00:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"Qwen Model Optimization and Format Clarification","url":"https://digestai.news/thread/qwen3-8-27b-gguf-updated-with-new-tensor-layout","storyCount":3},"cite":{"text":"Digest AI, \"Hugging Face adds GGUF support to Transformers for local inference\", 22 September 2026, https://digestai.news/story/hugging-face-adds-gguf-support-to-transformers-for-local-inference","publisher":"Digest AI","title":"Hugging Face adds GGUF support to Transformers for local inference","datePublished":"2026-09-22T00:00:00Z","url":"https://digestai.news/story/hugging-face-adds-gguf-support-to-transformers-for-local-inference"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}