Tag
Hugging Face announces native support for GGUF files in the transformers library, allowing easier use of quantized models with PyTorch tooling and performance comparable to llama.cpp.
Hugging Face's transformers library now supports GGUF models from llama.cpp, enabling efficient local inference on consumer hardware through familiar APIs.