@ClementDelangue: Decision models now run on device in llama.cpp. Free, fast, private! llama serve -hf ggml-org/Kev-4B-GGUF
Summary
ClementDelangue announces that decision models can now run on-device in llama.cpp, highlighting a free, fast, and private workflow using the Kev-4B GGUF release via `llama serve`.
View Cached Full Text
Cached at: 10/02/26, 06:45 PM
Decision models now run on device in llama.cpp. Free, fast, private!
llama serve -hf ggml-org/Kev-4B-GGUF https://t.co/ApnQQaKB4Z
Similar Articles
New in llama.cpp: Decision Models
llama.cpp server now supports 'decision models' via a new /v1/systemone endpoint, where a single forward pass scores typed options (choice/score/yes-no) with probabilities instead of generating text. Several ggml-org models (Julia-1, Laya, Kev-4B, lev, OpenJev) are available in GGUF, following TypeSafe's Jev System One format.
Google launches their own llama.cpp wrapper – llama.app (local AI models server)
Google launches llama.app, a wrapper for llama.cpp that serves as a local server for running AI models.
@julien_c: Llama.cpp has a new branding + official website. Run local models today! Now more than ever, open source must win. By @…
Llama.cpp has unveiled a new branding and official website, promoting the local execution of AI models and reinforcing the importance of open-source software.
@ErickSky: Forget about vLLM, llama.cpp, and expensive GPUs. [colibri] This runs GLM-5.2 (744B MoE) on ~25 GB of RAM with pure C a…
colibri is a pure C inference tool that runs the GLM-5.2 744B MoE model on ~25 GB RAM by streaming experts from disk, eliminating the need for expensive GPUs.
Llama.cpp version 0.2.0 is out!
Llama.cpp, a popular open-source tool for running LLaMA models, has released version 0.2.0 with changelog and pre-built binaries available on GitHub.