For everyone that uses OpenCode / Pi - Heres your promptprocessing fix!
Summary
A pull request for llama.cpp fixes the constant prompt processing issue that occurs when using OpenCode or Pi with the library.
Similar Articles
There's a new PR for llamacpp claiming to boost prompt processing with rocm by around 15%, also fixes a bug which makes Q2_K 28x faster
A new PR for llama.cpp boosts prompt processing on ROCm by ~15% and fixes a bug making Q2_K quantization 28x faster.
pi 0.81.0 adds support for llama.cpp
Pi 0.81.0 adds support for llama.cpp, enabling local LLM inference within the Pi coding agent harness.
PSA: If you haven’t updated Llama.cpp for a couple of days and find MTP to not be performing well, update llamacpp.
Update Llama.cpp for a significant token generation speed boost, up to 1.5-1.8x, and improved prompt processing.
Tip: use this llama.cpp PR to improve PP on Intel ARC
A llama.cpp PR significantly improves prompt processing speed on Intel ARC GPUs, with benchmark showing speed increase from 245t/s to 462t/s on a B580. The improvement currently works for F16 KV quantization, with plans to support other quants.
llama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference
A call for contributions to open pull requests in llama.cpp aimed at optimizing CPU, RAM, disk, and hybrid inference for faster performance, targeting completion by year-end.