Llama.cpp server running ~2 weeks straight. Loses its mind?

Reddit r/LocalLLaMA News

Summary

User reports that Qwen3.6 models running on llama.cpp server become significantly less capable after ~2 weeks of continuous operation, and restarting sessions does not resolve the issue.

I’ve got Qwen3.6 27b and Qwen3.6 35b running in two separate instances for over two weeks and they are considerably dumber now than when I launched them. is this a thing? am I going crazy? edit: sorry I’ve been using opencode and have started new sessions, which didn’t fix the situation.
Original Article

Similar Articles

Qwen3.6 27B more dumb in vLLM compared to llama.cpp

Reddit r/LocalLLaMA

A user reports that the Qwen3.6-27B model performs better and more reliably with llama.cpp than with vLLM, citing tool call errors and 'lobotomized' behavior in vLLM despite extensive configuration.

Qwen 3.6 27B flags/settings in llama.cpp

Reddit r/LocalLLaMA

A user shares their llama.cpp server configuration for running Qwen 3.6 27B on an RTX 5090, achieving 80-100 t/s, and asks the community for alternative settings and tips.

qwen3.6 just stops

Reddit r/LocalLLaMA

A user reports an issue where the Qwen 3.6 model stops mid-task when served via vLLM with specific Docker and speculative decoding configurations.