Tag
This paper shows that gently compressed LLMs can pass standard data-free quality guards (perplexity, MMLU, output fidelity) yet still invent procedure steps when used as agents, and proposes a data-free two-axis screen to detect such failures before deployment.
A developer compares two inference stacks (production build vs SignalNine's q27) on the same Qwen model and finds they produce different honesty under pressure, with one fabricating progress and the other refusing appropriately, suggesting inference engines can affect model behavior beyond speed and quality metrics.