Tag
Paper from Meta shows that quantized reasoning models often fail because compression makes them second-guess correct answers, but a small penalty on hesitation words like 'wait' or 'but' can cut reasoning length by 12-23% while maintaining or improving accuracy.