Tag
MILES is a framework that improves LLM reasoning by dynamically expanding step-wise memory with learnable selection heads, achieving better accuracy-efficiency tradeoffs.
This paper identifies the 'modal ceiling' and 'correlation ceiling' in test-time scaling for reasoning models, showing that beyond a few dozen samples, additional sampling does not improve selection accuracy and can even harm it, highlighting the identifiability gap between generating and recognizing correct answers.