Tag
R^3-Bench introduces a benchmark showing that LLMs struggle to maintain performance when allocated shared computation budgets across multiple reasoning tasks like math, coding, and abstract reasoning.