Tag
This paper introduces WM-SAR, a world-model correction method for agent planning that repairs causal subgraphs rather than visible symptoms, achieving better stabilization under token budgets compared to standard LLM correctors.
The paper introduces CreativityBench, a benchmark for evaluating large language models' ability to creatively repurpose tools based on affordance reasoning. It highlights that current models struggle with creative problem-solving despite strong general reasoning capabilities.