Tag
The paper proposes 'Space', a skill-guided adaptive action chunking method for long-horizon LLM agents, improving success rates by 7.0%–31.3% and reducing LLM decision rounds by up to 78.9%.
AdaPlanBench is a dynamic benchmark for evaluating LLM agents' ability to adaptively plan under progressively revealed world and user constraints through multi-turn interactions, showing current models struggle especially with user constraints.