Tag
Omni-Decision introduces an evidence-ledger planning approach for omni-modal agents to handle noisy multimodal observations, achieving state-of-the-art accuracy on OmniGAIA and WorldSense benchmarks.
The article proposes that agent plans should explicitly map task dependencies rather than relying on sequential lists, using a podcast workflow example to show how this makes waiting or repeating work more visible and efficient.
This paper introduces WM-SAR, a world-model correction method for agent planning that repairs causal subgraphs rather than visible symptoms, achieving better stabilization under token budgets compared to standard LLM correctors.
The paper introduces CreativityBench, a benchmark for evaluating large language models' ability to creatively repurpose tools based on affordance reasoning. It highlights that current models struggle with creative problem-solving despite strong general reasoning capabilities.