Tag
The article discusses trust boundaries in AI agent systems where an LLM plans an executable DAG of agents, highlighting challenges in validating plans and ensuring safety through human approval and per-agent permissions.
This paper introduces GAVEL, a framework that uses graph world models to verify and repair long-horizon LLM planning for robotic tasks, significantly improving success rates and efficiency in simulations.
GraphThink is a framework that integrates task graphs and scene graphs to enhance LLM-based planning for long-horizon embodied tasks, achieving state-of-the-art results on the ALFRED benchmark and improving generalization and closed-loop replanning.
AnovaX is a local-first desktop voice assistant that uses an LLM planner and multi-agent orchestrator to execute tasks on the user's computer. It features a safety layer, adaptive recovery, and a phone-friendly remote interface.
RoboTALES introduces a two-stage framework combining LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training, significantly outperforming existing methods on long-horizon manipulation tasks.
shadcn released the Agent Skill project improve, which has high-cost models perform code auditing and planning while low-cost models execute, forming an installable, orchestrated system with an execution closed loop.
This paper proposes a symbolic feedback-driven iterative self-refinement framework to improve the robustness and reliability of large language models in long-horizon planning tasks. The method uses natural language prompting, a symbolic verifier, and a plan recognizer to enhance feasibility and correctness.
This paper proposes UP-NRPA, an online framework that integrates user portraits with nested rollout policy adaptation using large language models to dynamically customize dialogue strategies without offline training, achieving 100% success on multiple dialogue tasks.
Introduces Simmer, a benchmark for evaluating latent failures in LLM-generated executable plans using a human-curated symbolic world model in the kitchen domain. Experiments show frontier LLMs achieve at most 17% error-free plans, with up to 56% containing latent failures, and counterfactual foresight simulation reduces failures significantly.