llm-planning

Tag

Cards List
#llm-planning

the model plans out an executable DAG of agents and someone has to decide how much to trust that before it actually runs anything

Reddit r/AI_Agents ↗ · 2026-09-16

The article discusses trust boundaries in AI agent systems where an LLM plans an executable DAG of agents, highlighting challenges in validating plans and ensuring safety through human approval and per-agent permissions.

0 favorites 0 likes
#llm-planning

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

Hugging Face Daily Papers ↗ · 2026-09-16 Cached

This paper introduces GAVEL, a framework that uses graph world models to verify and repair long-horizon LLM planning for robotic tasks, significantly improving success rates and efficiency in simulations.

0 favorites 0 likes
#llm-planning

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

arXiv cs.AI ↗ · 2026-08-11 Cached

GraphThink is a framework that integrates task graphs and scene graphs to enhance LLM-based planning for long-horizon embodied tasks, achieving state-of-the-art results on the ALFRED benchmark and improving generalization and closed-loop replanning.

0 favorites 0 likes
#llm-planning

AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery

arXiv cs.AI ↗ · 2026-07-20 Cached

AnovaX is a local-first desktop voice assistant that uses an LLM planner and multi-agent orchestrator to execute tasks on the user's computer. It features a safety layer, adaptive recovery, and a phone-friendly remote interface.

0 favorites 0 likes
#llm-planning

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

Hugging Face Daily Papers ↗ · 2026-07-07 Cached

RoboTALES introduces a two-stage framework combining LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training, significantly outperforming existing methods on long-horizon manipulation tasks.

0 favorites 0 likes
#llm-planning

@yibie: Recommend this project. shadcn (creator of shadcn/ui) built an Agent Skill — let the most expensive model handle auditing and planning, and let the cheap model write code. The idea itself isn't new, but shadcn turned it into an installable, orchestrated system with an execution closed loop. Every design decision is practical...

X AI KOLs Timeline ↗ · 2026-07-05 Cached

shadcn released the Agent Skill project improve, which has high-cost models perform code auditing and planning while low-cost models execute, forming an installable, orchestrated system with an execution closed loop.

0 favorites 0 likes
#llm-planning

Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework

arXiv cs.AI ↗ · 2026-06-29 Cached

This paper proposes a symbolic feedback-driven iterative self-refinement framework to improve the robustness and reliability of large language models in long-horizon planning tasks. The method uses natural language prompting, a symbolic verifier, and a plan recognizer to enhance feasibility and correctness.

0 favorites 0 likes
#llm-planning

UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems

arXiv cs.CL ↗ · 2026-06-15 Cached

This paper proposes UP-NRPA, an online framework that integrates user portraits with nested rollout policy adaptation using large language models to dynamically customize dialogue strategies without offline training, achieving 100% success on multiple dialogue tasks.

0 favorites 0 likes
#llm-planning

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

arXiv cs.CL ↗ · 2026-06-15 Cached

Introduces Simmer, a benchmark for evaluating latent failures in LLM-generated executable plans using a human-curated symbolic world model in the kitchen domain. Experiments show frontier LLMs achieve at most 17% error-free plans, with up to 56% containing latent failures, and counterfactual foresight simulation reduces failures significantly.

0 favorites 0 likes
← Back to home

Submit Feedback