GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

arXiv cs.AI Papers

Summary

GraphThink is a framework that integrates task graphs and scene graphs to enhance LLM-based planning for long-horizon embodied tasks, achieving state-of-the-art results on the ALFRED benchmark and improving generalization and closed-loop replanning.

arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose GraphThink, a novel framework that integrates a task graph to provide structured knowledge for robust planning and a scene graph to maintain environmental memory for event-driven replanning. Specifically, the task graph guides LLM thinking through contextual prompting and iterative refinement, effectively mitigating planning hallucinations. Furthermore, within the GRPO framework, the task graph offers delicate reward design to train the LLM planner, enhancing long-horizon planning capabilities and improving generalization. Finally, an event-driven replanning module, powered by the scene graph, enables closed-loop environment awareness and error correction. GraphThink achieves state-of-the-art performance on the ALFRED benchmark. In particular, our high-level planner surpasses leading API-based LLMs on both the validation set and held-out long-horizon tasks, underscoring its robust zero-shot and few-shot capabilities. Additional evaluations further demonstrate strong out-of-distribution generalization to novel tasks and environments.
Original Article
View Cached Full Text

Cached at: 08/11/26, 08:04 AM

# GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
Source: [https://arxiv.org/html/2608.07905](https://arxiv.org/html/2608.07905)
Chen Li1,3, Sijie Cheng4,6, Yuelin Zhang1,3, Junxi Li5, Maozhi Huang1,3, Yang Liu4, and Wenbing Huang🖂1,2,3🖂denotes the corresponding author: hwenbing@ruc\.edu\.cn\.1Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China\.2Beijing Academy of Artificial Intelligence, Beijing, China\.3Beijing Key Laboratory of Research on Large Models and Intelligent Governance, Beijing, China\.4Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China\.5Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR, China\.6RayNeo\.AI, Shenzhen, China\.

###### Abstract

Embodied agents using LLM\-based planners often struggle with physical hallucinations, poor generalization to long\-horizon tasks, and lack of environmental awareness\. We propose*GraphThink*, a novel framework that integrates a*task graph*to provide structured knowledge for robust planning and a*scene graph*to maintain environmental memory for event\-driven replanning\. Specifically, the task graph guides LLM thinking through contextual prompting and iterative refinement, effectively mitigating planning hallucinations\. Furthermore, within the GRPO framework, the task graph offers delicate reward design to train the LLM planner, enhancing long\-horizon planning capabilities and improving generalization\. Finally, an event\-driven replanning module, powered by the scene graph, enables closed\-loop environment awareness and error correction\. GraphThink achieves state\-of\-the\-art performance on the ALFRED benchmark\. In particular, our high\-level planner surpasses leading API\-based LLMs on both the validation set and held\-out long\-horizon tasks, underscoring its robust zero\-shot and few\-shot capabilities\. Additional evaluations further demonstrate strong out\-of\-distribution generalization to novel tasks and environments\.

## IIntroduction

There has been a growing exploration of embodied agents designed to execute long\-horizon everyday tasks given human instructions\[[47](https://arxiv.org/html/2608.07905#bib.bib33),[17](https://arxiv.org/html/2608.07905#bib.bib45),[20](https://arxiv.org/html/2608.07905#bib.bib46),[4](https://arxiv.org/html/2608.07905#bib.bib34)\]\. In the field of Embodied AI, instruction following necessitates that agents perform three key operations: interpreting natural language, using egocentric visual observations, and executing physical action to navigate and interact with the environment\. A straightforward approach\[[27](https://arxiv.org/html/2608.07905#bib.bib27),[37](https://arxiv.org/html/2608.07905#bib.bib28),[8](https://arxiv.org/html/2608.07905#bib.bib35)\]involves training agents in an end\-to\-end supervised manner using large\-scale datasets with annotated instructions and low\-level expert action sequences\. However, this paradigm is resource\-intensive: it relies heavily on task\-specific data and shows poor generalization to unseen scenarios\. In contrast, data\-efficient hierarchical methods\[[25](https://arxiv.org/html/2608.07905#bib.bib12),[2](https://arxiv.org/html/2608.07905#bib.bib48),[46](https://arxiv.org/html/2608.07905#bib.bib2),[19](https://arxiv.org/html/2608.07905#bib.bib8)\]have emerged as a promising alternative: the high\-level planner first decomposes instructions into subtasks, while the low\-level executor subsequently accesses the skill library to translate these subtasks into executable actions in the environment\.

Recently, hierarchical methods have increasingly harnessed Large Language Models \(LLMs\) for high\-level planning\[[1](https://arxiv.org/html/2608.07905#bib.bib38),[34](https://arxiv.org/html/2608.07905#bib.bib37),[30](https://arxiv.org/html/2608.07905#bib.bib39),[38](https://arxiv.org/html/2608.07905#bib.bib40)\], owing to their strong language understanding and reasoning capabilities\. When initial planning fails, these models can engage in replanning by incorporating observed environmental objects as contextual prompts\[[36](https://arxiv.org/html/2608.07905#bib.bib9),[5](https://arxiv.org/html/2608.07905#bib.bib5),[19](https://arxiv.org/html/2608.07905#bib.bib8)\]\. Nonetheless, long\-horizon planning remains a major challenge for embodied agents, with three key limitations: \(i\) Despite their strong reasoning capabilities, general\-purpose LLMs lack physical grounding, leading to instruction misinterpretation, planning hallucinations, and increased failure rates as task complexity grows\. \(ii\) Supervised fine\-tuning on limited in\-domain data results in poor generalization to unseen long\-horizon tasks, except when employing prohibitively expensive annotations\. \(iii\) Many existing replanning strategies that primarily trigger corrections for low\-level failures are prone to myopic decisions and inefficient retries, as they lack the environmental awareness to detect plans that are executable by low\-level actions but misaligned with the instruction\.

To address these challenges, we propose*GraphThink*,a dual\-graph enhanced closed\-loop planning system, as illustrated in Fig\.[1](https://arxiv.org/html/2608.07905#S1.F1)\. This framework integrates two core structured representations: a*task graph*to provide structured knowledge for robust planning, and a*scene graph*to maintain environmental memory for event\-driven replanning\. The*task graph*plays three essential roles in high\-level planning\. First, it guides the LLM’s plan generation by explicitly incorporating the task graph into the prompt\. Second, it serves as an external verifier to detect planning hallucinations and refine the subtask sequence iteratively\. Finally, we adopt Group Relative Policy Optimization \(GRPO\)\[[31](https://arxiv.org/html/2608.07905#bib.bib41)\]with task graph\-based reward to enhance the reasoning capabilities of LLMs\. Since embodied planning admits multiple valid solutions for a single goal, reward design becomes particularly challenging\. The task graph addresses this issue by encapsulating diverse feasible paths, enabling reward signals that accommodate multiple valid solutions\. With this delicate reward design, we are able to facilitate the alignment between the high\-level action space and instructions, leading to stronger reasoning and generalization capabilities\.

To enhance the LLM’s environmental awareness and correct instruction misalignment within a closed\-loop planning process, we design an event\-driven replanning module powered by the scene graph\. The dynamically updatedscene graphserves asa task\-centric memory moduleto focus reasoning on task\-relevant environmental cues\. Unlike general scene graphs\[[10](https://arxiv.org/html/2608.07905#bib.bib53),[39](https://arxiv.org/html/2608.07905#bib.bib54)\]that often include numerous irrelevant objects, our design keeps the graph size manageable, filtering out noise and enabling more efficient LLM reasoning\. When replanning is triggered by low\-level execution errors or proactive checks upon high\-level subtask completion, the current scene graph memory and plan execution progress are provided to the thinking LLM to support efficient tracking and adaptation of the plan\. To ensure the quality of replanning, all proposed revisions are constrained by the task graph to maintain feasibility\. Together, these components form a closed\-loop system that continuously aligns plan execution with user intent\.

We evaluate GraphThink on ALFRED\[[33](https://arxiv.org/html/2608.07905#bib.bib3)\], a challenging benchmark for vision\-language navigation and interaction\. Our hierarchical planning frameworkranks first on the official ALFRED leaderboard\. Particularly, we evaluate the performance of the high\-level planner andcontribute a long\-horizon datasetto better assess long\-horizon generalization\. The results show GraphThink outperforms various leading API\-based LLMs on both the validation set and unseen long\-horizon benchmark, demonstrating few\-shot and zero\-shot learning capabilities\.Further, GraphThink shows strong generalization to novel actions in AI2\-Thor\[[21](https://arxiv.org/html/2608.07905#bib.bib61)\]and effective cross\-environment transfer to VirtualHome\[[28](https://arxiv.org/html/2608.07905#bib.bib57)\]\.

![Refer to caption](https://arxiv.org/html/2608.07905v1/x1.png)Figure 1:GraphThink consists of three core modules: \(a\) the high\-level planner with task graph generates an initial plan, \(b\) the memory\-aware low\-level policy executes navigation and interaction actions for each subtask, and \(c\) dynamic replanning with scene graph adjusts plans during execution\. The example here shows replanning triggered upon the completion of subtask “put\(knife, sidetable\)”, and the revised subtasks become “pickup\(knife\)” and “put\(knife, dining table\)”\.
## IIRelated Work

Task Planning for Embodied Agents\.Prior work\[[33](https://arxiv.org/html/2608.07905#bib.bib3),[27](https://arxiv.org/html/2608.07905#bib.bib27),[37](https://arxiv.org/html/2608.07905#bib.bib28),[35](https://arxiv.org/html/2608.07905#bib.bib50)\]train agents end\-to\-end to directly generate low\-level actions given language instructions, but their performance in long\-horizon tasks remains limited\. Recently, hierarchical or modular planning\[[15](https://arxiv.org/html/2608.07905#bib.bib7),[32](https://arxiv.org/html/2608.07905#bib.bib11),[19](https://arxiv.org/html/2608.07905#bib.bib8)\]have proven effective by decomposing tasks into subtasks to bridge the gap between natural instructions and executable actions\. In the early stage, template\-based methods\[[25](https://arxiv.org/html/2608.07905#bib.bib12),[46](https://arxiv.org/html/2608.07905#bib.bib2)\]are limited to predefined tasks and struggle to generalize\. To address this problem, LLMs are being explored as high\-level planners, either through few\-shot in\-context prompting\[[36](https://arxiv.org/html/2608.07905#bib.bib9),[19](https://arxiv.org/html/2608.07905#bib.bib8)\]or by supervised training on specific datasets\[[49](https://arxiv.org/html/2608.07905#bib.bib4),[5](https://arxiv.org/html/2608.07905#bib.bib5)\]\.Some works\[[14](https://arxiv.org/html/2608.07905#bib.bib51),[36](https://arxiv.org/html/2608.07905#bib.bib9),[18](https://arxiv.org/html/2608.07905#bib.bib47),[19](https://arxiv.org/html/2608.07905#bib.bib8)\]further introduce replanning mechanisms to adjust actions by accepting environmental feedback, triggering local corrections to immediate errorsor predefined state differences\. Hence, we propose an event\-driven dynamic replanning mechanism to enhance both plan feasibility and instruction alignment\.

Complex Reasoning with LLMs\.To solve complex reasoning tasks, Chain\-of\-Thought \(CoT\) methods\[[43](https://arxiv.org/html/2608.07905#bib.bib22),[6](https://arxiv.org/html/2608.07905#bib.bib29),[26](https://arxiv.org/html/2608.07905#bib.bib23)\]prompt LLMs to generate intermediate reasoning steps\. However, as the number of steps increases, errors tend to accumulate\. Self\-correction methods\[[24](https://arxiv.org/html/2608.07905#bib.bib26),[11](https://arxiv.org/html/2608.07905#bib.bib30)\]leverage feedback to refine incorrect reasoning and improve accuracy\. Moreover, Retrieval\-Augmented Generation \(RAG\)\[[45](https://arxiv.org/html/2608.07905#bib.bib32),[42](https://arxiv.org/html/2608.07905#bib.bib24)\]and knowledge graphs\[[41](https://arxiv.org/html/2608.07905#bib.bib25),[50](https://arxiv.org/html/2608.07905#bib.bib31)\]methods enhance reasoning by integrating structured external knowledge to improve accuracy and reduce hallucinations\. To further enhance the performance through learning from interaction, recent work has shifted to supervised fine\-tuning \(SFT\)\[[48](https://arxiv.org/html/2608.07905#bib.bib20)\]and reinforcement learning \(RL\)\[[29](https://arxiv.org/html/2608.07905#bib.bib16)\]\. In our work, we further introduce a task graph\-based RL framework that enhances the model’s long\-horizon planning capability through reward signals derived from the task graph\.

## IIIMethodology

As illustrated in Fig\.[1](https://arxiv.org/html/2608.07905#S1.F1), GraphThink mainly consists of three components: \(a\) High\-level planning with task graph; \(b\) Memory\-aware low\-level action policy; \(c\) Dynamic replanning with scene graph memory\. We provide the details of each component in this section\.

### III\-AHigh\-Level Planning with Task Graph

To provide physical grounding for high\-level planning, we introduce a task graph that models logical and executability constraints over subtask transitions\. Rather than merely memorizing trajectories observed in the training data, we design an automatic task graph construction pipeline to instantiate feasible subtask\-transition priors, which provide structured guidance for generating executable subtask sequences\. The task graph is integrated across multiple stages of the high\-level planning pipeline: it guides LLM planning through carefully designed prompts that incorporate the task graph, supports the design of reward functions in GRPO, and acts as an external verifier to provide feedback for reasoning\.

#### III\-A1Task Graph Construction

Subtasks in embodied planning often exhibit ordering dependencies due to physical preconditions \(e\.g\.,Slice\(object\)typically requiresPickup\(knife\)first\), and we therefore represent feasible subtask transitions as a directed task graph𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\), where nodes𝒱=\{vi\}i=1N\\mathcal\{V\}=\\\{v\_\{i\}\\\}\_\{i=1\}^\{N\}denote subtasks and each edge\(vi,vj\)∈ℰ\(v\_\{i\},v\_\{j\}\)\\in\\mathcal\{E\}indicates thatvjv\_\{j\}can executably followviv\_\{i\}\. Theplanning objectiveis: given a natural language instructionτ\\tau, high\-level planning aims to generate a subtask sequenceS​\(τ\)=\(s1,s2,…,sT\)S\(\\tau\)=\(s\_\{1\},s\_\{2\},\\dots,s\_\{T\}\)that guides the agent to complete the task as instructed\. Here, the subtask sequence should form a feasible pathΠ=\(v1,v2,…,vT\)\\Pi=\(v\_\{1\},v\_\{2\},\\dots,v\_\{T\}\)in𝒢\\mathcal\{G\}, and each object insis\_\{i\}belongs to the environmental object set𝒪env\\mathcal\{O\}\_\{\\text\{env\}\}, which comprises all visible objects recognized by a pretrained object detection module\.

Subtasks as Nodes\.Nodeviv\_\{i\}representing a subtask is denoted asAi​\(oi\)A\_\{i\}\(o\_\{i\}\)orAi​\(oi,oj\)A\_\{i\}\(o\_\{i\},o\_\{j\}\), whereAiA\_\{i\}denotes a high\-level action andoi,ojo\_\{i\},o\_\{j\}refer to relevant objects\. Each subtask is grounded in the available low\-level skill set, thereby mapping the LLM’s open\-ended text generation to robot\-executable actions\. Following prior work\[[25](https://arxiv.org/html/2608.07905#bib.bib12),[46](https://arxiv.org/html/2608.07905#bib.bib2)\], our evaluation on the ALFRED benchmark uses 12 subtask nodes as an appropriate granularity for high\-level planning to executable low\-level policies\. Given the vast combinatorial space of subtask–object pairs, we employ meta\-classes𝒞≔\{obj,rec,mov,spec\}\\mathcal\{C\}\\coloneqq\\\{\\texttt\{obj\},\\texttt\{rec\},\\texttt\{mov\},\\texttt\{spec\}\\\}to abstract away specific object instances, whereobjdenotes pickupable items,recdesignates fixed containers,movindicates portable containers, andspecdenotes task\-specific operation targets required by certain subtasks\. For example,speccan be instantiated as Lamp forToggleObject, Fridge forCoolObject, Microwave forHeatObject, and SinkBasin forCleanObject\. This shifts the planner’s focus to functional matching between subtasks and objects, thereby facilitating generalization across environments among different objects with similar functionalities\. During plan generation, meta\-classes are instantiated as concrete object names from the environmental object set𝒪env\\mathcal\{O\}\_\{\\text\{env\}\}\. During execution, each high\-level subtask is expanded into a serialized low\-level action policy and dynamically instantiated through environmental perception\. For example,Heat\(object, microwave\)is not a single primitive action but a policy involving navigation, object placement, microwave operation, and object retrieval\.

Subtask Transition as Edges\.A task graph can be constructed by traversing training trajectories\. However, trajectory\-based construction only captures demonstrations observed in the dataset and may therefore miss feasible yet unobserved subtask transition paths\. To build a more general task graph, we construct edges through LLM\-assisted transition compatibility analysis, which considers both semantic compatibility and resource\-state feasibility between subtasks\. Specifically, for each subtaskviv\_\{i\}, the LLM receives its serialized low\-level policy and abstracts it into a structured transition interface⟨Pi,Ei⟩\\langle P\_\{i\},E\_\{i\}\\rangle, wherePiP\_\{i\}denotes the preconditions required to start the subtask andEiE\_\{i\}denotes the state or resource effects after completing it\. We add a directed edgevp→vqv\_\{p\}\\to v\_\{q\}when the effectsEpE\_\{p\}of the preceding subtask satisfy the preconditionsPqP\_\{q\}of the subsequent subtask, and the transition between the two subtasks is semantically valid\. For example, afterPickupObject\(obj\), the resulting effect includesholding\(obj\)and hand\-resource occupation, which satisfies the precondition ofPutObject\(obj, rec\)that the object should be held\. This mechanism supports generalization to novel tasks without requiring trajectory demonstrations\.

Scalability of Task Graph\.The task graph is designed to be easily extensible\. When generalizing to new tasks or environments \(Sec\.[IV\-D](https://arxiv.org/html/2608.07905#S4.SS4)\), new subtask nodes can be introduced according to specific task requirements as the low\-level policy set evolves\. Importantly, their incident edges can be generated through the same transition compatibility analysis, enabling task\-graph extension without requiring new trajectory demonstrations\. This design allows practitioners to choose an appropriate subtask granularity according to the target environment and task requirements\. Appendix A further verifies the reliability of this construction strategy and its robustness to noisy edge constraints\.

#### III\-A2Prompt with Task Graph

To mitigate planning hallucinations in LLM\-based task planning, we propose a novel prompting strategy incorporating robotic domain knowledge via a task graph\. This strategy directs the agent to generate subtask sequences that follow the graph topology, where each subsequent subtask is guided to follow a neighboring node of previous subtask\. Specifically, the nodes are encoded as action\-target pairs assigned unique labels \(*e\.g\.*, ‘\[A\] PickupObject\(object\)’\), while the edges are converted into symbolic connection relationships \(*e\.g\.*, ‘A → B’\)\. Static prompt components \(e\.g\., agent roles, planning rules, output format\) are also incorporated to facilitate in\-context learning\. The complete prompt template is shown in Appendix G\.

#### III\-A3GRPO with Task Graph

Recent work\[[29](https://arxiv.org/html/2608.07905#bib.bib16),[40](https://arxiv.org/html/2608.07905#bib.bib17)\]demonstrates that rule\-based rewards combined with RL significantly enhance reasoning\. For embodied task planning, reward design is challenging because a single instruction may admit multiple feasible solutions\. We address this by combining graph\-based structural rewards with instruction\-following rewards in GRPO\. Specifically,task graph rewards encourage exploration over feasible task\-graph paths, thereby avoiding dependence on a single annotated trajectory, whileinstruction rewards serve as a semantic constraintthat selects instruction\-aligned plans among feasible candidates\. This complementary design improves generalization while preventing reward hacking, such as generating plans that satisfy task\-graph constraints but deviate from the instruction semantics\. We next describe the detailed reward design used in GRPO\.

Format Reward\.Prior research\[[12](https://arxiv.org/html/2608.07905#bib.bib14),[44](https://arxiv.org/html/2608.07905#bib.bib15)\]has shown that the format rewardRfmtR\_\{\\text\{fmt\}\}effectively constrains the model’s output to match the expected structure\. We customized the format reward to suit the output structure of high\-level task planning\. The format reward is computed as:

Rfmt=\{1\.0,ifXMLandJSONvalid,0\.5,ifXMLvalid butJSONinvalid,0\.0,otherwise\.R\_\{\\text\{fmt\}\}=\\begin\{cases\}1\.0,&\\text\{if \{XML\} and \{JSON\} valid\},\\\\ 0\.5,&\\text\{if \{XML\} valid but \{JSON\} invalid\},\\\\ 0\.0,&\\text\{otherwise\}\.\\end\{cases\}\(1\)Rfmt=0\.5R\_\{\\text\{fmt\}\}=0\.5, if the model’s response followsXML\_format,*i\.e\.*,<think\> … </think\><answer\> … </answer\>;Rfmt=1R\_\{\\text\{fmt\}\}=1, only when the content inside<answer\> … </answer\>is aJSON\_objectincluding the specified key,*e\.g\.*, “high\_level”\.

Node\-Level Reward\.To encourage each generated subtasksis\_\{i\}to be valid, we compute a node\-level reward via dynamically weighted object\-validity scores across four meta\-classes in𝒞\\mathcal\{C\}:

Rnode=∑c∈𝒞wc⋅Rnodec,\\textstyle R\_\{\\text\{node\}\}=\\sum\_\{c\\in\\mathcal\{C\}\}w\_\{c\}\\cdot R\_\{\\text\{node\}\}^\{c\},\(2\)where the weightwcw\_\{c\}is dynamically adjusted based on whether the corresponding meta classccappears in the plan\. Weights of unused meta classes are set to zero, and the remaining weights are renormalized to ensure that all weights sum to 1\. For each meta class, the rewardRnodecR\_\{\\text\{node\}\}^\{c\}is defined as the validity rate of objects belonging to meta classccin the plan, where an object is considered valid when its name exists in the corresponding environmental list𝒪envc\\mathcal\{O\}\_\{\\rm env\}^\{c\}, then:

Rnodec=1\|𝒪Sc\|​∑o∈𝒪Sc𝕀​\(o∈𝒪envc\),\\textstyle R\_\{\\text\{node\}\}^\{c\}=\\frac\{1\}\{\|\\mathcal\{O\}\_\{S\}^\{c\}\|\}\\sum\_\{o\\in\\mathcal\{O\}\_\{S\}^\{c\}\}\\mathbb\{I\}\\left\(o\\in\\mathcal\{O\}\_\{\\text\{env\}\}^\{c\}\\right\),\(3\)where𝒪Sc\\mathcal\{O\}\_\{S\}^\{c\}denotes the set of all objects in planS​\(τ\)S\(\\tau\)that correspond to subtasks of classcc, and𝕀​\(⋅\)\\mathbb\{I\}\(\\cdot\)is an indicator function that returns 1 if the condition holds and 0 otherwise\.

Edge\-Level Reward\.To encourage subtask transitions conforming to the edgesℰ\\mathcal\{E\}in the task graph𝒢\\mathcal\{G\}, this reward traverses and checks all adjacent subtask pairs\(st,st\+1\)\(s\_\{t\},s\_\{t\+1\}\)in the generated plan\. The reward is defined as follows:

Redge=∏1≤t≤T−1𝕀​\(st,st\+1∈𝒱\)⋅𝕀​\(\(st,st\+1\)∈ℰ\),\\textstyle R\_\{\\text\{edge\}\}=\\prod\_\{\\begin\{subarray\}\{c\}1\\leq t\\leq T\-1\\end\{subarray\}\}\\mathbb\{I\}\(s\_\{t\},s\_\{t\+1\}\\in\\mathcal\{V\}\)\\cdot\\mathbb\{I\}\(\(s\_\{t\},s\_\{t\+1\}\)\\in\\mathcal\{E\}\),\(4\)where the reward equals 1 if both subtasks\(st,st\+1\)\(s\_\{t\},s\_\{t\+1\}\)are nodes in𝒱\\mathcal\{V\}and the transition between them satisfies the feasible\-transition constraint of𝒢\\mathcal\{G\}, and 0 otherwise\.

Instruction Following Reward\.The reward for instruction followingRinstR\_\{\\text\{inst\}\}encourages the model to generate plans that satisfy user instructions\. Unlike tasks with unique solutions \(*e\.g\.*, arithmetic or classification\), task planning is inherently multi\-solution, making exact\-match evaluation insufficient\. To address this, we employ a two\-stage reward based on the Ground\-Truth \(GT\) planS∗S^\{\*\}\. First, we extract critical subtasks𝒮key∗\\mathcal\{S\}^\{\*\}\_\{\\text\{key\}\}fromS∗S^\{\*\}that are essential for completing the instruction\. The generated planS​\(τ\)S\(\\tau\)receives a reward of 0\.5 if it covers all critical subtasks, and a full reward of 1\.0 only if it exactly matches the GT planS∗S^\{\*\}\. Otherwise, the reward is 0\. This two\-stage design ensures that the model is first driven to cover indispensable steps and then incentivized to produce fully precise plans\.

Overall Reward Function\.The overall reward function integrates rewards to guide the model in generating executable subtask sequences that meet instruction goals:

Rtotal=Rfmt\+Rnode\+Redge\+Rinst\.\\textstyle R\_\{\\text\{total\}\}=R\_\{\\text\{fmt\}\}\+R\_\{\\text\{node\}\}\+R\_\{\\text\{edge\}\}\+R\_\{\\text\{inst\}\}\.\\\\\(5\)

#### III\-A4Verification with Task Graph

Inspired by\[[9](https://arxiv.org/html/2608.07905#bib.bib1)\], we treat the task graph as an external validation tool and leverage its feedback to refine the responses of the LLM\. Since task\-graph constraints have already been learned through graph\-based rewards during GRPO, verification mainly serves as a lightweight safeguard at inference to correct occasional illegal transitions and reduce residual planning hallucinations\. The verification function is composed of two parts:

Node\-Level Legitimacy Verification\. To determine whether each object in the initial subtask sequence is valid, we perform a node\-level legitimacy verification:𝕍node=∏c∈𝒞,o∈𝒪Sc𝕀​\(o∈𝒪envc\)\\mathbb\{V\}\_\{\\text\{node\}\}=\\prod\_\{c\\in\\mathcal\{C\},o\\in\\mathcal\{O\}\_\{S\}^\{c\}\}\\mathbb\{I\}\(o\\in\\mathcal\{O\}\_\{\\text\{env\}\}^\{c\}\)\.

Edge\-Level Legitimacy Verification\.To verify that all action transitions satisfy the edge constraints of the task graph, the edge\-level legitimacy check is defined analogously to the edge\-level rewardRedgeR\_\{\\text\{edge\}\}\(see Eq\. \([4](https://arxiv.org/html/2608.07905#S3.E4)\)\):𝕍edge=Redge\\mathbb\{V\}\_\{\\text\{edge\}\}=R\_\{\\text\{edge\}\}\.

For any invalid subtask or transition, the corresponding error message is fed back to guide LLM correction\. ThisVerify→Feedback→Correctcycle iterates until the plan passes validation or reaches the maximum iteration limit \(*i\.e\.*3\)\.

### III\-BMemory\-Aware Low\-Level Action Policy

![Refer to caption](https://arxiv.org/html/2608.07905v1/x2.png)Figure 2:Qualitative example illustrating the benefits of the memory\-aware low\-level action policy\.The low\-level policy translates high\-level subtasks into executable primitive actions, dynamically grounded through environmental perception\. To enhance navigation efficiency and prevent redundant interactions, we develop amemory\-aware low\-level action policy\.While prior works \(e\.g\.,\[[16](https://arxiv.org/html/2608.07905#bib.bib6)\]\) have explored memory\-enhanced low\-level policy, our approach advances in proactive caching and instance\-level discrimination\.Unlike methods that primarily record past interactions, our agent proactively detects and caches information for objects that may require future interaction during execution\.When the agent encounters other task\-related objects on the way to the primary target object, it logs its current location as a candidate waypoint and caches the corresponding segmentation mask, in addition to saving the position and mask of the target object\. Thus, subsequent subtasks can directly utilize cached positions instead of re\-exploring, improving navigation consistency and efficiency\. For distinguishing multiple instances of the same type, individual interaction records are maintained to ensure precise identification and consistent placement\. As shown in Fig\.[2](https://arxiv.org/html/2608.07905#S3.F2), this memory mechanism allows detecting and recording secondary ‘Glassbottle2’ during the execution of ‘Pickup\(Glassbottle1\)’\. Subsequently, when executing ‘Put\(Glassbottle2, Shelf\)’, the agent prioritizes the shelf position retrieved from memory instead of re\-exploring, improving navigation efficiency and placement accuracy\.

### III\-CDynamic Replanning with Scene Graph

Although high\-level planning with the task graph yields high\-quality initial plans, it is not well grounded in dynamic environments and cannot readily adapt to newly emerged failures\. To address this, we design an event\-driven dynamic replanning module grounded by a dynamic scene graph, which invokes a replanning reasoning process upon either low\-level execution errors or high\-level subtask completion to determine whether the current plan should be kept or revised\.

#### III\-C1Task\-Aware Scene Graph Memory

To support closed\-loop reasoning and event\-driven replanning, GraphThink maintains a dynamic scene graph𝒢scene\\mathcal\{G\}\_\{\\rm scene\}as a task\-aware environmental memory\. Unlike complete scene representations\[[10](https://arxiv.org/html/2608.07905#bib.bib53),[39](https://arxiv.org/html/2608.07905#bib.bib54)\],𝒢scene\\mathcal\{G\}\_\{\\text\{scene\}\}is a compact, dynamically updated set of semantic triples\(oi,ri​j,oj\)\(o\_\{i\},r\_\{ij\},o\_\{j\}\)that describe relationsri​j∈ℛr\_\{ij\}\\in\\mathcal\{R\}\(*e\.g\.*, “on”, “next to”\) between task\-relevant objects, with key attributes \(*e\.g\.*, color and shape\) as node attributes\. This task\-focused representation reduces irrelevant visual noise and keeps the graph scale manageable \(typically fewer than 20 nodes in ALFRED\), while retaining the environmental evidence required for replanning\. The construction of𝒢scene\\mathcal\{G\}\_\{\\rm scene\}proceeds online during execution and does not require a pre\-built scene graph: \(i\)Key Object Extraction\.A language model parses the instructionτ\\tauto obtain𝒪key⊆𝒪env\\mathcal\{O\}\_\{\\text\{key\}\}\\subseteq\\mathcal\{O\}\_\{\\text\{env\}\}\. \(ii\)Viewpoint Capture\.During execution, record the agent’s egocentric views that contain any object in𝒪key\\mathcal\{O\}\_\{\\rm key\}\. To avoid unreliable relation captures from extreme viewpoints, we only perform relation extraction when the target object lies in the agent’s forward\-facing view, i\.e\., the angular deviation between the agent’s heading and the direction to the object center is within45∘45^\{\\circ\}\. \(iii\)Candidate Triple Generation\.The captured RGB image and the corresponding segmentation map are jointly provided to the VLM\. The segmentation map helps align VLM\-generated labels with environment object instances and reduces ambiguity caused by synonymous or visually similar objects\. The VLM then produces candidate relation triples and object attributes\. \(iv\)Scene Graph Update\.Candidate entries are filtered and incrementally merged into𝒢scene\\mathcal\{G\}\_\{\\rm scene\}as execution proceeds\. During execution, interaction\-induced changes also update the graph by replacing outdated relations, enabling the memory to adapt to object state changes and post\-interaction environmental updates\.

To reduce the impact of VLM hallucinations and perception errors, GraphThink adopts preventive filtering and conflict\-aware correction\. Before insertion, candidate entries are filtered by task relevance and rule\-based physical\-plausibility constraints; unreliable triples, such as\(SideTable,on,CounterTop\)\(\\texttt\{SideTable\},\\texttt\{on\},\\texttt\{CounterTop\}\), or task\-irrelevant entries are pruned\. When a new observation conflicts with existing relations, GraphThink does not keep mutually inconsistent entries simultaneously\. Instead, the agent re\-observes the scene from another viewpoint and asks the VLM to reassess the conflicting entries\. If some errors still enter the scene graph, they usually appear as execution failures or later contradictory observations, which trigger our event\-driven replanning module and subsequent scene graph updates\. This prevents corrupted memory from silently accumulating over long horizons\.

In this way,𝒢scene\\mathcal\{G\}\_\{\\text\{scene\}\}serves not as a static map, but as a dynamically corrected memory that provides compact and reliable grounding for downstream replanning\.

#### III\-C2Event\-driven Replanning

We aim for the agent to interleave execution with reasoning, continuously evaluating plan feasibility against environmental feedback and making necessary adjustments\. Unlike recent planning methods that primarily revise plans after execution failures\[[13](https://arxiv.org/html/2608.07905#bib.bib62),[19](https://arxiv.org/html/2608.07905#bib.bib8)\], GraphThink uses an LLM\-based replanning module driven by two complementary events: low\-level action errors and high\-level subtask completion \(Fig\.[3](https://arxiv.org/html/2608.07905#S3.F3)\)\. The former addresses actionable execution failures, while the latter proactively checks semantically valid but goal\-misaligned plans at subtask checkpoints\. Within this module, the task graph constrains static transition feasibility over executable subtasks, and the scene graph provides task\-relevant environmental context from dynamic observations, thereby reducing unconstrained LLM free\-generation during replanning\. When activated, the model assesses whether the plan should be adjusted based on the original plan, current subtask, feedback, and scene graph\. Execution continues if no adjustment is needed, or switches to the revised plan otherwise\.

![Refer to caption](https://arxiv.org/html/2608.07905v1/x3.png)Figure 3:The two examples illustrate two types of event\-driven dynamic replanning\.Low\-Level Action Error\.This trigger handles concrete execution failures, detected via feedback signals from the environment \(*e\.g\.*, interaction errors\) or from the navigation policy \(*e\.g\.*, exhaustive search failure\)\. When such an error occurs, the reasoning LLM integrates the current scene observations with commonsense knowledge to revise the plan\. For the instance in Fig\.[3](https://arxiv.org/html/2608.07905#S3.F3), if a potato cannot be found for slicing, the LLM might suggest searching inside closed containers like a refrigerator or microwave\. Notably, replanning is selectively invoked rather than triggered by every execution error\. Common recoverable failures, such as navigation collisions, are first handled by the low\-level policy, and only unresolved cases are escalated to the replanning module\. Detailed execution of this case is shown in Fig\.[4](https://arxiv.org/html/2608.07905#S3.F4)\.

![Refer to caption](https://arxiv.org/html/2608.07905v1/x4.png)Figure 4:An example of error\-triggered replanning\. When the agent fails to locate a potato during environmental exploration while executing slice\(potato\), GraphThink revises the plan and then attempts to open the refrigerator to search for one\.High\-Level Subtask Completion\.This trigger detects executable but goal\-misaligned plans through LLM\-based semantic assessment at subtask checkpoints\. The LLM assesses the current plan against the continuously updated scene graph memory and the plan execution progress to verify alignment with the final objective\. For example, for the goal “place the lettuce on the microwave oven table,” if the scene graph indicates the microwave is on a side table but the plan targets a countertop, the LLM identifies this mismatch and triggers a correction\. This allows proactive rectification of high\-level planning errors undetectable by low\-level error feedback, using subtask completion as checkpoints\.

## IVExperiments

In this section, we first validate GraphThink’s hierarchical framework for vision\-language navigation and interaction, and further assess the performance of its high\-level planner\.

TABLE I:Comparison with SOTA methods on ALFRED test set\. Baseline results are from the official leaderboard or papers\. The “Step\-by\-step Inst\.” column denotes whether step\-by\-step instructions are used in high\-level planning\.MethodStep\-by\-stepInst\.Tests SeenTests UnseenSR↑\\uparrowGC↑\\uparrowSR↑\\uparrowGC↑\\uparrowHLSM\[[3](https://arxiv.org/html/2608.07905#bib.bib13)\]✓29\.9441\.2120\.2730\.31LLM\-planner\[[36](https://arxiv.org/html/2608.07905#bib.bib9)\]✓18\.2026\.7716\.4223\.37CAPEAM\[[16](https://arxiv.org/html/2608.07905#bib.bib6)\]✓52\.5860\.9850\.3661\.40DISCO\[[46](https://arxiv.org/html/2608.07905#bib.bib2)\]✓59\.5966\.0656\.5566\.87Flare\[[19](https://arxiv.org/html/2608.07905#bib.bib8)\]✓40\.0548\.8440\.8851\.72Prompter\[[15](https://arxiv.org/html/2608.07905#bib.bib7)\]✗47\.9556\.9841\.5353\.69DISCO\[[46](https://arxiv.org/html/2608.07905#bib.bib2)\]✗58\.0564\.9654\.7765\.56OPEx\[[32](https://arxiv.org/html/2608.07905#bib.bib11)\]✗43\.5154\.2741\.2753\.82EPO\[[49](https://arxiv.org/html/2608.07905#bib.bib4)\]✗64\.7972\.3062\.3567\.52RoboGPT\[[5](https://arxiv.org/html/2608.07905#bib.bib5)\]✗59\.9267\.8362\.0072\.09GraphThink\(ours\)✗67\.7173\.9668\.5275\.76### IV\-AComparison with SOTA Methods on ALFRED

Benchmark and Metrics\.We conduct experiments on ALFRED\[[33](https://arxiv.org/html/2608.07905#bib.bib3)\], a challenging benchmark for robotics instruction following\. The language instructionL=\(Lhigh,Llow\)L=\(\{L\_\{\\text\{high\}\},L\_\{\\text\{low\}\}\}\)contains both high\-level goals and step\-by\-step guidance\. ALFRED is partitioned into ‘train’, ‘validation’ and ‘test’ sets\. Both ‘validation’ and ‘test’ are further divided into seen and unseen splits, where the unseen partitions consist of scenes absent from the training set\. The dataset includes 7 task types, 58 target object classes, and 26 receptacle classes distributed across 120 indoor scenes\. Objects within the same class often exhibit various visual appearances \(*e\.g\.*, there are 30 varieties of apples\), and the indoor scenes cover kitchens, bathrooms, bedrooms, and living rooms\. According to the official statistics, the training set consists of 21023 examples, the validation seen set contains 820 examples, the validation unseen set contains 821 examples, the test seen set has 1533 examples, and the test unseen set has 1529 examples\. We adopt two evaluation metrics\. The primary metric is Success Rate \(SR\), which measures the percentage of fully completed tasks\. Additionally, Goal\-Condition \(GC\) success rate evaluates the percentage of satisfied goal conditions\.

Performance\.To demonstrate the effectiveness of our approach, we compare GraphThink to competitive works in the test set reported on ALFRED public leaderboard\. Following\[[19](https://arxiv.org/html/2608.07905#bib.bib8),[5](https://arxiv.org/html/2608.07905#bib.bib5)\], we report baselines using 1\) only the goal instruction, and 2\) both the goal instruction and step\-by\-step instructions\. As shown in Table[I](https://arxiv.org/html/2608.07905#S4.T1), our approach significantly outperforms prior works by 6\.17 percentage points on unseen tasks, while achieving SOTA performance across all metrics in both unseen and seen environments \(reaching 67\.71% and 68\.52%, respectively\), demonstrating the effectiveness of our approach\. Moreover, using only the goal instruction, GraphThink outperforms prior methods under both settings, i\.e\., with and without step\-by\-step instructions, highlighting the superiority of our high\-level planning module\. It is worth emphasizing that GraphThink does not rely on metadata for environmental grounding, indicating its stronger generalization capability to real\-world scenarios\.

### IV\-BAblations of GraphThink

TABLE II:Ablation studies on three components of GraphThink’s hierarchical framework\.ComponentValid SeenValid UnseenSR↑\\uparrowGC↑\\uparrowSR↑\\uparrowGC↑\\uparrowreplaced planner40\.7046\.1941\.5947\.82w/o replan60\.1868\.3959\.8268\.10w/o replan\_low64\.3172\.1364\.1071\.58w/o replan\_high63\.5871\.6564\.2572\.10w/o memory67\.2674\.3966\.5974\.13GraphThink68\.3675\.3268\.7276\.01We conduct an ablation study of GraphThink’s components, with experimental results on the ALFRED validation set shown in Table[II](https://arxiv.org/html/2608.07905#S4.T2)\. \(i\) Replacing our task graph\-based planner with a RAG\-enhanced Qwen2\.5\-7B\-Instruct planner \(Row 1\) leads to significant performance drops of 27\.66% and 27\.13% on seen and unseen splits, respectively, underscoring the importance of high\-quality initial planning\. \(ii\) Removing the replanning module results in open\-loop execution that cannot adapt to unexpected errors or environmental feedback, reducing success rates by 8\.18% and 8\.90%\. We then separately ablate the two types of event\-driven replanning\. Without low\-level execution error driver \(Row 3\), the agent fails to handle execution errors \(e\.g\., wandering due to objects in closed containers or detection failures\), reducing SR by 4\.05% and 4\.62%\. Without high\-level subtask completion driver \(Row 4\), the agent cannot verify and correct plans against environmental information, causing SR drops of 4\.78% and 4\.47%\. \(iii\) Disabling the object state and location memory in the low\-level policy \(Row 5\) degrades performance by impairing object localization and introducing interaction errors\. In tasks requiring placement of two identical objects into the same container, the agent may move one object repeatedly or misplace them in different locations, resulting in task failures\. These results demonstrate that GraphThink ensures robust execution through: task graph\-based planning for initial plans, dynamic replanning for maintaining executability and instruction alignment, and memory\-aware navigation for reliable interaction\.

### IV\-CEvaluations of High\-Level Planning

This subsection evaluates GraphThink’s high\-level planner by comparing it with strong LLM\-based planners and conducting extensive ablation studies on its key components and training design\.

#### IV\-C1Comparison with SOTA planners

![Refer to caption](https://arxiv.org/html/2608.07905v1/x5.png)Figure 5:Success rates for high\-level planning on ALFRED validation set and long\-horizon tasks\. For brevity, we adopt the abbreviations Zs for Zero\-shot and Fs for Few\-shot\.Since ALFRED has limited task diversity, we construct a long\-horizon dataset of 1396 samples by extending original 7 short task types with 17 complex types to better evaluate unseen long\-horizon planning \(details in Appendix B\)\. To evaluate GraphThink’s high\-level planner, we compare GraphThink against multiple strong LLMs \(*i\.e\.*, Deepseek\-R1, Gemini\-2\.5\-Pro, and GPT\-5\.2\) under diverse inference settings \(*i\.e\.*, zero\-shot, CoT, few\-shot, RAG\)\. For fair comparison, we conduct controlled evaluations on the same Qwen2\.5\-7B\-Instruct backbone, comparing GraphThink with multiple inference strategies, SFT, and task\-graph search baselines \(Appendix C\)\. For high\-level planning, success rate is the primary metric\. As a complementary evaluation, Section[IV\-E](https://arxiv.org/html/2608.07905#S4.SS5)employs multiple fine\-grained metrics to validate GraphThink’s robustness, efficiency, and reasoning quality across increasingly complex tasks\. Since multiple feasible paths often exist in task planning and certain GT annotations in datasets may not accurately match the instructions, relying solely on the single GT in the validation set is insufficient\. Therefore, we incorporate the task graph as a verification tool and propose a more comprehensive validation method\. The procedure begins with graph\-constrained verification and GT matching\. If both checks pass, the plan is successful\. If graph verification fails, the plan is marked as failed\. When the plan satisfies graph constraints but diverges from GT, it undergoes further validation using an LLM\.

How does GraphThink compare to state\-of\-the\-art LLMs?As shown in Fig\.[5](https://arxiv.org/html/2608.07905#S4.F5), GraphThink achieves superior performance across all settings, with 96\.22% on valid seen, 96\.56% on valid unseen and 90\.04% on long horizon tasks\. Although prompt\-driven variants such as few\-shot CoT and RAG yield substantial improvements upon stronger LLMs like GPT\-5\.2, GraphThink still surpasses all such methods by a significant margin\.

Does GraphThink improve long\-horizon reasoning?Across baselines, performance drops markedly on long\-horizon tasks, whereas GraphThink remains robust\. Detailed analysis in Section[IV\-E](https://arxiv.org/html/2608.07905#S4.SS5)further shows that, as task horizons increase, the baselines suffer pronounced degradation in planning success and reasoning fidelity, while GraphThink maintains stronger robustness and high inference efficiency\. This highlights the benefit of the task graph for reliable and efficient long\-horizon embodied planning\.

TABLE III:Qualitative comparison of long\-horizon task planning performance with different training methods\.Instruction:Put a cleaned tomato slice in a bowl on the top shelf, and put a pan containing a heated potato on the dining table\.GraphThink \(Ours\):\(PickupObject, Knife\), \(SliceObject, Tomato\), \(PutObject, \(Knife, SinkBasin\)\), \(PickupObject, TomatoSliced\), \(CleanObject, \(TomatoSliced, SinkBasin\)\), \(PutPickObject, \(TomatoSliced, Bowl\)\), \(PutObject, \(Bowl, Shelf\)\), \(PickupObject, Potato\), \(HeatObject, \(Potato, Microwave\)\), \(PutPickObject, \(Potato, Pan\)\), \(PutObject, \(Pan, DiningTable\)\)Success: all goals are satisfied with valid transitions\.GRPO without Task Graph Rewards:\(PickupObject, Knife\), \(PickupObject, Potato\), \(HeatObject, \(Potato, Microwave\)\), \(PutObject, \(Potato, Pan\)\), \(PickupObject, Knife\), \(PickupObject, Tomato\), \(SliceObject, Tomato\), \(PutObject, \(Knife, SinkBasin\)\), \(PickupObject, TomatoSliced\), \(CleanObject, \(TomatoSliced, SinkBasin\)\), \(PutObject, \(TomatoSliced, Bowl\)\), \(PutObject, \(Bowl, Shelf\)\), \(PutObject, \(Pan, DiningTable\)\)Error: invalid consecutive pickup/put transitions and missing pan pickup before table placement\.SFT:\(PickupObject, Knife\), \(SliceObject, Tomato\), \(PutObject, \(TomatoSliced, Bowl\)\), \(PutObject, \(Bowl, Shelf\)\), \(PickupObject, Potato\), \(HeatObject, \(Potato, Microwave\)\), \(PutPickObject, \(Potato, Pan\)\), \(PutObject, \(Pan, DiningTable\)\)Error: missing tomato cleaning and knife placement, with an invalid consecutive put transition\.Does GraphThink outperform other strategies under the same backbone?With the Qwen2\.5\-7B\-Instruct backbone, SFT is strong on validation but drops sharply to 20\.57% on long\-horizon tasks\. Due to high similarity between ALFRED’s training and validation sets, SFT handles in\-domain short\-term planning adequately but struggles with complex long\-horizon tasks where planning hallucinations and missing steps prevail\. In contrast, GraphThink achieves 90\.04% on long\-horizon tasks, demonstrating that task graph\-based RL generalizes effectively to complex task compositions without requiring costly supervised fine\-tuning on expert\-annotated long\-horizon trajectories\. To further qualitatively analyze long\-horizon reasoning behavior, Table[III](https://arxiv.org/html/2608.07905#S4.T3)compares GraphThink with two in\-domain training baselines, including SFT and GRPO without task graph rewards\. The cases show that GraphThink reduces planning hallucinations and missing steps in complex instructions, demonstrating the effectiveness of task graph constraints and graph\-based reward signals\. See Appendix D for more cases\. We also compare with task‑graph search baselines \(Appendix C\), where stepwise expansion often accumulates myopic errors and leads to sub‑optimal paths in long‑horizon settings, making naive search markedly less effective than our integrated LLM‑based planning\.

#### IV\-C2Ablations of our high\-level planner

TABLE IV:Success rate contributions of individual components in the proposed planner\.MethodValidSeenValidUnseenLongHorizonw/o Task\-Graph Prompt77\.9374\.1140\.69w/o GRPO49\.8852\.6433\.67w/o Verification95\.1295\.0986\.25All96\.2296\.5690\.04We ablate the three task\-graph modules in the high\-level planner to assess their contributions \(Table[IV](https://arxiv.org/html/2608.07905#S4.T4)\): \(i\)Prompt with Task Graph\.Removing the task graph from the prompt during inference significantly reduces planning accuracy, showing that the graph enhances the LLM’s ability to generate grounded plans and reduces hallucinations\. \(ii\)Graph GRPO with Task Graph\.Replacing the GRPO\-trained model with the base Qwen2\.5\-7B\-Instruct causes significant performance degradation, highlighting the importance of our task graph\-based reinforcement learning framework\. \(iii\)Verification with Task Graph\.Verification feedback has little effect on the validation set, since initial plans generally satisfy graph constraints\. However, on long\-horizon tasks where constraints are sometimes violated, error feedback improves performance by 3\.79%\. The verification loop is also bounded to at most 3 iterations, making it a low\-cost guardrail against residual planning hallucinations\. Notably, even without refinement, our planner remains competitive and still outperforms baselines\.

#### IV\-C3Ablations of GRPO

TABLE V:Ablation studies of GRPO components and the two\-stage instruction reward design\.SettingValidSeenValidUnseenLongHorizonAblations of GRPO Componentsw/oRnode\+RedgeR\_\{\\text\{node\}\}\+R\_\{\\text\{edge\}\}93\.0589\.3344\.63w/oRnodeR\_\{\\text\{node\}\}93\.5491\.6676\.50w/oRedgeR\_\{\\text\{edge\}\}89\.8885\.0350\.50w/oRinstR\_\{\\text\{inst\}\}84\.2784\.6655\.59w/oRformatR\_\{\\text\{format\}\}91\.7189\.8270\.63w/o Task\-Graph Prompt85\.0083\.4459\.46Ablations of the two\-stageRinstR\_\{\\text\{inst\}\}Critical\-onlyRinstR\_\{\\text\{inst\}\}89\.7690\.6772\.85GT\-onlyRinstR\_\{\\text\{inst\}\}91\.7190\.0668\.84Full GRPO95\.1295\.0986\.25![Refer to caption](https://arxiv.org/html/2608.07905v1/x6.png)Figure 6:Ablation studies on training data volume\.We conduct a quantitative ablation study to analyze key components in GRPO and further ablate the two\-stage design of the instruction\-following rewardRinstR\_\{\\text\{inst\}\}, as shown in Table[V](https://arxiv.org/html/2608.07905#S4.T5)\. The “Full GRPO” setting indicates success rates with full rewards and the task graph prompt\. Results here differ from Fig\.[5](https://arxiv.org/html/2608.07905#S4.F5)as we remove validation feedback to isolate training effects\. Our findings are: \(i\)Task graph reward enhances generalization and zero\-shot learning:Without the graph reward \(i\.e\., “w/oRnode\+RedgeR\_\{\\text\{node\}\}\+R\_\{\\text\{edge\}\}”\), reliance on only the instruction reward restricts learning to task paths in the dataset\. By contrast, the graph reward encourages exploration of all feasible subtask transitions, as the task graph encapsulates diverse possible paths\. Its absence hinders long\-horizon planning, confirming that the graph reward significantly improves generalization tocomplex task combinations beyond training coverage\. Even without the instruction reward \(i\.e\., “w/oRinstR\_\{\\text\{inst\}\}”\), using only unlabeled data and graph rewards achieves an average 84\.47% on the valid set, demonstrating the graph reward’s ability to guide high\-level action learning and impart basic planning capabilities in zero\-shot scenarios\. \(ii\)Edge\-level transition constraints are especially important:The fine\-grained ablations ofRnodeR\_\{\\text\{node\}\}andRedgeR\_\{\\text\{edge\}\}further show that removing either the node\-level reward or the edge\-level reward results in performance decline\. Notably, the absence ofRedgeR\_\{\\text\{edge\}\}causes a more significant drop, indicating that constraining subtask transitions through graph edges is crucial for grounded subtask decomposition\. \(iii\)Instruction reward aligns user intent:Without the instruction following reward, the model may engage in reward hacking by solely generating plans that satisfy graph constraints while disregarding actual instruction requirements\. These results underscore the essential role of the instruction following reward in ensuring semantic alignment between the generated plans and the original task instruction\. Furthermore, using only critical\-step coverage \(i\.e\., “Critical\-only”\) or only exact match \(i\.e\., “GT\-only”\) both underperform compared to the progressive design, especially on long\-horizon tasks\. This indicates that the two\-stage design ofRinstR\_\{\\text\{inst\}\}improves semantic alignment without inducing overfitting to annotated paths\. \(iv\)Format reward supports structural output consistency:When the format reward is ablated, a slight performance drop is observed\. This suggests that although other rewards can partially guide the model to produce structured outputs, explicit format supervision helps ensure structure correctness\. \(v\)Incorporating the task graph into prompts enhances training effectiveness:When we remove the explicit task graph from the prompt while keeping all other training and inference settings unchanged \(i\.e\., “w/o Task\-Graph Prompt”\), performance drops markedly on both the validation set and long\-horizon tasks \(average declines of 10\.89% and 26\.79%, respectively\)\. These results confirm that the task graph not only aids model reasoning but also substantially improves training effectiveness\.

Impact of training data volume\.To evaluate the impact of training data volume, we train the planning model using subsets corresponding to 1%, 10%, 20%, and 50% of the full training samples\. These subsets cover all seven task types to ensure a fair representation of the training set\. As shown in Fig\.[6](https://arxiv.org/html/2608.07905#S4.F6), model performance improves with increasing amounts of training data\. The performance gain is relatively modest on the validation set, indicating that our planning method can effectively strengthen short\-task \(*i\.e\.*, validation\-set\) planning even with limited data\. In contrast, the improvement is more pronounced on long\-horizon tasks, suggesting that scaling the training data helps LLMs learn the richer action space required for complex planning\.

![Refer to caption](https://arxiv.org/html/2608.07905v1/x7.png)Figure 7:An example of GraphThink executing a newly composed task involving unseen action primitives in AI2\-Thor\. The new tasks ‘DirtyObject’ and ‘BreakObject’ are highlighted in blue\.

### IV\-DGeneralization to Novel Tasks and Environments

To systematically evaluate GraphThink’s generalizability beyond benchmark\-specific settings, we conduct experiments along two complementary dimensions: \(i\) generalization to novel tasks in AI2\-Thor, testing the ability to incorporate unseen action primitives; and \(ii\) cross\-environment generalization to different task settings \(i\.e\., WAH\-NL\[[7](https://arxiv.org/html/2608.07905#bib.bib58)\]and VirtualHome\-HG\[[23](https://arxiv.org/html/2608.07905#bib.bib59)\]in the VirtualHome environment\), evaluating the transferability of our planning methodology to new domains\. Following the best\-performing baselines in Sec\.[IV\-C1](https://arxiv.org/html/2608.07905#S4.SS3.SSS1), we compare GraphThink with: \(i\) the leading RAG\-enhanced LLM \(GPT\-5\.2\); and \(ii\) methods using the same backbone \(Qwen2\.5\-7B\-Instruct\), including supervised fine\-tuning \(SFT\) and few\-shot prompting augmented with chain\-of\-thought \(CoT\) reasoning, to ensure a fair comparison\.

Generalization to novel tasks within AI2\-Thor\.We test GraphThink’s ability to handle new action primitives by expanding its task graph with 9 new actions supported by the underlying AI2\-Thor simulator but absent from ALFRED \(e\.g\., BreakObject, FillObject, ThrowObject\)\. Following a similar construction methodology as the Long\-Horizon Dataset, we create 10 new task types combining these novel skills \(e\.g\., Break&Throw, Clean&Fill&Heat/Cool&Place\), with 400 samples\. By extending the task graph with transition compatibility constraints, we evaluate GraphThink’s zero\-shot generalization to these novel tasks\. As shown in Table[VI](https://arxiv.org/html/2608.07905#S4.T6), GraphThink substantially outperforms baselines on the newly composed tasks\. Fig\.[7](https://arxiv.org/html/2608.07905#S4.F7)visualizes an example execution of the “Dirty&Clean&Break&Cool&Place” task in the scene, illustrating how GraphThink composes newly introduced actions with existing manipulation skills\.

TABLE VI:Generalization performance across new tasks\.MethodModelNewTasksRAGGPT\-5\.265\.75Fs\+CoTQwen2\.5\-7B\-Instruct16\.75SFTQwen2\.5\-7B\-Instruct16\.25GraphThinkQwen2\.5\-7B\-Instruct80\.50TABLE VII:Cross\-Environment Generalization to VirtualHome\.MethodModelWAH\-NLVirtualHome\-HGRAGGPT\-5\.292\.0088\.89Fs\+CoTQwen2\.5\-7B\-Instruct32\.0017\.78SFTQwen2\.5\-7B\-Instruct87\.0078\.89Ours\(ALF\)Qwen2\.5\-7B\-Instruct90\.0076\.67Ours\(VH\)Qwen2\.5\-7B\-Instruct93\.0088\.89![Refer to caption](https://arxiv.org/html/2608.07905v1/x8.png)Figure 8:An example of GraphThink for the task ‘Heat a clean apple and then put it on the dining table’ in VirtualHome\.Cross\-Environment Generalization to VirtualHome\.To further validate the versatility of the proposed GraphThink in broader robotic task applications, we test it on two additional benchmarks built on the VirtualHome simulator, including 100 rearrangement tasks from WAH‑NL and 90 VirtualHome‑HG tasks covering cooking, cleaning, and laundry\. We test two adaptation strategies:\(i\) Domain\-specific adaptation:VirtualHome has a partially different action space and interaction primitives from ALFRED\. In this case, we construct a new task graph from VirtualHome’s action primitives, and the planner is trained with our framework using in\-domain data\.\(ii\) Cross\-domain transfer:Although the low\-level primitives differ, most high\-level subgoal semantics can be shared across embodied environments\. We apply the ALFRED\-trained planner and map abstract subtasks to VirtualHome’s action programs via a predefined translation layer that aligns semantic action types between domains\. If novel subtask types that do not exist are encountered \(e\.g\., laundry\), the task graph can be extended with corresponding nodes to support the new skills\.

Baselines are fairly compared using corresponding in\-domain VirtualHome data\. Results in Table[VII](https://arxiv.org/html/2608.07905#S4.T7)show that GraphThink with domain\-specific adaptation \(VH\) delivers the best performance, while the cross\-domain variant \(ALF\) remains competitive, indicating that our high\-level planner generalizes well to VirtualHome despite differences in action space and task semantics\. These results further suggest that \(i\) high\-level planning priors learned in the ALFRED domain are reusable across embodied simulators, and \(ii\) re\-instantiating the planning framework for the target domain further enhances cross\-environment generalization, validating the effective cross\-domain transferability of GraphThink\. Fig\.[8](https://arxiv.org/html/2608.07905#S4.F8)further provides a qualitative execution example in the VirtualHome environment\.

### IV\-EPerformance Analysis across Varying Task Horizons

To comprehensively evaluate the robustness and scalability of our approach, we analyze GraphThink’s performance across tasks of increasing complexity \(7–12 subtasks\), comparing it against two strong baselines: GPT\-5\.2 enhanced with RAG and a SFT model based on the same Qwen2\.5\-7B\-Instruct backbone\. The results are summarized in Fig\.[9](https://arxiv.org/html/2608.07905#S4.F9)\. All metrics are reported as averages\. We use the following abbreviations: Acc \(Accuracy\), GPR \(Graph Pass Rate\), MSR \(Missing Step Rate\), ASR \(Additional Step Rate\), WTR \(Wrong Transfer Rate\), and AER \(Affordance Error Rate\)\. MSR, ASR, WTR, and AER are computed per task instance and then averaged over all tasks with the same horizon length, reflecting per\-task reasoning fidelity\.

![Refer to caption](https://arxiv.org/html/2608.07905v1/x9.png)Figure 9:Performance comparison across increasing task horizons\. Subfigure \(a\) reports planning success rate by accuracy, and subfigure \(b\) shows graph pass rate\. Subfigure \(c\) compares computational costs measured by inference time\. Subfigure \(d\) presents fine\-grained reasoning\-fidelity metrics, including missing step rate \(MSR\), additional step rate \(ASR\), wrong transfer rate \(WTR\), and affordance error rate \(AER\)\. Qwen2\.5\-7B\-Instruct\+GraphThink shows stronger robustness and planning fidelity than the baselines, while achieving substantially lower inference time than GPT\-5\.2\+RAG on long\-horizon tasks\.\(i\)Performance Scaling: As shown in Fig\.[9](https://arxiv.org/html/2608.07905#S4.F9)\(a\), Qwen2\.5\-7B\-Instruct\+GraphThink shows only moderate degradation as task horizons increase, with accuracy decreasing from 98\.73% to 80\.71%\. Meanwhile, Fig\.[9](https://arxiv.org/html/2608.07905#S4.F9)\(b\) shows that its graph pass rate remains at or above 85% across all horizons\. This contrasts sharply with the baselines: GPT\-5\.2 with RAG drops from 82\.59% to 22\.40%, while Qwen2\.5\-7B\-Instruct\+SFT drops from 40\.63% and remains near zero from 9 subtasks onward\.

\(ii\)Computational Costs: As shown in Fig\.[9](https://arxiv.org/html/2608.07905#S4.F9)\(c\), Qwen2\.5\-7B\-Instruct\+GraphThink maintains stable inference time \(∼\\sim7–10s\) for tasks with 7–11 subtasks, demonstrating well\-controlled overhead\. Even at 12 subtasks, the average inference time only increases to 23\.7s, which remains lower than GPT\-5\.2 with RAG \(32\.78–46\.90s\)\. Although Qwen2\.5\-7B\-Instruct\+SFT has low inference time, its success rate and reasoning fidelity degrade severely on long\-horizon tasks\. These results collectively validate the robustness, efficiency, and reasoning quality of Qwen2\.5\-7B\-Instruct\+GraphThink on increasingly complex long\-horizon tasks\.

\(iii\)Reasoning Fidelity: Fig\.[9](https://arxiv.org/html/2608.07905#S4.F9)\(d\) reports four fine\-grained error metrics following\[[22](https://arxiv.org/html/2608.07905#bib.bib60)\], which help identify specific weaknesses in LLM planning\. Qwen2\.5\-7B\-Instruct\+GraphThink consistently achieves superior planning quality, with remarkably low wrong transfer rates \(0\.02–2\.16%\) and near\-zero affordance errors, indicating accurate action–object grounding under graph constraints\. Although Qwen2\.5\-7B\-Instruct\+GraphThink shows moderate fluctuations in missing and additional step rates on longer horizons, these errors remain well controlled compared with the baselines\. For example, its ASR peaks at 8\.81% at 11 subtasks, whereas GPT\-5\.2 with RAG reaches a much higher ASR of 46\.67% at the same horizon\.

## VConclusion

We introduce GraphThink, a general graph\-enhanced planning framework that improves LLM\-based embodied task planning\. Our method reduces planning hallucinations by grounding reasoning in the structured task graph and ensuring plan executability through event\-driven replanning with dynamic scene graphs\. Experimental results on ALFRED validate the effectiveness of our components, while tests on a new long\-horizon dataset demonstrate strong reasoning and generalization capabilities\. Additional experiments show promising zero\-shot learning potential\.Future work could extend this approach to asynchronous planning or multi\-robot collaborative systems through temporal\-state edge augmentation or multi\-layer graph construction\.

## References

- \[1\]M\. Ahn, A\. Brohan, N\. Brown, Y\. Chebotar, O\. Cortes, B\. David, C\. Finn, C\. Fu, K\. Gopalakrishnan, K\. Hausman,et al\.\(2022\)Do as i can, not as i say: grounding language in robotic affordances\.arXiv preprint arXiv:2204\.01691\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p2.1)\.
- \[2\]\(2023\)Multi\-level compositional reasoning for interactive instruction following\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.37,pp\. 223–231\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1)\.
- \[3\]V\. Blukis, C\. Paxton, D\. Fox, A\. Garg, and Y\. Artzi\(2022\)A persistent spatial semantic representation for high\-level natural language instruction execution\.InConference on Robot Learning,pp\. 706–717\.Cited by:[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.6.1)\.
- \[4\]M\. Cai, X\. Chen, Y\. An, J\. Zhang, X\. Wang, W\. Xu, W\. Zhang, and T\. Liu\(2025\)CookBench: a long\-horizon embodied planning benchmark for complex cooking scenarios\.arXiv preprint arXiv:2508\.03232\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1)\.
- \[5\]Y\. Chen, W\. Cui, Y\. Chen, M\. Tan, X\. Zhang, J\. Liu, H\. Li, D\. Zhao, and H\. Wang\(2025\)Robogpt: an llm\-based long\-term decision\-making embodied agent for instruction following tasks\.IEEE Transactions on Cognitive and Developmental Systems\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p2.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[§IV\-A](https://arxiv.org/html/2608.07905#S4.SS1.p2.1),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.15.1)\.
- \[6\]S\. Cheng, Z\. Wu, J\. Chen, Z\. Li, Y\. Liu, and L\. Kong\(2023\)Unsupervised explanation generation via correct instantiations\.InProceedings of the AAAI conference on artificial intelligence,Vol\.37, number 11,pp\. 12700–12708\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[7\]J\. Choi, Y\. Yoon, H\. Ong, J\. Kim, and M\. Jang\(2024\)LoTa\-bench: benchmarking language\-oriented task planners for embodied agents\.InInternational Conference on Learning Representations \(ICLR\) 2024,pp\. 1–27\.Cited by:[§IV\-D](https://arxiv.org/html/2608.07905#S4.SS4.p1.1)\.
- \[8\]K\. Ehsani, T\. Gupta, R\. Hendrix, J\. Salvador, L\. Weihs, K\. Zeng, K\. P\. Singh, Y\. Kim, W\. Han, A\. Herrasti,et al\.\(2024\)Spoc: imitating shortest paths in simulation enables effective navigation and manipulation in the real world\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 16238–16250\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1)\.
- \[9\]Z\. Gou, Z\. Shao, Y\. Gong, Y\. Shen, Y\. Yang, N\. Duan, and W\. Chen\(2023\)Critic: large language models can self\-correct with tool\-interactive critiquing\.arXiv preprint arXiv:2305\.11738\.Cited by:[§III\-A4](https://arxiv.org/html/2608.07905#S3.SS1.SSS4.p1.1)\.
- \[10\]Q\. Gu, A\. Kuwajerwala, S\. Morin, K\. M\. Jatavallabhula, B\. Sen, A\. Agarwal, C\. Rivera, W\. Paul, K\. Ellis, R\. Chellappa,et al\.\(2024\)Conceptgraphs: open\-vocabulary 3d scene graphs for perception and planning\.In2024 IEEE International Conference on Robotics and Automation \(ICRA\),pp\. 5021–5028\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p4.1),[§III\-C1](https://arxiv.org/html/2608.07905#S3.SS3.SSS1.p1.10)\.
- \[11\]J\. Guan, W\. Wu, P\. Xu, H\. Wang, M\. Huang,et al\.\(2024\)Amor: a recipe for building adaptable modular knowledge agents through process feedback\.Advances in Neural Information Processing Systems37,pp\. 126118–126148\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[12\]D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi,et al\.\(2025\)Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§III\-A3](https://arxiv.org/html/2608.07905#S3.SS1.SSS3.p2.1)\.
- \[13\]J\. Huang, Y\. Xiao, Z\. Zhang, M\. Coates, J\. Hao, and Y\. Zhang\(2025\)One demo is all it takes: planning domain derivation with llms from a single demonstration\.arXiv preprint arXiv:2505\.18382\.Cited by:[§III\-C2](https://arxiv.org/html/2608.07905#S3.SS3.SSS2.p1.1)\.
- \[14\]W\. Huang, F\. Xia, T\. Xiao, H\. Chan, J\. Liang, P\. Florence, A\. Zeng, J\. Tompson, I\. Mordatch, Y\. Chebotar,et al\.\(2022\)Inner monologue: embodied reasoning through planning with language models\.arXiv preprint arXiv:2207\.05608\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p1.1.2)\.
- \[15\]Y\. Inoue and H\. Ohashi\(2022\)Prompter: utilizing large language model prompting for a data efficient embodied instruction following\.arXiv preprint arXiv:2211\.03267\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.11.1)\.
- \[16\]B\. Kim, J\. Kim, Y\. Kim, C\. Min, and J\. Choi\(2023\)Context\-aware planning and environment\-aware memory for instruction following embodied agents\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 10936–10946\.Cited by:[§III\-B](https://arxiv.org/html/2608.07905#S3.SS2.p1.1.2),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.8.1)\.
- \[17\]B\. Kim, M\. Seo, and J\. Choi\(2024\)Online continual learning for interactive instruction following agents\.arXiv preprint arXiv:2403\.07548\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1)\.
- \[18\]J\. Kim, C\. Min, B\. Kim, and J\. Choi\(2024\)Pre\-emptive action revision by environmental feedback for embodied instruction following agents\.In8th Annual Conference on Robot Learning,Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p1.1.2)\.
- \[19\]T\. Kim, B\. Kim, and J\. Choi\(2025\)Multi\-modal grounded planning and efficient replanning for learning embodied agents with a few examples\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39, number 4,pp\. 4329–4337\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1),[§I](https://arxiv.org/html/2608.07905#S1.p2.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1.2),[§III\-C2](https://arxiv.org/html/2608.07905#S3.SS3.SSS2.p1.1),[§IV\-A](https://arxiv.org/html/2608.07905#S4.SS1.p2.1),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.10.1)\.
- \[20\]T\. Kim, C\. Min, B\. Kim, J\. Kim, W\. Jeung, and J\. Choi\(2024\)ReALFRED: an embodied instruction following benchmark in photo\-realistic environments\.InEuropean Conference on Computer Vision,pp\. 346–364\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1)\.
- \[21\]E\. Kolve, R\. Mottaghi, W\. Han, E\. VanderBilt, L\. Weihs, A\. Herrasti, M\. Deitke, K\. Ehsani, D\. Gordon, Y\. Zhu,et al\.\(2017\)Ai2\-thor: an interactive 3d environment for visual ai\.arXiv preprint arXiv:1712\.05474\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p5.1.3)\.
- \[22\]M\. Li, S\. Zhao, Q\. Wang, K\. Wang, Y\. Zhou, S\. Srivastava, C\. Gokmen, T\. Lee, E\. L\. Li, R\. Zhang,et al\.\(2024\)Embodied agent interface: benchmarking llms for embodied decision making\.Advances in Neural Information Processing Systems37,pp\. 100428–100534\.Cited by:[§IV\-E](https://arxiv.org/html/2608.07905#S4.SS5.p4.1.1)\.
- \[23\]P\. Liu, L\. P\. Kaelbling, J\. B\. Tenenbaum, and J\. Mao\(2025\)Lifelong experience abstraction and planning\.InICML 2025 Workshop on Programmatic Representations for Agent Learning,Cited by:[§IV\-D](https://arxiv.org/html/2608.07905#S4.SS4.p1.1)\.
- \[24\]A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang,et al\.\(2023\)Self\-refine: iterative refinement with self\-feedback\.Advances in Neural Information Processing Systems36,pp\. 46534–46594\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[25\]S\. Y\. Min, D\. S\. Chaplot, P\. Ravikumar, Y\. Bisk, and R\. Salakhutdinov\(2021\)Film: following instructions in language with modular methods\.arXiv preprint arXiv:2110\.07342\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[§III\-A1](https://arxiv.org/html/2608.07905#S3.SS1.SSS1.p2.7)\.
- \[26\]I\. Obi, V\. L\. Venkatesh, W\. Wang, R\. Wang, D\. Suh, T\. I\. Amosa, W\. Jo, and B\. Min\(2025\)Safeplan: leveraging formal logic and chain\-of\-thought reasoning for enhanced safety in llm\-based robotic task planning\.arXiv preprint arXiv:2503\.06892\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[27\]A\. Pashevich, C\. Schmid, and C\. Sun\(2021\)Episodic transformer for vision\-and\-language navigation\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 15942–15952\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1)\.
- \[28\]X\. Puig, K\. Ra, M\. Boben, J\. Li, T\. Wang, S\. Fidler, and A\. Torralba\(2018\)Virtualhome: simulating household activities via programs\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 8494–8502\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p5.1.3)\.
- \[29\]C\. Qian, E\. C\. Acikgoz, Q\. He, H\. Wang, X\. Chen, D\. Hakkani\-Tür, G\. Tur, and H\. Ji\(2025\)Toolrl: reward is all tool learning needs\.arXiv preprint arXiv:2504\.13958\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1),[§III\-A3](https://arxiv.org/html/2608.07905#S3.SS1.SSS3.p1.1)\.
- \[30\]K\. Rana, J\. Haviland, S\. Garg, J\. Abou\-Chakra, I\. Reid, and N\. Suenderhauf\(2023\)Sayplan: grounding large language models using 3d scene graphs for scalable robot task planning\.arXiv preprint arXiv:2307\.06135\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p2.1)\.
- \[31\]Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p3.1)\.
- \[32\]H\. Shi, Z\. Sun, X\. Yuan, M\. Côté, and B\. Liu\(2024\)Opex: a component\-wise analysis of llm\-centric agents in embodied instruction following\.arXiv preprint arXiv:2403\.03017\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.13.1)\.
- \[33\]M\. Shridhar, J\. Thomason, D\. Gordon, Y\. Bisk, W\. Han, R\. Mottaghi, L\. Zettlemoyer, and D\. Fox\(2020\)Alfred: a benchmark for interpreting grounded instructions for everyday tasks\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 10740–10749\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p5.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[§IV\-A](https://arxiv.org/html/2608.07905#S4.SS1.p1.1)\.
- \[34\]I\. Singh, V\. Blukis, A\. Mousavian, A\. Goyal, D\. Xu, J\. Tremblay, D\. Fox, J\. Thomason, and A\. Garg\(2023\)ProgPrompt: generating situated robot task plans using large language models\.In2023 IEEE International Conference on Robotics and Automation \(ICRA\),pp\. 11523–11530\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p2.1)\.
- \[35\]K\. P\. Singh, S\. Bhambri, B\. Kim, R\. Mottaghi, and J\. Choi\(2021\)Factorizing perception and policy for interactive instruction following\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 1888–1897\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p1.1)\.
- \[36\]C\. H\. Song, J\. Wu, C\. Washington, B\. M\. Sadler, W\. Chao, and Y\. Su\(2023\)Llm\-planner: few\-shot grounded planning for embodied agents with large language models\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 2998–3009\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p2.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1.2),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.7.1)\.
- \[37\]M\. Suganuma, T\. Okatani,et al\.\(2021\)Look wide and interpret twice: improving performance on interactive instruction\-following tasks\.In30th International Joint Conference on Artificial Intelligence, IJCAI 2021,pp\. 923–930\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1)\.
- \[38\]C\. Sun, S\. Huang, H\. Liu, J\. Gong, and D\. Pompili\(2025\)Retrieval\-augmented hierarchical in\-context reinforcement learning and hindsight modular reflections for task planning with llms\.In2025 IEEE International Conference on Robotics and Automation \(ICRA\),pp\. 1217–1224\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p2.1)\.
- \[39\]A\. Takmaz, A\. Delitzas, R\. W\. Sumner, F\. Engelmann, J\. Wald, and F\. Tombari\(2025\)Search3d: hierarchical open\-vocabulary 3d segmentation\.IEEE Robotics and Automation Letters\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p4.1),[§III\-C1](https://arxiv.org/html/2608.07905#S3.SS3.SSS1.p1.10)\.
- \[40\]M\. Vojnovic and S\. Yun\(2025\)What is the alignment objective of grpo?\.arXiv preprint arXiv:2502\.18548\.Cited by:[§III\-A3](https://arxiv.org/html/2608.07905#S3.SS1.SSS3.p1.1)\.
- \[41\]H\. Wang, S\. Zhang, S\. Wang, T\. Jiang, and Y\. Ge\(2025\)Double\-feedback: enhancing large language models reasoning in robotic tasks by knowledge graphs\.IEEE Robotics and Automation Letters\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[42\]Z\. Wang, S\. X\. Teo, J\. J\. Chew, and W\. Shi\(2025\)Instructrag: leveraging retrieval\-augmented generation on instruction graphs for llm\-based task planning\.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1413–1422\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[43\]J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[44\]T\. Xie, Z\. Gao, Q\. Ren, H\. Luo, Y\. Hong, B\. Dai, J\. Zhou, K\. Qiu, Z\. Wu, and C\. Luo\(2025\)Logic\-rl: unleashing llm reasoning with rule\-based reinforcement learning\.arXiv preprint arXiv:2502\.14768\.Cited by:[§III\-A3](https://arxiv.org/html/2608.07905#S3.SS1.SSS3.p2.1)\.
- \[45\]W\. Xu, M\. Wang, W\. Zhou, and H\. Li\(2024\)P\-rag: progressive retrieval augmented generation for planning on embodied everyday task\.InProceedings of the 32nd ACM International Conference on Multimedia,pp\. 6969–6978\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[46\]Y\. Yang\(2024\)DISCO: embodied navigation and interaction via differentiable scene semantics and dual\-level control\.InEuropean Conference on Computer Vision, ECCV 2024 \(29/09/2024\-04/10/2024, Milan\),Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1),[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[§III\-A1](https://arxiv.org/html/2608.07905#S3.SS1.SSS1.p2.7),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.12.1),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.9.1)\.
- \[47\]S\. Zhang, Z\. Xu, P\. Liu, X\. Yu, Y\. Li, Q\. Gao, Z\. Fei, Z\. Yin, Z\. Wu, Y\. Jiang,et al\.\(2024\)Vlabench: a large\-scale benchmark for language\-conditioned robotics manipulation with long\-horizon reasoning tasks\.arXiv preprint arXiv:2412\.18194\.Cited by:[§I](https://arxiv.org/html/2608.07905#S1.p1.1)\.
- \[48\]Z\. Zhang and A\. Zhang\(2023\)You only look at screens: multimodal chain\-of\-action agents\.arXiv preprint arXiv:2309\.11436\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.
- \[49\]Q\. Zhao, H\. Fu, C\. Sun, and G\. Konidaris\(2024\)EPO: hierarchical llm agents with environment preference optimization\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 6401–6415\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p1.1),[TABLE I](https://arxiv.org/html/2608.07905#S4.T1.4.14.1)\.
- \[50\]J\. Zhu, Y\. Liu, M\. Bao, K\. Zhang, Y\. Zhang, and Q\. Liu\(2025\)Self\-reflective planning with knowledge graphs: enhancing llm reasoning reliability for question answering\.arXiv preprint arXiv:2505\.19410\.Cited by:[§II](https://arxiv.org/html/2608.07905#S2.p2.1)\.

Similar Articles