A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning

arXiv cs.AI Papers

Summary

A fully differentiable neuro-soft-symbolic framework integrates visual perception and task planning in a single computational graph, achieving higher success rates on benchmarks like Blocksworld and PlanBench with reduced computation.

arXiv:2609.21221v1 Announce Type: new Abstract: Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and planning valid action sequences following logical rules. Conventional methods convert perception into discrete symbolic facts and then plan, discarding perceptual uncertainty and severing task-level feedback to perception. We introduce a generic, fully differentiable neuro-soft-symbolic framework that connects visual perception and task planning within a single computational graph. The framework maintains a continuous soft symbolic state, lifts domain rules into a differentiable soft-$T_P$ transition operator, and optimizes action logits over a short planning horizon. Gradients from the planning objective can also update the perception parameters, allowing task-relevant perceptual representations to be refined during planning. On Blocksworld, our method solves 40/40 LatPlan-40 tasks and 596/600 PlanBench-600 tasks, compared with 33/40 for LatPlan and 587/600 for the reasoning-model baseline, while requiring substantially less computation and time. In the perceptual-uncertainty ablation, our method improves the success rate from 59\% with frozen perception to 83\%. We further conduct task-and-motion simulations on Blocksworld scenes, providing an execution-level validation of the compatibility between decoded task plans and downstream robotic motion execution.
Original Article
View Cached Full Text

Cached at: 09/21/26, 09:19 AM

# A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning
Source: [https://arxiv.org/html/2609.21221](https://arxiv.org/html/2609.21221)
###### Abstract

Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and planning valid action sequences following logical rules\. Conventional methods convert perception into discrete symbolic facts and then plan, discarding perceptual uncertainty and severing task\-level feedback to perception\. We introduce a generic, fully differentiable neuro\-soft\-symbolic framework that connects visual perception and task planning within a single computational graph\. The framework maintains a continuous soft symbolic state, lifts domain rules into a differentiable soft\-TPT\_\{P\}transition operator, and optimizes action logits over a short planning horizon\. Gradients from the planning objective can also update the perception parameters, allowing task\-relevant perceptual representations to be refined during planning\. On Blocksworld, our method solves 40/40 LatPlan\-40 tasks and 596/600 PlanBench\-600 tasks, compared with 33/40 for LatPlan and 587/600 for the reasoning\-model baseline, while requiring substantially less computation and time\. In the perceptual\-uncertainty ablation, our method improves the success rate from 59% with frozen perception to 83%\. We further conduct task\-and\-motion simulations on Blocksworld scenes, providing an execution\-level validation of the compatibility between decoded task plans and downstream robotic motion execution\.

## IIntroduction

Visual task planning requires two complementary capabilities\. The system must infer a structured symbolic state from visual observations, where perceptual information may be uncertain, and must generate an action sequence that satisfies the logical rules and constraints of the task\. Neural networks provide effective perceptual representations but do not inherently guarantee valid symbolic reasoning\[[1](https://arxiv.org/html/2609.21221#bib.bib23)\]\. Conversely, symbolic planning methods provide explicit transition semantics and verifiable logical constraints\[[2](https://arxiv.org/html/2609.21221#bib.bib12),[3](https://arxiv.org/html/2609.21221#bib.bib13),[4](https://arxiv.org/html/2609.21221#bib.bib14)\]but typically assume that an accurate symbolic state is already available\. A suitable planning framework must therefore leverage perceptual information while enforcing the task’s logical rules and constraints\.

Existing approaches expose different limitations at the interface between perception and planning, see Figure[1](https://arxiv.org/html/2609.21221#S1.F1)\. \(1\) Explicit symbolic planning operates on deterministic facts and applies a symbolic solver to generate an action sequence, but perception is outside the planning computation and the method requires an explicit, accurate symbolic state\. \(2\) Language\- and vision\-language\-based methods \(VLM/LLM/LRM\) instead infer semantic descriptions, subgoals, or candidate actions from visual or language inputs through inductive reasoning\. This process can produce hallucinated or semantically inconsistent predictions and does not guarantee that the resulting actions satisfy all task constraints\. \(3\) The hard neuro\-symbolic method introduces perception into the pipeline but separates it from planning through a hard symbolic commitment: perceptual errors cannot be revised through the subsequent planning computation; the planning module cannot leverage the graded uncertainty present in the neural perceptual output; and planning objectives cannot shape the representation learned by the perception module because the symbolic handoff blocks task\-level gradient feedback\.

![Refer to caption](https://arxiv.org/html/2609.21221v1/figures/introduction_motivation.png)

Fig\. 1:Comparison of current planning methods\.We address this limitation with a fully differentiable neuro\-soft\-symbolic planning framework\. Instead of committing to a single symbolic state, the framework represents perceptual relations as continuous\-valued symbolic representations\. These representations retain graded perceptual information while providing a structured interface for logical reasoning\. Domain rules are lifted into an action\-conditioned soft\-TPT\_\{P\}transition operator, which propagates the soft state under candidate action distributions across a short\-horizon planning window\. The resulting planning objective provides gradients to the action variables and, in the perception\-connected setting, to the perception parameters\. In this way, the framework preserves perceptual uncertainty, enforces logical constraints during planning, and allows task\-level objectives to shape the learned perceptual representation\.

We evaluate the framework on visual Blocksworld planning tasks\. The method solves 40/40 \(100%\) LatPlan instances and 596/600 \(99\.33%\) PlanBench instances, compared with 33/40 \(82\.50%\) for LatPlan and 587/600 \(97\.83%\) for the PlanBench reasoning\-model baseline\. It requires an average of 9\.96 s per PlanBench instance, compared with 40\.43 s for the evaluated o1\-preview LRM\.

The contributions of this work are:

- •A generic, task\-agnostic, fully differentiable neuro\-soft\-symbolic framework that unifies visual perception and task planning within a single computational graph\. It maintains continuous symbolic interpretations throughout rule\-guided planning, allowing task objectives to influence both action optimization and perceptual representations\.
- •A logic\-guided soft\-TPT\_\{P\}transition operator that replaces hard symbolic commitment with continuous soft states, which form the basis for a fully differentiable perception–planning computation, while lifting domain rules and action schemas into differentiable temporal propagation under candidate action distributions and preserving perceptual uncertainty\.

## IIRelated Work

Classical symbolic planning\.Classical planning originated with explicit symbolic preconditions and effects\[[2](https://arxiv.org/html/2609.21221#bib.bib12)\]and was later standardized through languages such as PDDL\[[3](https://arxiv.org/html/2609.21221#bib.bib13)\]\. Modern planners provide efficient search, explicit transition semantics, and verifiable plans\[[4](https://arxiv.org/html/2609.21221#bib.bib14)\]; integrated task\-and\-motion planning extends this formulation with geometric feasibility and motion constraints\[[5](https://arxiv.org/html/2609.21221#bib.bib6)\]\. These methods assume the current task state is already available as reliable symbolic facts\. Consequently, perception uncertainty is either handled by a separate front end or omitted from the planning goal\.

Neuro\-symbolic reasoning\.Logic Tensor Networks represent logical theories with differentiable tensor semantics\[[6](https://arxiv.org/html/2609.21221#bib.bib16)\], while DeepProbLog combines neural predicates with probabilistic logic programs\[[7](https://arxiv.org/html/2609.21221#bib.bib15)\]\. Neural Logic Machines further demonstrate how relational structure and logical operations can be learned with neural modules\[[8](https://arxiv.org/html/2609.21221#bib.bib17)\]\. Recent soft answer\-set and differentiable propagation methods provide additional foundations for retaining logical structure in continuous computation\[[9](https://arxiv.org/html/2609.21221#bib.bib4),[10](https://arxiv.org/html/2609.21221#bib.bib5)\]\. These works inspire our use of continuous symbolic interpretations and differentiable rule lifts, but they primarily address static deduction or state propagation rather than temporal planning with execution\.

LatPlan learns a propositional representation from images and performs classical planning in the learned latent space\[[11](https://arxiv.org/html/2609.21221#bib.bib1)\]\. Model\-based visual control methods learn latent dynamics and optimize imagined rollouts from pixels\[[12](https://arxiv.org/html/2609.21221#bib.bib18),[13](https://arxiv.org/html/2609.21221#bib.bib19)\], while other approaches learn latent plans from demonstrations or play\[[14](https://arxiv.org/html/2609.21221#bib.bib20)\]\. These methods demonstrate the value of planning over representations, but their latent states do not generally expose verifiable rule\-level transition semantics\.

Language\- and vision\-language\-based planning\.Language models provide a semantic mechanism for task decomposition and action generation\. LLM\+P translates natural\-language problems into PDDL and delegates search to a classical planner\[[15](https://arxiv.org/html/2609.21221#bib.bib9)\], while SayCan grounds language\-model suggestions in learned robotic affordances\[[16](https://arxiv.org/html/2609.21221#bib.bib11)\]\. Other language\-conditioned systems use code generation, embodied feedback, or hierarchical task planning for robotic tasks\[[17](https://arxiv.org/html/2609.21221#bib.bib21),[18](https://arxiv.org/html/2609.21221#bib.bib22),[19](https://arxiv.org/html/2609.21221#bib.bib10)\]\. PlanBench evaluates language models on planning and reasoning about change\[[20](https://arxiv.org/html/2609.21221#bib.bib2)\], while later evaluation shows that additional reasoning\-model computation does not by itself guarantee executable and verifiable plans\[[21](https://arxiv.org/html/2609.21221#bib.bib3)\]\. These methods rely on language\-mediated semantic prediction, translation, parsing, or external verification, and they do not provide a differentiable route from a task objective through rule\-constrained temporal state updates to a visual representation\. Our method does not require an LLM/LRM, or a natural\-language intermediate representation; task semantics are specified directly through domain rules\.

Task and motion planning\.Classical task\-and\-motion planning couples symbolic task decisions with geometric feasibility, collision avoidance, and robot motion generation\[[5](https://arxiv.org/html/2609.21221#bib.bib6)\]\. Differentiable task\-and\-motion and trajectory\-optimization methods introduce gradients into parts of this process\[[22](https://arxiv.org/html/2609.21221#bib.bib7)\]\. Our current framework addresses the high\-level task\-planning component, with actions represented as relational operators such as moving an entity to a support target\. This formulation provides a natural basis for future integration with differentiable motion optimizers\.

Overall, existing neuro\-symbolic methods primarily study static inference or rule evaluation, whereas planning methods generally operate on committed symbolic states or use non\-differentiable interfaces between perception, semantics, and search\. Our work addresses this gap with a fully differentiable neuro\-soft\-symbolic framework in which continuous symbolic representations remain connected to logical state transitions, action optimization, and the task objective within one computational graph\.

![Refer to caption](https://arxiv.org/html/2609.21221v1/figures/neuro_soft_symbolic_framework_updated_gradients.png)Fig\. 2:Overview of the proposed fully differentiable neuro\-soft\-symbolic planning framework\.
## IIIMethod

### III\-AFramework Overview

Figure[2](https://arxiv.org/html/2609.21221#S2.F2)gives the complete pipeline\. The perception module maps current and goal observations to continuous support interpretations\. Domain rules are then lifted into an action\-conditioned soft\-TPT\_\{P\}transition operator, which propagates the interpretation under soft action distributions across a short planning window\. The resulting terminal state is evaluated by a planning objective, whose gradients reach the action logits and, in the perception\-connected setting, the perception parameters\. After optimization, the action distribution is decoded into a legal planning action sequence; execution produces the next observation and triggers the next short\-horizon replanning window\.

### III\-BProblem Formulation

We consider a deterministic symbolic task represented by a domain transition relation and a visual observation pair\(I0,Ig\)\(I\_\{0\},I\_\{g\}\)\. LetNNentities interact withKKbase slots\. For Blocksworld,K=1K=1and the base slot is an unlimited\-capacity table\. A support state is

𝐅∈\[0,1\]N×\(N\+K\),∑y𝐅x,y=1,\\mathbf\{F\}\\in\[0,1\]^\{N\\times\(N\+K\)\},\\qquad\\sum\_\{y\}\\mathbf\{F\}\_\{x,y\}=1,\(1\)where𝐅x,y\\mathbf\{F\}\_\{x,y\}denotes the belief that entityxxdirectly rests on targetyy\.

At each planning steptt, the optimizer maintains action logitsat∈ℝN×\(N\+K\+1\)a\_\{t\}\\in\\mathbb\{R\}^\{N\\times\(N\+K\+1\)\}\. The extra column represents STAY\. The corresponding action distribution isMt=softmax⁡\(at/τ\)M\_\{t\}=\\operatorname\{softmax\}\(a\_\{t\}/\\tau\)\. Given an initial support state and a goal support state, the short\-term planning problem is

mina0:T−1,θℒ\(𝐅T,𝐅g\)s\.t\.𝐅T=Φθ\(T\)\(𝐅0,a0:T−1\)\.\\small\\min\_\{a\_\{0:T\-1\},\\,\\theta\}\\mathcal\{L\}\(\\mathbf\{F\}\_\{T\},\\mathbf\{F\}\_\{g\}\)\\hskip 9\.24994pt\\text\{s\.t\.\}\\hskip 9\.24994pt\\mathbf\{F\}\_\{T\}=\\Phi\_\{\\theta\}^\{\(T\)\}\(\\mathbf\{F\}\_\{0\},a\_\{0:T\-1\}\)\.\(2\)whereθ\\thetadenotes perception parameters in the perception\-connected variant\.

### III\-CVisual\-to\-Soft\-Symbolic Perception

The perception network uses a CNN backbone followed by region\-of\-interest pooling for each entity\. Letrxr\_\{x\}be the embedding of entityxxand letryr\_\{y\}be the embedding of a candidate support target, including a learned base\-slot embedding\. A directional bilinear head assigns

s\(x,y\)=rx⊤Wry\+b,𝐅x,:=softmaxy\(s\(x,y\)\)\.s\(x,y\)=r\_\{x\}^\{\\top\}Wr\_\{y\}\+b,\\qquad\\mathbf\{F\}\_\{x,:\}=\\operatorname\{softmax\}\_\{y\}\(s\(x,y\)\)\.\(3\)The matrixWWis not constrained to be symmetric, because “xxonyy” and “yyonxx” are different relations\. Self\-support entries and padded entities are masked before the softmax\. This produces a continuous symbolic interpretation of the image without an intermediate argmax\.

TABLE I:Correspondence between Blocksworld logic rules and the differentiable soft\-TPT\_\{P\}lift used in our implementation\.
### III\-DDifferentiable Soft\-TPT\_\{P\}Transition

The soft\-TPT\_\{P\}transition operator provides the domain\-rule interface for the shared differentiable planning machinery\. We describe its Blocksworld instantiation\. Each block has one support, a block target supports at most one block, the table has unlimited capacity, and only a clear block may move\. STAY is represented explicitly in the action distribution\.

Let𝒴\\mathcal\{Y\}denote the physical support targets, excluding STAY\. The soft clear value of blockxxis

ct​\(x\)=∏z≠x\(1−𝐅t,z,x\)\.c\_\{t\}\(x\)=\\prod\_\{z\\neq x\}\\left\(1\-\\mathbf\{F\}\_\{t,z,x\}\\right\)\.\(4\)
For a candidate movex→yx\\\!\\rightarrow\\\!y, we first compute the raw legality\-gated move mass

g~t​\(x,y\)=Mt​\(x,y\)​ct​\(x\)​c¯t​\(y\)​et​\(x,y\),\\tilde\{g\}\_\{t\}\(x,y\)=M\_\{t\}\(x,y\)\\,c\_\{t\}\(x\)\\,\\bar\{c\}\_\{t\}\(y\)\\,e\_\{t\}\(x,y\),\(5\)wherec¯t​\(y\)=ct​\(y\)\\bar\{c\}\_\{t\}\(y\)=c\_\{t\}\(y\)for block targets andc¯t​\(y\)=1\\bar\{c\}\_\{t\}\(y\)=1for the table\. The termet​\(x,y\)e\_\{t\}\(x,y\)masks self\-support and other domain\-specific invalid targets\. Because at most one block move at each step, we compute

ρt​\(x\)=∑y∈𝒴g~t​\(x,y\),qt​\(x\)=∏x′≠x\(1−ρt​\(x′\)\),\\rho\_\{t\}\(x\)=\\sum\_\{y\\in\\mathcal\{Y\}\}\\tilde\{g\}\_\{t\}\(x,y\),\\qquad q\_\{t\}\(x\)=\\prod\_\{x^\{\\prime\}\\neq x\}\\left\(1\-\\rho\_\{t\}\(x^\{\\prime\}\)\\right\),\(6\)and define the committed move mass as

gt​\(x,y\)=g~t​\(x,y\)​qt​\(x\),mt​\(x\)=∑y∈𝒴gt​\(x,y\)\.g\_\{t\}\(x,y\)=\\tilde\{g\}\_\{t\}\(x,y\)\\,q\_\{t\}\(x\),\\qquad m\_\{t\}\(x\)=\\sum\_\{y\\in\\mathcal\{Y\}\}g\_\{t\}\(x,y\)\.\(7\)This provides differentiable competition among move proposals\.

The support state is then updated by

𝐅t\+1,x,y=\(1−mt​\(x\)\)​𝐅t,x,y\+gt​\(x,y\),y∈𝒴\.\\mathbf\{F\}\_\{t\+1,x,y\}=\\left\(1\-m\_\{t\}\(x\)\\right\)\\mathbf\{F\}\_\{t,x,y\}\+g\_\{t\}\(x,y\),\\qquad y\\in\\mathcal\{Y\}\.\(8\)The explicit STAY probability and all move mass rejected by the logical gates remain in the residual mass1−mt​\(x\)1\-m\_\{t\}\(x\), thereby preserving the current support distribution\. Thus,𝐅t\+1=Φ⁡\(𝐅t,Mt\)\\mathbf\{F\}\_\{t\+1\}=\\Phi\(\\mathbf\{F\}\_\{t\},M\_\{t\}\)is differentiable with respect to both the soft state and the action logits\.

For block targets, we optionally enforce target\-capacity consistency with

𝒯Ptarg​\(𝐅\)z,y=𝐅z,y​∏z′≠z\(1−𝐅z′,y\),\\mathcal\{T\}\_\{P\}^\{\\mathrm\{targ\}\}\(\\mathbf\{F\}\)\_\{z,y\}=\\mathbf\{F\}\_\{z,y\}\\prod\_\{z^\{\\prime\}\\neq z\}\\left\(1\-\\mathbf\{F\}\_\{z^\{\\prime\},y\}\\right\),\(9\)followed by row normalization and fixed\-point iteration\.

The equations above instantiate the proposed framework for Blocksworld and provide a concrete example of how domain rules are implemented\. The overall framework is generic: across domains, the action\-logit optimization, soft temporal rollout, planning objective, gradient propagation, decoding procedure, and short\-horizon replanning remain unchanged\. Only the rule\-specific transition predicates and action schemas used to instantiateΦ\\Phiare replaced\. Thus, the same framework can be applied to different planning domains without redesigning the optimization procedure\.

### III\-EDifferentiable Perception–Planning Optimization

For each short\-horizon window, the action logits and the perception parametersθ\\thetaare updated through the same differentiable planning objective\. The visual states are recomputed from the raw images at every gradient step, ensuring that each parameter update is used by the next rollout\.

In addition to the goal, movement, and cycle terms, we penalize rolled\-out states that have not yet reached a fixed point of the target\-exclusivity operatorTPT\_\{P\}\.

ℒf​p=∑t=1T‖𝐅t−TP​\(𝐅t\)‖2\.\\displaystyle\\mathcal\{L\}\_\{fp\}=\\sum\_\{t=1\}^\{T\}\\bigl\\\|\\mathbf\{F\}\_\{t\}\-T\_\{P\}\(\\mathbf\{F\}\_\{t\}\)\\bigr\\\|^\{2\}\.\(10\)ℒf​p\\mathcal\{L\}\_\{fp\}is evaluated on each rolled\-out state*before*the fixed\-point refinement is applied, so it directly rewards a transition that is already close to self\-consistent, rather than relying on refinement alone\. The full planning objective, shared by every variant, is

ℒ=\\displaystyle\\mathcal\{L\}=\{\}ℒg​o​a​l​\(𝐅T,𝐅g\)\+λm​ℒm​o​v​e\+λc​ℒc​y​c​l​e\+λf​p​ℒf​p\.\\displaystyle\\mathcal\{L\}\_\{goal\}\(\\mathbf\{F\}\_\{T\},\\mathbf\{F\}\_\{g\}\)\+\\lambda\_\{m\}\\mathcal\{L\}\_\{move\}\+\\lambda\_\{c\}\\mathcal\{L\}\_\{cycle\}\+\\lambda\_\{fp\}\\mathcal\{L\}\_\{fp\}\.\(11\)The goal term is a cross\-entropy for partial\-goal tasks\. The movement term favors parsimonious plans, the cycle term penalizes direct two\-cycles, and the fixed\-point term encourages a self\-consistent rollout\. For the perception\-connected variant, task\-level gradients also update the perception parametersθ\\theta\. To limit deviation from the initial perceptual predictions, we use KL anchors:

ℒP​C=ℒ\+λa​n​c,0​D0\+λa​n​c,g​Dg,\\mathcal\{L\}\_\{PC\}=\\mathcal\{L\}\+\\lambda\_\{anc,0\}D\_\{0\}\+\\lambda\_\{anc,g\}D\_\{g\},\(12\)whereD0=DK​L\(𝐅0∥𝐅0\(0\)\)D\_\{0\}=D\_\{KL\}\(\\mathbf\{F\}\_\{0\}\\\|\\mathbf\{F\}\_\{0\}^\{\(0\)\}\)andDg=DK​L\(𝐅g∥𝐅g\(0\)\)D\_\{g\}=D\_\{KL\}\(\\mathbf\{F\}\_\{g\}\\\|\\mathbf\{F\}\_\{g\}^\{\(0\)\}\)\.𝐅0\(0\)\\mathbf\{F\}\_\{0\}^\{\(0\)\}and𝐅g\(0\)\\mathbf\{F\}\_\{g\}^\{\(0\)\}are the reference soft states, while𝐅0=Pθ​\(I0\)\\mathbf\{F\}\_\{0\}=P\_\{\\theta\}\(I\_\{0\}\)and𝐅g=Pθ​\(Ig\)\\mathbf\{F\}\_\{g\}=P\_\{\\theta\}\(I\_\{g\}\)remain functions ofθ\\theta\. Thus, the anchors regularize perception updates without breaking the differentiable perception–planning pathway\.

TABLE II:Main planning results\. All reported Ours results on Blocksworld use 3 independent planning seeds\.Redshows the improvement of our method over the corresponding baseline\.Our Cost is estimated from measured GPU time \(9\.96 s/instance, measured while running our 3 evaluation seeds concurrently on a single A100\) assuming on\-demand cloud A100 pricing \($2\.745/GPU\-hour, AWS EC2 p4d\.24xlarge, and $3\.673/GPU\-hour, Google Cloud a2\-highgpu\-1g; both accessed September 2026\)\.

The resulting gradients update the action logits, and, for the perception\-connected variant, the perception parameters, inside the same short\-term window\. The action\-logit learning rate is0\.10\.1and the perception learning rate is10−410^\{\-4\}\. Independently of theℒf​p\\mathcal\{L\}\_\{fp\}term above, we also iteratively refine each rolled\-out state toward the same fixed point before the next transition step, which improves optimization stability for longer windows\.

### III\-FShort\-Horizon Planning and Replanning

The optimizer plans over a short window, decodes the planning action sequence, and executes it against the current real symbolic state\. The next window starts from the resulting state, following an iterative short\-horizon planning procedure also used by learned latent\-world\-model planners\[[23](https://arxiv.org/html/2609.21221#bib.bib8)\]\. Within each short\-term planning window, the forward computation, planning objective, and gradient updates form a differentiable path from the visual inputs through the soft\-TPT\_\{P\}rollout to the perception parameters and action logits\. The action sequence is decoded, legally executed under the task rules, and used to initialize the next planning window\.

TABLE III:PlanBench Blocksworld error analysis\. The PlanBench baseline results are reported overall; our method is additionally broken down by number of blocks\. Optimal is reported as a percentage of valid plans\.TABLE IV:Ablation Study for Differentiable Perception\-Planning

## IVExperiments

### IV\-AExperiments Setup

Dataset\.We evaluate on image\-based Blocksworld and on Logistics and Sokoban tasks\. Blocksworld contains the 40 image instances associated with LatPlan and 600 official PlanBench instances\. The 600\-instance set contains 100 3\-block, 445 4\-block and 55 5\-block problems\. Experiments were run on a single NVIDIA A100\.

Baseline\.We compare with LatPlan on the 40 image\-based Blocksworld instances\. On the 600\-instance PlanBench Blocksworld benchmark, we include the reported zero\-shot results of Claude 3\.5 Sonnet, LLaMA 3\.1 405B, o1\-mini, and o1\-preview\[[21](https://arxiv.org/html/2609.21221#bib.bib3)\]\. These results correspond to the best reported performance for each listed model under the evaluation setting used in PlanBench\. For the Logistics and Sokoban domains, we compare with the o1\-preview results reported in the same benchmark evaluation\.

Evaluation Metrics\.Plans are decoded into action sequences and validated by executing them from initial state\. We report validity, optimality, inexecutable plans, non\-goal\-reaching plans, average solve time, and estimated computational cost\. A plan is valid if every action satisfies its preconditions and the resulting final state satisfies the goal\. This plan\-validation evaluates the decoded actions under the task’s transition rules, thereby exposing both execution failures and failures to reach the goal\.

### IV\-BMain Results

Table[II](https://arxiv.org/html/2609.21221#S3.T2)reports the main results, with optimality reported as the percentage of valid plans\. All results on the Blocksworld LatPlan\-40 and PlanBench\-600 dataset use three independent planning seeds\. On LatPlan\-40, our neuro\-soft\-symbolic perception–planning method achieves 40/40 valid solutions, including 31/40 BFS\-optimal solutions, compared with 33/40 for LatPlan\. On PlanBench\-600, our method achieves 596/600 valid solutions \(99\.33%\), surpassing the 587/600 \(97\.83%\) result of o1\-preview\. All four unsuccessful PlanBench instances are non\-goal\-reaching; none contains an inexecutable action sequence\.

Our method also provides better efficiency on the 600\-instance Blocksworld benchmark\. The average solve time is reduced from 40\.43 s to 9\.96 s per instance, while the estimated cost per 100 instances decreases from $42\.12 to $0\.89\. This lower latency and computational cost are important for practical planning systems\.

### IV\-CPlan Validation and Distribution Analysis

Table[III](https://arxiv.org/html/2609.21221#S3.T3)reports the PlanBench baseline error distribution together with our block\-count breakdown\. The baseline row follows the published PlanBench Blocksworld result, while the 3\-block, 4\-block, and 5\-block rows show our method’s behavior\. Our four failures occur only in the most difficult 5\-block tier, as expected for a longer and more coupled planning problem\. More importantly, our zero inexecutable count shows that every submitted action satisfied the modeled Blocksworld preconditions under plan validation, which is important for logic\-constrained and safety\-critical tasks\.

The distribution of task difficulty provides a complementary view of the PlanBench\-600 results\. The optimal step counts range from 1 to 8 and are concentrated between 3 and 6 steps \(Figure[3](https://arxiv.org/html/2609.21221#S4.F3)\)\. Among valid\-but\-suboptimal instances, the mean excess is 2\.47 steps, with a range of 1–7 steps\. Optimal step counts on LatPlan\-40 range from 0 to 4 \(mean 2\.12\); our method matches the optimum on 31/40 instances, with a mean excess of 2\.67 steps \(range 1–4\) among the remaining 9\.

![Refer to caption](https://arxiv.org/html/2609.21221v1/figures/planbench_step_count_distribution.png)Fig\. 3:Distribution of PlanBench\-600 instances by optimal step count\.
### IV\-DAblation Study for Differentiable Perception\-Planning

To evaluate our method’s robustness, we conduct an ablation study on the 600\-instance Blocksworld using three configurations: ground\-truth symbolic state, frozen perception, and fully differentiable perception–planning\. Frozen perception computes the soft symbolic state with fixed perception parameters during planning, whereas fully differentiable perception–planning allows planning gradients to update these parameters\. As reported in Table[IV](https://arxiv.org/html/2609.21221#S3.T4), on clear Blocksworld images, all three settings achieve596/600596/600valid plans\. This demonstrates that our fully differentiable method—which maintains a continuous soft symbolic representation across perception and planning—reaches the performance of ground\-truth symbolic inputs without requiring external solvers or reasoning models\. Because Blocksworld scenes are relatively simple, the perception model nears saturation on clear images, making the benefits of updating perception parameters difficult to observe under clean inputs\. To evaluate the perception–planning connection under perceptual uncertainty, we construct a controlled image\-level degradation experiment\. We select a fixed 100\-instance subset from PlanBench\-600, preserving the original block\-count proportions \(17 from 3\-block, 74 from 4\-block, and 9 from 5\-block instances\)\. We apply a 6\-pixel radius Gaussian blur to the initial and goal images before passing them to the perception backbone\. All other settings \(soft\-TPT\_\{P\}transition, window size, perception checkpoint, seeds, and planning hyperparameters\) remain unchanged, and the perception model is neither retrained nor fine\-tuned on the degraded images\.

On the clear 100\-instance subset, both frozen perception and fully differentiable perception–planning achieve100/100100/100valid plans, confirming no difficulty bias in the subset\. However, under image degradation, performance clearly diverges: frozen perception achieves59/10059/100valid plans, while our fully differentiable perception–planning method reaches83/10083/100\(improved 24%\)\. This difference reflects how the two methods handle perceptual errors\. Frozen perception treats the perceptual output from degraded images as a fixed planning input; if this output contains incorrect relational judgments, subsequent planning cannot revise them\. In contrast, our method allows gradients from the planning objective to propagate through the soft symbolic state to the perception parameters, thereby correcting some perceptual errors during optimization\. These results indicate that a differentiable perception–planning connection provides a positive contribution when perception is uncertain\.

### IV\-EAblation Study for Short\-Horizon Window Size

The short\-horizon window size used for the Blocksworld results in Table[II](https://arxiv.org/html/2609.21221#S3.T2)isT=3T=3, selected based on this ablation\. We evaluate different values ofTTon 150 independently rendered Blocksworld instances prepared for this experiment, while keeping all other hyperparameters fixed\. Because the settings solve different numbers of instances, the optimal rate is computed over all 150 instances rather than only the valid plans\. As shown in Table[V](https://arxiv.org/html/2609.21221#S4.T5),T=3T=3provides the best overall trade\-off, achieving the highest validity \(100\.0%\), the highest optimal rate \(76\.7%\), the lowest average time \(7\.77 s\), and no failures\. The shorter window \(T=1T=1\) more often terminates before reaching the goal, whereas the longer window \(T=5T=5\) increases the cost of each optimization round and reduces validity\. No setting produces an inexecutable action sequence, consistent with the zero\-inexecutable result on the full 600\-instance evaluation \(Table[III](https://arxiv.org/html/2609.21221#S3.T3)\)\.

TABLE V:Ablation study on short\-horizon window\-size sensitivity\.![Refer to caption](https://arxiv.org/html/2609.21221v1/figures/instance50_qualitative_trace.png)Fig\. 4:Case study of the proposed perception\-to\-planning procedure on a 4\-block PlanBench instance\.
### IV\-FCase Study on Blocksworld Instance

Figure[4](https://arxiv.org/html/2609.21221#S4.F4)shows the complete computation for a specific four\-block Blocksworld problem from PlanBench\. The initial state isaaon the table,bboncc,ccon the table, andddon the table\. The goal state requiresaaonbb,bboncc,ccondd, andddon the table\. As shown byF0F\_\{0\}, the perception module processes the initial image into soft probabilities in\[0,1\]\[0,1\]over continuous support relations\. In the first round, corresponding tot=0,1,2t=0,1,2, no candidate action att=2t=2reaches the decoder confidence threshold of0\.30\.3for the Blocksworld task\. Therefore, only the actions att=0t=0andt=1t=1are submitted, and the next planning round starts from the stateF2F\_\{2\}obtained in the first round\.

This example highlights how the differentiable rule constraints determine action ordering\. Although the relation “bboncc” already satisfies part of the goal,cccannot be moved toddwhilebbremains oncc\. From Eq\.[4](https://arxiv.org/html/2609.21221#S3.E4),Fb,c≈1F\_\{b,c\}\\approx 1drivesct​\(c\)c\_\{t\}\(c\)close to zero\. Consequently, the raw move massg~t​\(c,d\)\\tilde\{g\}\_\{t\}\(c,d\)in Eq\.[5](https://arxiv.org/html/2609.21221#S3.E5)remains negligible regardless of how strongly the action logits favormove⁡\(c→d\)\\mathrm\{move\}\(c\\\!\\rightarrow\\\!d\)\. The optimizer must therefore first assign mass tomove⁡\(b→table\)\\mathrm\{move\}\(b\\\!\\rightarrow\\\!\\mathrm\{table\}\), which increases the subsequent clear value ofccand makesmove⁡\(c→d\)\\mathrm\{move\}\(c\\\!\\rightarrow\\\!d\)feasible\.

Across the two planning windows, the planner produces the four\-action:move⁡\(b→table\)\\mathrm\{move\}\(b\\\!\\rightarrow\\\!\\mathrm\{table\}\),move⁡\(c→d\)\\mathrm\{move\}\(c\\\!\\rightarrow\\\!d\),move⁡\(b→c\)\\mathrm\{move\}\(b\\\!\\rightarrow\\\!c\), andmove⁡\(a→b\)\\mathrm\{move\}\(a\\\!\\rightarrow\\\!b\)\. The first window commits the first two actions and reachesF2F\_\{2\}\. The planner then replans fromF2F\_\{2\}and commits the remaining two actions in the second window\. Plan validation replays the complete decoded sequence from the initial state, confirming that all four actions are executable and the final state satisfies the goal\.

![Refer to caption](https://arxiv.org/html/2609.21221v1/figures/integrated_task_motion_simulation.png)Fig\. 5:Integrated task\-and\-motion simulation on four\-block and five\-block Blocksworld instances\.
### IV\-GTask\-and\-Motion Simulation

To evaluate whether the decoded task actions can be grounded in downstream motion execution, we integrated the action sequences from PlanBench Blocksworld instances into a PyBullet simulation of a Franka Emika Panda arm\. Each relational actionmove​\(x,y\)\\texttt\{move\}\(x,y\)is realized as a pick\-and\-place primitive, with inverse kinematics converting Cartesian waypoints for approach, grasp, transport, and release into the corresponding arm joint configurations\. This experiment provides an execution\-level validation of the task\-to\-motion connection between high\-level task planning and downstream motion execution\. The simulation results show that the relational action sequences generated by our method can be translated into concrete arm motions and successfully execute the corresponding block manipulations\. Figure[5](https://arxiv.org/html/2609.21221#S4.F5)shows key simulation frames from two instances, including the pick\-and\-place operation for each action step\.

TABLE VI:Cross\-domain comparison on Logistics and Sokoban\.
### IV\-HCross\-Domain Evaluation

We also apply the soft\-symbolic planning machinery to two additional domains of PlanBench\. Logistics tests transport and fleet\-routing structure, while Sokoban tests irreversible manipulation actions and deadlock\-sensitive planning in constrained spaces\. On the same 200\-instance Logistics set, our method solves 196/200 instances \(98\.00%\), compared with 188/200 \(94\.00%\) for o1\-preview\. In Sokoban, our method solves 33/55 instances \(60\.00%\), compared with 7/55 \(12\.73%\) for o1\-preview\. In both domains, our method produces no inexecutable action sequence; the remaining failures are non\-goal\-reaching\. The corresponding error counts and average times are reported in Table[VI](https://arxiv.org/html/2609.21221#S4.T6)\. These domains use domain\-specific transition rules while retaining the same differentiable action\-optimization principle\.

## VConclusion and Future Work

We presented a generic, fully differentiable neuro\-soft\-symbolic framework for perceptual task planning\. It maintains continuous symbolic interpretations, preserves perceptual uncertainty, and connects perception to planning through a differentiable soft\-TPT\_\{P\}rollout within each optimization window\. On the Blocksworld benchmark, our method achieves higher planning validity than the baselines, while substantially reducing solve time and estimated cost relative to the LLM/LRM baseline\. These efficiency gains provide a practical advantage for embodied planning systems operating under latency and resource constraints\. The perceptual\-uncertainty ablation further shows that the differentiable perception–planning connection improves robustness to perceptual uncertainty\.

The current evaluation focuses on compact relational domains with a manageable number of entities and action types\. In future work, we will extend the framework to more demanding settings with many interacting entities, long chains of dependencies, richer action schemas, partial observability, and geometric constraints\. These extensions require scalable relational representations, structured decomposition, and efficient rule propagation to broaden the framework’s applicability to embodied AI and robotic planning while preserving the differentiable perception–planning connection\.

## References

- \[1\]S\. Levine, C\. Finn, T\. Darrell, and P\. Abbeel\(2016\)End\-to\-end training of deep visuomotor policies\.InJournal of Machine Learning Research,Vol\.17,pp\. 1–40\.Cited by:[§I](https://arxiv.org/html/2609.21221#S1.p1.1)\.
- \[2\]R\. E\. Fikes and N\. J\. Nilsson\(1971\)STRIPS: a new approach to the application of theorem proving to problem solving\.InProceedings of the 2nd International Joint Conference on Artificial Intelligence,pp\. 608–620\.Cited by:[§I](https://arxiv.org/html/2609.21221#S1.p1.1),[§II](https://arxiv.org/html/2609.21221#S2.p1.1)\.
- \[3\]D\. McDermott, M\. Ghallab, A\. E\. Howe, C\. A\. Knoblock, A\. Ram, M\. Veloso, D\. S\. Weld, and D\. E\. Wilkins\(1998\)PDDL—the planning domain definition language\.Technical reportTechnical ReportCVC TR\-98\-003/DCS TR\-1165,Yale Center for Computational Vision and Control\.Cited by:[§I](https://arxiv.org/html/2609.21221#S1.p1.1),[§II](https://arxiv.org/html/2609.21221#S2.p1.1)\.
- \[4\]M\. Helmert\(2006\)The fast downward planning system\.Journal of Artificial Intelligence Research26,pp\. 191–246\.Cited by:[§I](https://arxiv.org/html/2609.21221#S1.p1.1),[§II](https://arxiv.org/html/2609.21221#S2.p1.1)\.
- \[5\]C\. R\. Garrett, R\. Chitnis, R\. Holladay, B\. Kim, T\. Silver, L\. P\. Kaelbling, and T\. Lozano\-Pérez\(2021\)Integrated task and motion planning\.Annual Review of Control, Robotics, and Autonomous Systems4\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p1.1),[§II](https://arxiv.org/html/2609.21221#S2.p5.1)\.
- \[6\]L\. Serafini and A\. d’Avila Garcez\(2016\)Logic tensor networks: theory, practice, and application\.InProceedings of the 2016 International Conference on Computational Logic,Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p2.1)\.
- \[7\]R\. Manhaeve, S\. Dumančić, A\. Kimmig, T\. Demeester, and L\. D\. Raedt\(2018\)DeepProbLog: neural probabilistic logic programming\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p2.1)\.
- \[8\]H\. Dong, J\. Mao, T\. Wu, Y\. Yao, B\. C\. Ellis, B\. Jalaian, S\. Seo, Y\. Zhang, and J\. B\. Tenenbaum\(2019\)Neural logic machines\.InInternational Conference on Learning Representations,Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p2.1)\.
- \[9\]A\. Takemura and K\. Inoue\(2024\)Differentiable logic programming for distant supervision\.InEuropean Conference on Artificial Intelligence,Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p2.1)\.
- \[10\]T\. Eiter, K\. Inoue, and S\. Moriyama\(2026\)Neural decision–propagation for answer set programming\.arXiv preprint arXiv:2605\.01797\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p2.1)\.
- \[11\]M\. Asai, H\. Kajino, A\. Fukunaga, and C\. Muise\(2022\)Classical planning in deep latent space\.Journal of Artificial Intelligence Research74,pp\. 1599–1686\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p3.1),[TABLE II](https://arxiv.org/html/2609.21221#S3.T2.3.2.2.1)\.
- \[12\]F\. Ebert, C\. Finn, A\. X\. Lee, and S\. Levine\(2018\)Visual foresight: model\-based deep reinforcement learning for vision\-based robotic control\.InarXiv preprint arXiv:1812\.00568,Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p3.1)\.
- \[13\]D\. Hafner, T\. Lillicrap, I\. Fischer, R\. Villegas, D\. Ha, H\. Lee, and M\. Norouzi\(2020\)Dream to control: learning behaviors by latent imagination\.InInternational Conference on Learning Representations,Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p3.1)\.
- \[14\]C\. Lynch, M\. Khansari, T\. Xiao, V\. Kumar, J\. Tompson, S\. Levine, and P\. Sermanet\(2020\)Learning latent plans from play\.InConference on Robot Learning,Proceedings of Machine Learning Research, Vol\.100,pp\. 1113–1132\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p3.1)\.
- \[15\]B\. Liu, Y\. Jiang, X\. Zhang, Q\. Liu, S\. Zhang, J\. Biswas, and P\. Stone\(2023\)LLM\+P: empowering large language models with optimal planning proficiency\.arXiv preprint arXiv:2304\.11477\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p4.1)\.
- \[16\]M\. Ahn, A\. Brohan, N\. Brown, Y\. Chebotar, C\. Cortes, B\. David, C\. Finn, K\. Fu, K\. Gopal, K\. Hausman, A\. Herzog, D\. Hsu, J\. Ibarz, B\. Ichter, A\. Irpan, E\. Jang, R\. Ruano, D\. Sadigh, P\. Sermanet, N\. Sievers, C\. Tan, K\. Tien, V\. Vanhoucke, F\. Xia, T\. Xiao, P\. Xu, S\. Yan, and A\. Zeng\(2022\)Do as i can, not as i say: grounding language in robotic affordances\.arXiv preprint arXiv:2204\.01691\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p4.1)\.
- \[17\]J\. Liang, W\. Huang, F\. Xia, P\. Xu, K\. Hausman, M\. I\. Jordan, and S\. Levine\(2023\)Code as policies: language model programs for embodied control\.InProceedings of the 2023 IEEE International Conference on Robotics and Automation,pp\. 9493–9500\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p4.1)\.
- \[18\]W\. Huang, F\. Xia, T\. Xiao, H\. Chan, J\. Liang, P\. Stone, B\. Ichter, D\. Driess, A\. Brohan, Y\. Lu, and S\. Levine\(2023\)Inner monologue: embodied reasoning through planning with language models\.InConference on Robot Learning,pp\. 1769–1782\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p4.1)\.
- \[19\]M\. Kwon, Y\. Kim, and Y\. J\. Kim\(2025\)Fast and accurate task planning using neuro\-symbolic language models and multi\-level goal decomposition\.arXiv preprint arXiv:2409\.19250\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p4.1)\.
- \[20\]K\. Valmeekam, M\. Marquez, A\. Olmo, S\. Sreedharan, and S\. Kambhampati\(2023\)PlanBench: an extensible benchmark for evaluating large language models on planning and reasoning about change\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 38975–38987\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p4.1)\.
- \[21\]K\. Valmeekam, K\. Stechly, A\. Gundawar, and S\. Kambhampati\(2025\)A systematic evaluation of the planning and scheduling abilities of the reasoning model o1\.Transactions on Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=FkKBxp0FhR)Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p4.1),[TABLE II](https://arxiv.org/html/2609.21221#S3.T2.3.4.2.1),[TABLE II](https://arxiv.org/html/2609.21221#S3.T2.3.5.1.1),[TABLE II](https://arxiv.org/html/2609.21221#S3.T2.3.6.1.1),[TABLE II](https://arxiv.org/html/2609.21221#S3.T2.3.7.1.1),[TABLE III](https://arxiv.org/html/2609.21221#S3.T3.1.2.2.1),[TABLE III](https://arxiv.org/html/2609.21221#S3.T3.1.3.2.1),[§IV\-A](https://arxiv.org/html/2609.21221#S4.SS1.p2.1)\.
- \[22\]W\. Shen, C\. Garrett, A\. Goyal, T\. Hermans, and F\. Ramos\(2024\)Differentiable gpu\-parallelized task and motion planning\.arXiv preprint arXiv:2411\.11833\.Cited by:[§II](https://arxiv.org/html/2609.21221#S2.p5.1)\.
- \[23\]L\. Maes, Q\. L\. Lidec, D\. Scieur, Y\. LeCun, and R\. Balestriero\(2026\)LeWorldModel: stable end\-to\-end joint\-embedding predictive architecture from pixels\.arXiv preprint arXiv:2603\.19312\.Cited by:[§III\-F](https://arxiv.org/html/2609.21221#S3.SS6.p1.1)\.

Similar Articles

A better method for planning complex visual tasks

MIT News — Artificial Intelligence

MIT researchers developed VLMFP, a two-stage generative AI approach combining vision-language models with formal planning software to achieve 70% success rate on complex visual planning tasks like robot navigation, nearly 2.3x better than existing baselines. The method automatically translates visual scenarios into planning files that classical solvers can process, enabling effective long-horizon planning in novel environments.

Neuro-Inspired Inverse Learning for Planning and Control

arXiv cs.AI

This paper introduces a neuro-inspired framework called Inverter that uses Inverse Learning (IL) for fast and efficient planning and control, achieving significant improvements on D4RL benchmarks and quantum gate synthesis with orders of magnitude less inference computation.