Knowing When to Yield: Grounded Arbitration of User Corrections in Text-Based Embodied Agents

arXiv cs.AI Papers

Summary

The paper introduces GAVA, a framework that arbitrates whether an embodied agent should accept, reject, inspect the world, or ask for clarification when given a user correction, evaluated in text-only ALFWorld with training-derived object-location priors that reduce interaction and joint cost against uniform baselines.

arXiv:2610.00282v1 Announce Type: new Abstract: How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, 162 checkpoints produce 972 paired true and false interventions. Complete local inspections give GAVA and always verify 100 percent correction accuracy, establishing the evidence contract rather than a comparative advantage. In same-episode execution, GAVA reduces interaction cost against always verify but ties a cost threshold under a perfect speaker. An exploratory training-only object-location prior lowers interaction and declared joint cost on 340 unseen scenarios by 0.490 and 0.420 relative to uniform GAVA. After freezing the policy, costs, baselines, and multiplicity plan, the gains replicate on 77 non-overlapping seen checkpoints, covering 308 scenarios: 0.595 and 0.517, with both 95 percent checkpoint-bootstrap confidence intervals excluding zero. Joint cost also improves over an identical-prior fixed policy, while the matched calibrated no-VOI comparison remains inconclusive. Semantic GAVA makes four factual errors in each cohort, corresponding to 98.8 percent and 98.7 percent accuracy, and all methods complete every task. Results support selective information gathering with semantic priors under declared costs, but do not establish a general advantage of environmental value of information over clarification. The study uses normalized claims, complete symbolic observations, and controlled speakers; it evaluates neither human participants, visual input, nor physical robots.
Original Article
View Cached Full Text

Cached at: 10/02/26, 09:45 AM

# Knowing When to Yield: Grounded Arbitration of User Corrections in Text-Based Embodied Agents
Source: [https://arxiv.org/html/2610.00282](https://arxiv.org/html/2610.00282)
Conference:Proceedings of the 22nd ACM/IEEE International Conference on Human\-Robot Interaction; March 8–12, 2027; Santa Clara, CA, USACCS:Computing methodologies Planning under uncertaintyCCS:Human\-centered computing HCI theory, concepts and models,Runjia DuAffiliation:Independent Researcher,Zeming LiuAffiliation:Brown University,Hang LyuAffiliation:Brown University,Zehua YangAffiliation:Independent ResearcherandBojun LinAffiliation:Pinterest Inc

© , 2027

###### Abstract\.

How should an embodied agent respond when a person’s correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker\. GAVA implements this interface with observation\-bounded evidence, legal probes, and a one\-step expected\-loss rule\. We contribute an executable protocol and a controlled account of checking, clarification, and prior knowledge\. In text\-only ALFWorld, 162 checkpoints produce 972 paired true/false interventions\. Complete local inspections give GAVA and always verify 100% correction accuracy, establishing the evidence contract rather than a comparative advantage\. Same\-episode execution with an independent train\-prior planner counts realized actions and clarification exchanges\. With a perfect speaker, GAVA reduces interaction cost against always verify but ties a cost threshold\. A locally frozen fallible\-speaker study provides only nominal evidence for reliability calibration under a uniform prior\. An exploratory extension uses training\-only object–location counts as a claim prior\. On 340 unseen scenarios it lowers interaction and declared joint cost against uniform GAVA by−0\.490\-0\.490and−0\.420\-0\.420\. We then freeze the policy, costs, matched baselines, and multiplicity plan before internal replication on 77 non\-overlapping seen checkpoints \(308 scenarios\)\. The respective gains replicate:−0\.595\-0\.595and−0\.517\-0\.517, with both 95% checkpoint\-bootstrap CIs excluding zero\. Joint cost also improves against an identical\-prior fixed policy \(−0\.639\-0\.639, CI \[−1\.275,−0\.166\-1\.275,\-0\.166\]\)\. The identical\-prior calibrated no\-VOI comparison remains inconclusive \(−0\.208\-0\.208, CI \[−0\.630,0\.224\-0\.630,0\.224\]\); it establishes neither superiority nor equivalence\. Semantic GAVA makes four factual errors in each cohort \(98\.8%/98\.7% accuracy\)\. All methods complete every task\. The results support selective information gathering with semantic priors under declared costs\. They do not isolate a general benefit of environmental VOI over clarification\. The study uses normalized claims, complete symbolic observations, and controlled speakers; no human participants, visual input, or physical robot are evaluated\.

###### Keywords:

human–robot interaction, language correction, embodied agents, active verification, sycophancy, simulation

## 1\.Introduction

Language offers an appealing interface for correcting a robot while it works\. Prior systems map utterances such as “not there” or “do this instead” into plan edits, policy updates, or revised actions\([Thierauf et al\., 2024](https://arxiv.org/html/2610.00282#bib.bib7);[Sharma et al\., 2022](https://arxiv.org/html/2610.00282#bib.bib8);[Cui et al\., 2023](https://arxiv.org/html/2610.00282#bib.bib11);[Liu et al\., 2023](https://arxiv.org/html/2610.00282#bib.bib9);[Shi et al\., 2024](https://arxiv.org/html/2610.00282#bib.bib10)\)\. This literature largely studies whether a robot can*use*a correction\. It usually does not test whether the correction should be trusted\. That assumption is consequential: people can misremember an object location, refer to the wrong receptacle, or propose an action whose preconditions are not yet known\. A robot that treats every correction as supervision may exchange one planning error for another\.

Language\-model research calls uncritical agreement with a user’s false premise*sycophancy*, and distinguishes it from stubborn rejection of a valid correction\([Sinha, 2026](https://arxiv.org/html/2610.00282#bib.bib4);[Pi et al\., 2025](https://arxiv.org/html/2610.00282#bib.bib5);[Beigi et al\., 2025](https://arxiv.org/html/2610.00282#bib.bib6)\)\. A recent HRI scoping review frames robotic sycophancy as an emerging interaction concern and calls for dedicated measurement\([Seaborn and Yalçın, 2026](https://arxiv.org/html/2610.00282#bib.bib3)\)\. Existing benchmarks, however, mostly score an answer to a static question\. An embodied agent can act to obtain evidence, and its choice has step, dialogue, and physical consequences\. Thus the relevant question is not simply whether the model agrees; it is whether the agent chooses an appropriate, grounded response before changing the world\.

We study this question without visual rendering or human participants\. The text\-only ALFWorld interface aligns symbolic household tasks with an interactive language environment\([Shridhar et al\., 2021](https://arxiv.org/html/2610.00282#bib.bib1)\)\. It exposes legal actions and text observations while retaining hidden state for evaluation, allowing a strict separation between what a deployable method observes and what the experimenter uses to label a correction\.

This paper contributes:

1. \(1\)a paired protocol for factual corrections at matched early checkpoints across six task families\. Observable inputs are separated from hidden labels; three wording styles check the normalized\-claim interface;
2. \(2\)GAVA \(Grounded Advice Verification and Arbitration\), an auditable one\-step rule foraccept,reject,verify, andclarify\. Its complete\-listing contract makes positive and claim\-local negative evidence executable; and
3. \(3\)a controlled comparison separating evidence resolution from interaction efficiency: static ablations test the observation contract; online runs count actions and dialogue; an exploratory semantic prior and a prospectively frozen internal replication test cost savings against matched policies\.

Semantic priors reduce checking relative to uniform GAVA and always verify; the full policy lowers declared joint cost relative to prior\-only decisions\. Added value over matched calibrated clarification remains uncertain\. The HRI contribution is a correction\-handling interface and evaluation protocol\. Simulated users supply controlled interventions; the study does not establish human trust, natural dialogue competence, or physical safety\.

## 2\.Related Work

### 2\.1\.Language correction in robotics

Natural\-language feedback has been used to revise robot cost functions\([Sharma et al\., 2022](https://arxiv.org/html/2610.00282#bib.bib8)\), alter actions and goals\([Thierauf et al\., 2024](https://arxiv.org/html/2610.00282#bib.bib7)\), update policies\([Liu et al\., 2023](https://arxiv.org/html/2610.00282#bib.bib9);[Shi et al\., 2024](https://arxiv.org/html/2610.00282#bib.bib10)\), and provide shared\-autonomy corrections\([Cui et al\., 2023](https://arxiv.org/html/2610.00282#bib.bib11)\)\. These systems establish that verbal correction can be useful, but typically treat it as task\-relevant supervision\. Thierauf et al\. explicitly identify unsuitable or invalid task\-level corrections and uncertainty about the desired response as open limitations\([Thierauf et al\., 2024](https://arxiv.org/html/2610.00282#bib.bib7)\)\. GAVA addresses that boundary: it arbitrates factual/procedural claims before a downstream learner or planner uses them\.

### 2\.2\.Grounded feedback and closed\-loop planning

Embodied language agents improve when they feed environment observations back into planning\([Huang et al\., 2022](https://arxiv.org/html/2610.00282#bib.bib12);[Bhat et al\., 2024](https://arxiv.org/html/2610.00282#bib.bib13)\)\. Those feedback channels report execution progress or state; they are not conversational claims that may be true or false\. We combine closed\-loop state feedback with an explicit source distinction: observations update the belief store, whereas a correction remains a hypothesis until entailed, contradicted, or actively checked\.

### 2\.3\.Active query and clarification

Information\-seeking interaction is established in HRI\. Deits et al\. select clarifying questions by expected entropy reduction and merge human answers into a grounded command model\([Deits et al\., 2013](https://arxiv.org/html/2610.00282#bib.bib14)\)\. Cakmak and Thomaz show that robot query types impose different human costs and preferences\([Cakmak and Thomaz, 2012](https://arxiv.org/html/2610.00282#bib.bib15)\); grounded clarification has also been evaluated end to end\. Text agents can also learn when to query an external knowledge source\([Liu et al\., 2022](https://arxiv.org/html/2610.00282#bib.bib16)\), while ReSpAct interleaves reasoning, dialogue, and action in ALFWorld\([Dongre et al\., 2025](https://arxiv.org/html/2610.00282#bib.bib17)\), and SECURE uses embodied conversation to acquire previously missing concepts\([Rubavicius et al\., 2026](https://arxiv.org/html/2610.00282#bib.bib18)\)\. Thus neither VOI nor clarification is our novelty\. The distinct object here is a correction whose factual content may itself be wrong, and the choice between querying the world and querying the speaker: GAVA checks that source before using it as supervision, while the benchmark pairs true/false claims and isolates runtime observations from label state\.

### 2\.4\.Sycophancy and correction selectivity

SycoBench\-600 measures whether language models accept correct feedback and resist incorrect feedback\([Sinha, 2026](https://arxiv.org/html/2610.00282#bib.bib4)\); related work exposes a sycophancy–stubbornness trade\-off in visual question answering\([Pi et al\., 2025](https://arxiv.org/html/2610.00282#bib.bib5)\)and uses uncertainty\-aware reasoning to mitigate textual sycophancy\([Beigi et al\., 2025](https://arxiv.org/html/2610.00282#bib.bib6)\)\. Our work does not claim to originate*correction selectivity*\. It operationalizes selectivity in executable, partially observed household tasks, adds legal evidence\-gathering actions and their cost, and evaluates state\-grounded consequences rather than answer text\.

Table[1](https://arxiv.org/html/2610.00282#acmlabel1)makes the resulting gap explicit\. Robot\-correction work supplies the interaction setting but normally assumes useful guidance; sycophancy benchmarks vary feedback truth but lack actions; closed\-loop planners act on observations but do not arbitrate a fallible conversational source\. GAVA is intended as a bridge among these capabilities, not a replacement for the learning or planning methods in the first two rows\.

Table 1\.Relationship to adjacent research\. “Yes” means the capability is central to the cited evaluation, not merely possible in principle\.A comparison table showing that prior robot correction, closed\-loop planning, and sycophancy benchmarks each cover only subsets of executable action, fallible user feedback, active verification, and explicit observable\-state isolation; this work covers all four in a text simulator\.

## 3\.Grounded Correction Arbitration

Figure[1](https://arxiv.org/html/2610.00282#acmlabel2)summarizes the information boundary\. The method\-facing API omits the evaluation record, and the scorer never participates in action selection\. In integrated runs, a separate evaluation process owns the oracle; policy workers hold neither the label record nor the simulated\-user object\. Their only reply channel is a one\-shot capability bound to a specific clarification event\. This protects the audited trusted\-code execution path; it does not claim filesystem or OS sandboxing of adversarial method code\.

Public runtime:correctionctc\_\{t\}→\\rightarrowatomic claimsP⁡\(ct\)P\(c\_\{t\}\)→\\rightarrowobservable beliefBtB\_\{t\}→\\rightarrowevidence \{entailed, contradicted, unknown\}→\\rightarrowexpected\-loss/VOI arbitration→\\rightarrowaccept, reject, legal probe, or clarifyLegal feedback loop:probe action→\\rightarrowTextWorld observation→Bt⊕o\\rightarrow B\_\{t\}\\oplus o→\\rightarrowre\-arbitrationEvaluation\-only path \(separate file/process\):hidden facts→\\rightarrowprivate twin labels→\\rightarrowpost\-clarify text service and post\-hoc scoring

Figure 1\.GAVA pipeline and audited trusted\-code oracle boundary\. Hidden facts generate and score interventions but are absent from method inputs\.A three\-line flow diagram\. Public corrections become atomic claims, observable evidence, and arbitration decisions\. Legal actions return text observations to the belief state\. A separate evaluation path uses hidden facts only for labels and scoring\.At steptt, an agent has goalgg, observation historyhth\_\{t\}, and an observation\-bounded belief stateBtB\_\{t\}\. The interface represents correctionctc\_\{t\}as atomic propositionsP⁡\(ct\)=\{p1,…,pk\}P\(c\_\{t\}\)=\\\{p\_\{1\},\\ldots,p\_\{k\}\\\}and optional edits\. The benchmark supplies normalized claims; learned parsing is not evaluated\. For eachpip\_\{i\}, the verifier returns

\(1\)ei∈\{entailed,contradicted,unknown\}\.e\_\{i\}\\in\\\{\\textsc\{entailed\},\\textsc\{contradicted\},\\textsc\{unknown\}\\\}\.Missing evidence isunknown, not false\. Functional predicates add only ontology\-licensed contradictions\. We do not declare ALFWorld’sa​tatrelation globally functional: co\-located receptacle names can both list the same object\. A location claim is contradicted only by a complete local listing of the claimed receptacle that omits the object\.

The response set isD=\{accept,reject,verify,clarify\}D=\\\{\\textsc\{accept\},\\textsc\{reject\},\\textsc\{verify\},\\textsc\{clarify\}\\\}\. For an unknown claim with model priorπ=Pr⁡\(pi∣Bt\)\\pi=\\Pr\(p\_\{i\}\\mid B\_\{t\}\), GAVA compares one\-step expected asymmetric loss within this response class:

\(2\)RA\\displaystyle R\_\{A\}=\(1−π\)​CF​A,\\displaystyle=\(1\-\\pi\)C\_\{FA\},RR\\displaystyle R\_\{R\}=π​CT​R,\\displaystyle=\\pi C\_\{TR\},\(3\)RF\\displaystyle R\_\{F\}=min⁡\{RA,RR\},\\displaystyle=\\min\\\{R\_\{A\},R\_\{R\}\\\},RC​L\\displaystyle R\_\{CL\}=CC​L\+\(1−γ\)​RF\\displaystyle=C\_\{CL\}\+\(1\-\\gamma\)R\_\{F\}\(4\)\+γ⁡\(1−η\)​\[π​CT​R\+\(1−π\)​CF​A\],\\displaystyle\\quad\+\\gamma\(1\-\\eta\)\[\\pi C\_\{TR\}\+\(1\-\\pi\)C\_\{FA\}\],whereCF​AC\_\{FA\}andCT​RC\_\{TR\}penalize false acceptance and true rejection;CC​LC\_\{CL\}is clarification cost\. Resolution probabilityγ\\gammaand conditional answer accuracyη\\etaare assumed for either truth value\. Equation \([4](https://arxiv.org/html/2610.00282#S3.E4)\) evaluates a restricted policy: follow a resolved reply; otherwise use the best fixed response\. It does not perform a posterior Bayes decision on the reply\. Letr0=min⁡\{RA,RR,RC​L\}r\_\{0\}=\\min\\\{R\_\{A\},R\_\{R\},R\_\{CL\}\\\}\. A public probeqqhas a legal macro, costC⁡\(q\)C\(q\), and probabilityρq\\rho\_\{q\}of complete, truthful resolution\. Non\-resolution is modeled as leaving the prior unchanged and incurring fallback riskr0r\_\{0\}\. The one\-step probe risk and value are

\(5\)Rq=C⁡\(q\)\+\(1−ρq\)​r0,VOI⁡\(q\)=r0−Rq\.R\_\{q\}=C\(q\)\+\(1\-\\rho\_\{q\}\)r\_\{0\},\\qquad\\mathrm\{VOI\}\(q\)=r\_\{0\}\-R\_\{q\}\.GAVA selects theqqwith minimumRqR\_\{q\}only whenVOI⁡\(q\)\>τ\\mathrm\{VOI\}\(q\)\>\\tau; otherwise it returns the minimum\-risk non\-probe decision\. Equal fixed risks favor accept; equality between clarification and the fixed fallback favors the fixed response\. In this studyτ=0\\tau=0\. Static response selection assumesγ=η=1\\gamma=\\eta=1but stops at a clarification request \(Section[6\.3](https://arxiv.org/html/2610.00282#S6.SS3)\); integrated runs execute the reply\. The frozen reliability study varies both parameters by checkpoint\. Uniform\-prior runs setπ=\.5\\pi=\.5\. The exploratory semantic variant setsπ\\piwithin the exposed pair to\(n⁡\(x,r\)\+1\)/\(n⁡\(x,r\)\+n⁡\(x,r′\)\+2\)\(n\(x,r\)\+1\)/\(n\(x,r\)\+n\(x,r^\{\\prime\}\)\+2\), wherennis the frozen ALFWorld\-training object–location count andr′r^\{\\prime\}is the public symmetric mate; it reads no test label\. This is a smoothed, pair\-conditioned estimate; ranking accuracy does not establish probability calibration\. Entailed and contradicted evidence is hard under the benchmark’s complete\-listing contract \(π=1\\pi=1and00, respectively\), so its residual fixed\-decision loss is zero\. A noisy deployment would replace this contract with a calibrated observation modelP⁡\(o∣Bt,q\)P\(o\\mid B\_\{t\},q\)and the full expectation; that extension is not evaluated here\.

### 3\.1\.Decision procedure

GAVA queriesBtB\_\{t\}: entailed claims are accepted and contradicted claims rejected without action\. Unknown claims use the best fallback over accept, reject, and clarify\. A location probe requires admissiblego to r\. Its macro navigates torrand opens it only ifopen ris then admissible\. Complete listings licenseρq=1\\rho\_\{q\}=1from public commands and observation semantics, without hidden facts\.

For complete probes, Equation \([5](https://arxiv.org/html/2610.00282#S3.E5)\) givesRq=C⁡\(q\)R\_\{q\}=C\(q\): inspection is chosen exactly when its cost is below the best non\-probe risk\. With symmetric lossCF​A=CT​R=CC\_\{FA\}=C\_\{TR\}=C, this condition becomes

\(6\)C⁡\(q\)<min⁡\{C​min⁡\(π,1−π\),RC​L\}\.C\(q\)<\\min\\\{C\\min\(\\pi,1\-\\pi\),R\_\{CL\}\\\}\.This explains the perfect\-speaker cost\-threshold tie\. Semantic priors and fallible replies alter the fallback risk, not the decision rule\.

Figure[2](https://arxiv.org/html/2610.00282#acmlabel3)shows complete\-listing execution; only returned observations updateBtB\_\{t\}\. Incomplete descriptors useρq<1\\rho\_\{q\}<1and Equation \([5](https://arxiv.org/html/2610.00282#S3.E5)\)’s one\-step fallback; repeated unresolved probing is not evaluated\. This local risk omits downstream search costs and does not guarantee minimum episode cost\.

Input:correction claimsP⁡\(ct\)P\(c\_\{t\}\), observable beliefBtB\_\{t\}, admissible actionsAtA\_\{t\}, and costs\. For each claimpp:\(1\)Querye←Bt​\(p\)e\\leftarrow B\_\{t\}\(p\)\.\(2\)If entailed, accept; if contradicted, reject\.\(3\)Otherwise compute the minimum\-risk non\-probe fallback over accept, reject, and clarify\.\(4\)Construct probesqqfrom admissible commandsAtA\_\{t\}that can revealpp\.\(5\)If the best probe has strictly positive VOI, execute it, updateBtB\_\{t\}only from its returned observation, and go to step 1\.\(6\)Otherwise return the non\-probe fallback\.

Figure 2\.GAVA under complete\-listing probes \(ρq=1\\rho\_\{q\}=1\)\. Evaluation\-only facts are not inputs\.Pseudocode that queries observable evidence, handles entailed and contradicted claims, compares unknown\-claim fallbacks with legal probes, updates only from returned observations, and repeats\.
### 3\.2\.Local negative evidence

Open\-world reasoning creates a subtle failure mode: observing a listed object is positive evidence, but global absence is not evidence of falsehood\. ALFWorld nevertheless returns a complete visible\-content enumeration after reaching an open surface or opening a container\. GAVA therefore licenses a negative¬a​t​\(x,r\)\\neg at\(x,r\)only when \(1\) the agent legally inspectedrr, \(2\) the environment returned a complete local listing, and \(3\)xxwas absent\. It does not infer that an unmentioned object is absent elsewhere\. The implementation also binds the tightly scoped anaphor “in it” only to a container confirmed as opened in the same observation\. It likewise does not infer¬a​t​\(x,r2\)\\neg at\(x,r\_\{2\}\)merely from seeingxxatr1r\_\{1\}, because ALFWorld can expose one object through multiple co\-located receptacle names\.

### 3\.3\.Claim\-level decisions

Decisions are made per proposition\. A multi\-claim correction can therefore be partially accepted: supported claims enter the planner, contradicted claims are rejected, and unknown claims remain pending\. The present confirmatory benchmark uses one controlled location claim per intervention to isolate arbitration from semantic parsing; multi\-claim behavior is covered by executable unit tests and left for a broader benchmark\.

### 3\.4\.Worked trace

One unseen task asks the agent to place a saltshaker in a drawer\. The simulated correction claims thatsaltshaker 1is atcabinet 2\. Initially the proposition is unknown and both accepting and rejecting have expected loss 0\.5\. With clarification cost 0\.25 and normalized inspection\-macro cost 0\.10, GAVA selectsverify\(VOI 0\.15\)\. The first legal action reports that the cabinet is closed, which is not location evidence\. GAVA opens it; the complete response is “In it, you see a glassbottle 1, and a plate 2\.” The contextual “it” is bound to the cabinet confirmed open in that same observation\. Because the complete local list omits the saltshaker, the claim becomes contradicted and is rejected\. At no point does the runtime access the oracle’s true location\.

## 4\.Paired Embodied\-Correction Benchmark

### 4\.1\.Environment and tasks

We use ALFWorld’s TextWorld interface, not its Unity renderer\([Côté et al\., 2018](https://arxiv.org/html/2610.00282#bib.bib2)\)\. The six task families require locating, moving, examining, cleaning, cooling, heating, or placing household objects\. Each record stores the goal, observation, legal commands, replay prefix, correction text, and normalized claim\. The method never receives simulator predicates or the expert policy\.

### 4\.2\.Intervention construction

For each accepted game, the evaluation\-only generator reads the initial expert policy and hidden facts\. If the policy begins by navigating to and taking a target object from a factually validated receptacle, it creates a true location claim and a minimally different false twin naming another currently reachable receptacle\. To prevent the false\-member construction from encoding the label, all legal locations are hash\-ranked from public checkpoint fields and paired by an adjacent, fixed\-point\-free involution: ifAAmaps toBB, the same public rule mapsBBtoAA\. With an odd number of locations, the last location is unmatched; a checkpoint whose true location is unmatched is excluded rather than given a directional fallback\. Both twins share the goal, checkpoint, object, legal\-action set, and wording template\. We render neutral, doubt, and authority styles\. Templates change pragmatic pressure but not propositional content\.

Pilot games are excluded from the unseen test set by path\. We evaluate separate ALFWorldvalid\_seenandvalid\_unseensplits; these names describe ALFWorld’s learned\-agent split, not newly sampled domains\. GAVA fits no policy, but its semantic variant estimates a prior from training counts\. Invalid cases are rejected before any method output is inspected: short policies, non\-location openings, policy/fact mismatches, absence of a legal false location, or an unmatched true location\. Checkpoint, pair, and skip counts plus file hashes are saved in machine\-readable manifests\.

Table[2](https://arxiv.org/html/2610.00282#acmlabel4)reports the complete selection flow\. All 134 unseen games were enumerated; after removing 26 pilot games, all remaining 108 were inspected\. For the secondary seen analysis, predeclared per\-family quotas caused inspection to stop after 111 of 140 games, leaving 29 uninspected\. The 21 seen timeouts, rather than being scored as failures, are outside the selected stratum and are a stated validity threat\.

Table 2\.Checkpoint selection flow\. “Other” includes non\-location policy openings, policy/fact mismatches, and unmatched locations\.The complete flow from all candidate games through pilot or quota exclusions, inspection, rejection, and final accepted checkpoints\.Table 3\.Frozen benchmark composition\. “Pairs” counts pressure\-specific true/false twins; the statistical unit is the checkpoint\.The unseen split contains 85 checkpoints, 255 true\-false pairs, and 510 interventions\. The seen split contains 77 checkpoints, 231 pairs, and 462 interventions\.
### 4\.3\.Why the twins remain partially observed

The intervention is injected before the base policy’s first navigation action\. The reset observation enumerates reachable furniture but not the target’s location\. Consequently, both the true and false member begin as unknown to the runtime method even though the generator can validate them against hidden facts\. This design rules out a trivial strategy that simply restates an already observed location\. It creates the same observation\-state obligation for both twins, but does not assume equal semantic plausibility: training\-world knowledge may favor one member, which we audit and use in a claim\-adaptive extension\. Once a complete local listing is returned, positive and claim\-bounded negative evidence are equally actionable\.

### 4\.4\.Oracle isolation and leakage controls

Public and oracle JSONL files are physically separate\. Public records contain the claim but no truth label, true location, hidden fact list, or expert policy\. Their IDs are 24\-hex hashes of public content, and rows are sorted by opaque ID rather than private twin order\. Tests over both released splits forbid truth\-bearing keys/tokens and enforce the format: all 2,916 ID fields pass, and the lexicographically first member is true in 250/486 pairs\. Every released pair passes the public involution in both directions, so either member is construction\-consistent as the true claim\. As negative controls, the former minimum\-location rule scores 50\.6% unseen and 49\.4% seen, while a deterministic public pair\-hash guess scores 51\.4% and 49\.8%; all Wilson intervals include 50%\. These checks rule out the identified deterministic construction decoder, not semantic inference from ordinary task knowledge\. Runners accept only the public path; evaluation code alone accepts both files\. Tests also verify that the adapter requests observations and legal commands, not simulator facts\. For integrated clarification, a dedicated service process owns oracle and simulated\-user state\. Policy workers hold only authenticated transport data and one scenario/claim\-bound, one\-shot capability\. Simulated\-user state remains server\-side; unknown, reused, premature, or mismatched capabilities are rejected\. The service is invoked lazily after choosingclarify\. This is a hardened audited boundary for trusted method code, not an OS sandbox against malicious code or direct access to installed ALFWorld files\.

## 5\.Evaluation

### 5\.1\.Research questions and outcomes

We ask: \(RQ1\) does grounded arbitration resolve paired true/false claims; \(RQ2\) which observation and probe mechanisms enable resolution; \(RQ3\) does the rule choose lower model risk across heterogeneous evidence, probes, reliability, cost, and loss; \(RQ4\) how do responses affect task success and action cost; and \(RQ5\) can a training\-only semantic prior reduce checking, replicate internally, and improve on policies with the same prior but no environmental VOI?

The static primary outcome is accuracy among resolved accept/reject responses: accept true claims or reject false ones\. Coverage is separate\. A clarification request is pending, not an answer\. Online runs execute probes and replies before scoring correctness\. Other outcomes are false acceptance, false rejection, their union as incorrect fixed\-decision \(IFD\) rate, legal actions, extra steps, model risk, and task success\.

For perfect\-speaker RQ4, the primary comparison is realized interaction cost: actions plus cost\-weighted, counted clarification exchanges\. The locally frozen fallible\-speaker study uses declared joint cost: interaction cost plus realized false\-accept or true\-reject loss\. These are declared decision units, not measured harm\. Components and correctness are reported separately; eventual task recovery can conceal correction errors\.

Method comparisons measure uncertainty at the base checkpoint, averaging its twins and timings\. Wording styles check the normalized\-proposition interface, not human persuasion\.

### 5\.2\.Systems and ablations

We compare full GAVA with always accept, always reject, always verify, and a feasibility\-only policy that accepts any correction whose proposed navigation is legal; the stress test also uses the minimum\-risk fixed\-only fallback, including clarification\. We ablate complete\-listing negative evidence and automatic opening of a closed receptacle\. A cost sweep variesC⁡\(q\)C\(q\)around the VOI threshold\. A public\-input\-only component stress test instantiates 11 declared decision contexts at each checkpoint, including already known evidence, absent or unreliable probes, competing probes, the VOI boundary, and asymmetric losses\.

For RQ4, we use a deterministic reactive planner that receives only the public goal, normalized target object, current text observation, and admissible commands\. Its initial search order uses object–receptacle counts estimated from 6,353 ALFWorld training trajectories; held\-out expert policies and oracle facts are unavailable during execution\. After every legal action it replans from the returned text, searches the next receptacle after a failed hypothesis, and implements all six task families\. This deliberately transparent planner isolates the downstream effect of arbitration; it is not presented as a learned or frontier LLM planner\. In the integrated condition, arbitration runs online in the same TextWorld episode either before planning or after the first train\-prior navigation action\. Accepting adds the claimed location as a search preference, rejecting adds none, verifying executes the currently legal navigation/opening macro and re\-arbitrates from its text, and clarification returns one controlled text reply whose parsed answer updates the preference\. The reliability study additionally compares calibrated GAVA with the same policy assuming a perfect speaker, a calibrated policy without environmental VOI, and the grounded raw\-cost threshold\. The claim\-adaptive extension changes onlyπ\\piusing the frozen train prior\. After its exploratory unseen run, we freeze a non\-overlapping seen replication with two matched baselines: prior\-only fixed Bayes choice and calibrated accept/reject/clarify without environmental VOI\. All methods run online; four primitive arms additionally validate a scenario\-complete policy replay\.

### 5\.3\.Implementation and analysis

All correction decisions use deterministic rules\. Uniform GAVA usesCF​A=CT​R=1C\_\{FA\}=C\_\{TR\}=1, priorPr⁡\(p\)=0\.5\\Pr\(p\)=0\.5when unknown, clarification cost0\.250\.25, and normalized inspection\-macro cost0\.100\.10unless varied\. This macro may contain a navigation action and, for a closed receptacle, one opening action; actual environment steps are reported separately, and the stress test varies resolution probability and loss\. Entailed and contradicted propositions are hard evidence under the complete\-listing contract\. The integrated experiment instead uses an environment\-action cost of 1, clarification cost 1\.75, andCF​A=CT​R=6C\_\{FA\}=C\_\{TR\}=6\. A fixed SHA\-256\-ranked sample of 100 training games \(1,199 legal receptacle probes\) estimates expected probe length by receptacle type; held\-out execution records the realized one\- or two\-action macro\. This makes a surface check cheaper than dialogue while the expected cost of a closed/openable container can exceed it\. These are declared interaction units, not estimates of human time or physical risk\. We locally froze the reliability design before generating its outcomes; it was not preregistered or externally timestamped\. We SHA\-256\-ranked the original 89 unseen checkpoints into a fixed3×33\\times 3design withCC​L∈\{0\.75,1\.75,3\.0\}C\_\{CL\}\\in\\\{0\.75,1\.75,3\.0\\\}and low/medium/high\(γ,η\)\(\\gamma,\\eta\)values of\(0\.45,0\.80\)\(0\.45,0\.80\),\(0\.70,0\.90\)\(0\.70,0\.90\), and\(0\.95,0\.98\)\(0\.95,0\.98\)\. Cell assignment used only public checkpoint paths\. After the symmetric\-pair repair, we removed the four ineligible checkpoints without reassigning any survivor; the final cells contain 8–10 checkpoints\. Deterministic hashes of the public, truth\-independent pair ID instantiate whether a lazily queried reply resolves and is correct, coupling both twins without exposing which is true\. An unresolved reply takes the same minimum\-risk fixed fallback used inRC​LR\_\{CL\}\. These parameters are controlled stress levels, not human estimates\. “Calibrated” means supplied speaker reliability here, not calibration learned from people or calibration of the semantic prior\. The test suite contains 72 tests covering open\-world semantics, local anaphora, arbitration, record validation, lazy reply gating, action legality, planner state continuity, and oracle leakage\. We report paired checkpoint bootstrap intervals for method differences and do not treat the three templates as independent samples\. The initial semantic extension and its sweeps remain post\-review exploratory\. Before new seen policy outcomes for this extension, we froze its 77\-checkpoint internal replication, matched baselines, six comparisons, and Holm family; this is prospective only for those outcomes, not a preregistration or new domain\. Closed\-loop comparisons average the true/false twins—and, for integrated runs, both intervention timings—within each checkpoint before 10,000 bootstrap draws and a 20,000\-draw paired sign\-flip test \(seed 20270909\)\. Scenario accuracy is reported as error counts and descriptive percentages: twins and timings are correlated, so scenario\-level binomial intervals are not used as uncertainty estimates\. Prior\-ranking Wilson intervals use one pair per checkpoint\.

The main runtime uses Python 3\.11\.14 \(x86\_64 under macOS\), ALFWorld 0\.4\.2, TextWorld 1\.7\.0, and a project\-local data directory\. Seeds, record hashes, validation reasons, and raw action/dialogue traces are saved\. The statistical script uses 10,000 checkpoint bootstrap draws with seed 20270918\.

#### Baseline semantics\.

Always accept and always reject resolve without inspecting\. Feasibility\-only accepts if the correction’s proposed navigation is currently admissible; because both twins name reachable receptacles, it tests whether executability alone can establish truth\. No correction follows the train\-prior search order in closed\-loop evaluation\. Always accept makes the claimed location the first search preference; always reject ignores it\. Always verify follows the same legal information\-gathering trajectory whenever one is available but does not compare its cost to dialogue\. Always clarify requests the controlled reply in every case\. A transparent grounded cost heuristic verifies iff the train\-estimated probe length is below clarification cost; under the present hard\-evidence, symmetric\-loss contract it is the closed\-form special case of GAVA and is expected to tie it when the speaker is perfect\. In the reliability study, uncalibrated GAVA likewise assumesγ=η=1\\gamma=\\eta=1; calibrated no\-VOI receives the declared\(γ,η\)\(\\gamma,\\eta\)but cannot inspect the environment\.

## 6\.Results

### 6\.1\.Dataset yield and validation

The frozen evaluation contains 85 unseen and 77 seen checkpoints, or 162 independent checkpoints total\. Each checkpoint produces a true/false twin under three wording styles, yielding 972 interventions\. The unseen generator enumerated 134 games, excluded the 26\-game pilot, and rejected 19 of the 108 inspected candidates; four otherwise eligible checkpoints were excluded by the symmetric matching rule\. The seen generator inspected 111 of 140 games under predeclared per\-family quotas: 21 exceeded the 10\-second bound, nine had a non\-location opening, three had policy/fact mismatches, and one eligible checkpoint was unmatched; 29 were uninspected after quotas filled\. Independent reconstruction produced byte\-identical public and oracle files\. All 72 tests passed\.

### 6\.2\.Correction resolution

Static coverage counts terminal accept/reject responses; accuracy is conditional on coverage\. Clarification requests remain pending\. Full GAVA and always verify achieve 100% coverage and accuracy with 0% incorrect fixed decisions\. GAVA uses 1\.14/1\.23 probe actions per unseen/seen intervention\. Always accept, always reject, and feasibility\-only each have 50% accuracy and 50% incorrect decisions: twins have opposite truth but both locations are reachable\.

Without local negative evidence, true claims are accepted after observation but no false claim is rejected\. Coverage is 50%, with every resolved response correct\. Without opening containers, coverage falls to 85\.9% unseen \(checkpoint\-bootstrap 95% CI \[80\.6, 90\.6\]\) and 77\.3% seen \(\[71\.4, 82\.5\]\); resolved decisions remain correct\. Accuracy alone would conceal this non\-resolution\.

These ablations test the observation contract\. Complete local evidence permits resolution; removing it reduces coverage\. Feasibility alone provides no truth evidence\. At the default low cost, GAVA ties always verify because every unknown claim is legally inspectable\. The experiment does not test noisy perception or natural\-language understanding\.

### 6\.3\.Heterogeneous arbitration

With symmetric loss,π=\.5\\pi=\.5, and clarification cost 0\.25, GAVA verifies at probe costs 0\.00–0\.24\. At 0\.25 or above it selectsclarify\. The static harness stops at this request: coverage and incorrect fixed decisions are zero\. Equation \([4](https://arxiv.org/html/2610.00282#S3.E4)\) prices a subsequent answer that is not executed here\. This tests response selection, not realized dialogue loss\. The following online experiments execute and score replies\.

The 11 public descriptors per checkpoint \(1,782 decisions\) vary evidence, probe availability, reliability, cost, the VOI boundary, and asymmetric loss\. Mean model risk is 0\.176 for GAVA, 0\.192 for always verify, and 0\.205 for the best fixed fallback\. Losses use the same model that selects responses\. This checks the rule on designed cases, rather than independently validating task gains or probability calibration\.

### 6\.4\.Integrated online arbitration and planning

The online experiment invokes arbitration in the same TextWorld episode as the independent planner\. Every neutral true/false twin is run at both initial and post\-first\-navigation timings: 340 unseen and 308 seen scenarios per method\. The 100\-game training sample yields 1,199 resolved probe macros\. On held\-out always\-verify runs, 266/48 unseen probes and 212/70 seen probes take one/two actions; three unseen unknown claims have no currently legal probe\. Thus costs and availability vary through environment interaction rather than declared stress\-test descriptors\.

The perfectly informative\-speaker boundary first establishes online plumbing\. GAVA lowers mean interaction cost relative to always verify by−0\.144\-0\.144unseen \(95% CI \[−0\.278,−0\.028\-0\.278,\-0\.028\]\) and−0\.367\-0\.367seen \(\[−0\.714,−0\.156\-0\.714,\-0\.156\]\), while completing every task without invalid actions\. It nevertheless ties the grounded cost threshold exactly on both splits\. Its unseen comparison with always accept is uncertain \(\+0\.018\+0\.018, CI \[−0\.149,0\.174\-0\.149,0\.174\]\)\. Final\-response mutations change planner preference and trajectory in 12/12 unseen and 6/6 seen cases, excluding a logging\-only explanation\.

### 6\.5\.Frozen fallible\-speaker evaluation

The locally pre\-outcome\-frozen reliability study assigns each unseen checkpoint to one of nine dialogue\-cost/informativeness cells without reading oracle labels or method outputs\. All eight methods run on the same 340 paired scenarios, yielding 2,720 main outcomes\. Every method again has 100% task success and zero invalid actions\. Interaction cost alone cannot distinguish a cheap but credulous response\. Figure[3](https://arxiv.org/html/2610.00282#acmlabel6)summarizes the replicated effects and the seen\-split cost–error frontier for the subsequent semantic\-prior extension\.

Figure 3\.Semantic arbitration results\. \(a\) Paired mean differences and 95% checkpoint\-bootstrap CIs; negative values favor Semantic GAVA\. U is the unseen exploratory cohort and S the prospectively frozen seen replication\. Matched baselines appear only in S\. \(b\) Seen interaction cost versus incorrect fixed\-decision \(IFD\) rate\. Dashed contours show equal declared joint cost\. All methods complete every task\.Panel a is a forest plot\. Semantic GAVA has lower interaction and declared joint cost than uniform GAVA and always verify on both cohorts, and lower declared joint cost than prior\-only on the seen replication\. Its interval relative to calibrated no\-VOI crosses zero\. Panel b plots interaction cost against the percentage of incorrect fixed decisions\. Semantic GAVA lies near the lower\-left corner; no\-VOI uses slightly less interaction but makes more factual errors\.Uniform\-prior GAVA initially accepts 21 cases, clarifies 10, rejects three, and verifies 306; after evidence or dialogue it makes 170 accepts, six corrected responses, and 164 rejections, all factually correct\. Uncalibrated GAVA and the raw cost heuristic coincide, but make 7\.6% incorrect fixed decisions\. Uniform\-prior GAVA’s delta in declared joint cost against either is−0\.412\-0\.412\(CI \[−0\.793,−0\.029\-0\.793,\-0\.029\], rawp=\.035p=\.035\), but 48/85 checkpoints tie and the four\-comparison Holm\-adjustedp=\.104p=\.104makes this nominal evidence\. Interaction cost is indistinguishable \(\+0\.047\+0\.047, CI \[−0\.168,0\.313\-0\.168,0\.313\]\)\.

Versus calibrated no\-VOI, uniform\-prior GAVA’s delta in declared joint cost is−1\.796\-1\.796\(CI \[−2\.150,−1\.393\-2\.150,\-1\.393\]\), supporting environmental evidence within the uniform\-prior setting\. Against always verify, it is−0\.073\-0\.073\(CI \[−0\.174,−0\.004\-0\.174,\-0\.004\]\)\. Although this checkpoint\-bootstrap interval for the mean excludes zero, the predeclared paired sign\-flip test does not reject the paired null \(p=\.122p=\.122\); with 81/85 ties, we therefore do not treat this as robust evidence of broad dominance\.

The high\-information, cost\-0\.75 cell favors the uncalibrated policy by 0\.863 units of declared joint cost, so calibration is not uniformly beneficial\. All 793 lazy reply calls equal counted clarification exchanges\.

### 6\.6\.Semantic prior and robustness extension

A frozen training\-only object–location prior identifies the true member in 71/85 unseen pairs \(83\.5%, Wilson CI \[74\.2,89\.9\]\) and 69/77 seen pairs \(89\.6%, \[80\.8,94\.6\]\)\. These checkpoint\-level rates show ranking information, not calibrated probabilities\. Construction symmetry blocks the generator decoder while retaining semantic information that a \.5 prior discards\.

On 340 unseen scenarios, Semantic GAVA’s deltas in interaction cost and declared joint cost versus uniform GAVA are−0\.490\-0\.490\(CI \[−0\.609,−0\.363\-0\.609,\-0\.363\]\) and−0\.420\-0\.420\(\[−0\.579,−0\.218\-0\.579,\-0\.218\]\)\. Against always verify, the interaction\-cost delta is−0\.563\-0\.563\(CI \[−0\.659,−0\.469\-0\.659,\-0\.469\]\); the declared\-joint\-cost delta is−0\.493\-0\.493\(CI \[−0\.638,−0\.305\-0\.638,\-0\.305\]\)\. All four Holm\-adjustedp<\.001p<\.001\. It makes 4/340 factual errors \(98\.8% descriptive accuracy\) while completing every task without invalid actions\. Its online responses, trajectories, dialogue, and costs match primitive\-arm replay in all 340 cases\.

We then froze the same policy, costs, eight methods, and six\-test Holm family before a full run on the 77 non\-overlapping seen checkpoints\. Across 308 scenarios, Semantic GAVA again lowers interaction and declared joint cost versus uniform GAVA by−0\.595\-0\.595and−0\.517\-0\.517\(CIs \[−0\.726,−0\.459\-0\.726,\-0\.459\] and \[−0\.714,−0\.252\-0\.714,\-0\.252\]\)\. Versus always verify, the corresponding reductions are−0\.655\-0\.655and−0\.577\-0\.577\(\[−0\.773,−0\.537\-0\.773,\-0\.537\] and \[−0\.765,−0\.321\-0\.765,\-0\.321\]\); all four adjustedp=\.00030p=\.00030\. Against an identical\-prior fixed Bayes rule, its delta in declared joint cost is−0\.639\-0\.639\(\[−1\.275,−0\.166\-1\.275,\-0\.166\], adjustedp=\.0169p=\.0169\) with 4 versus 32 factual errors\. Against an identical\-prior, calibrated no\-VOI policy, however, the delta in declared joint cost is−0\.208\-0\.208\(\[−0\.630,0\.224\-0\.630,0\.224\],p=\.359p=\.359; 63/77 ties\)\. This establishes neither superiority nor equivalence\. Semantic GAVA’s 4/308 factual errors give 98\.7% descriptive accuracy\. Online and primitive replay match in 308/308 cases; all methods again achieve complete task success and zero invalid actions\.

The cost trade\-off is directly interpretable: four errors at loss 6 add24/340=0\.07124/340=0\.071unseen and24/308=0\.07824/308=0\.078seen to mean interaction cost, explaining the smaller joint\-cost gains after rounding\. These are arithmetic consequences of the reported counts, not additional outcomes\. The earlier calibration/loss sweeps remain synthetic sensitivity analyses\. Static wording invariance is an implementation check, not evidence about human language\.

## 7\.Discussion

### 7\.1\.What the comparisons establish

The interface separates using a claim from gathering evidence\. Accept changes search preference; reject withholds it; verify inspects the world; clarify asks the speaker\. For HRI, this makes environmental checking and speaker consultation distinct, costed choices before using factual advice\. Whether people understand, prefer, or trust these responses requires human evaluation\.

Coverage matters: without local negative evidence, every resolved response is correct, yet no false claim is rejected\. Full GAVA resolves both twins through inspection\. Non\-resolution does not establish repeated questioning or stalling: the static harness stops at a clarification request\. Online runs measure the actions and executed replies that connect response choices to planner behavior\.

The experiments separate three questions\. Static ablations establish that complete local evidence resolves factual claims\. Semantic priors then reduce interaction and joint cost against uniform GAVA and always verify in both cohorts\. Finally, the full policy lowers joint cost against an identical\-prior fixed rule in the seen replication: prior\-only commitment leaves avoidable declared loss\.

Whether inspection adds value beyond calibrated clarification remains open: 63/77 checkpoints tie and the matched CI spans zero\. The evidence supports the configurable interface and its measured cost–accuracy trade\-off, not a general advantage for environmental VOI\. This matched baseline and the perfect\-speaker threshold tie define the limits of the contribution\.

### 7\.2\.Pathway to physical deployment

The claim and arbitration interfaces could connect navigation, manipulation, and a perception\-backed belief store\. Physical deployment would require calibrated sensing, skill\-success monitoring, collision and force constraints, irreversibility\-aware loss, and recovery policies\. These requirements describe an implementation pathway; transfer beyond the present planner and simulation remains untested\.

### 7\.3\.Limitations

The benchmark uses templated English, one normalized location claim, and two early intervention timings\. Complete symbolic listings bypass uncertain perception\. Speech recognition, ambiguous reference, deceptive intent, asynchronous intervention, and social cues are absent\. False twins need not resemble natural mistakes\. Navigation abstracts geometry and makes checking cheap\.

Online clarification supplies one controlled reply, sometimes a correct replacement location, from an isolated service\. Reliability levels and action/error exchange rates are synthetic, not human estimates\. “Calibrated” policies receive speaker parameters rather than learning them\. Equation \([4](https://arxiv.org/html/2610.00282#S3.E4)\) follows resolved replies instead of optimally combining each answer with its prior\. Dialogue cost and error loss do not measure physical harm\.

The semantic prior is a smoothed, pair\-conditioned training estimate\. Ranking accuracy does not validate its probabilities or transfer to unpaired claims, changed correction frequencies, or shifted locations\. Multi\-claim behavior is covered only by software tests\. Incomplete probes are descriptors, not noisy trajectories; the one\-step rule is not a full belief\-space planner\.

The planner is deterministic and hand engineered, using training counts\. The seen replication is prospective for new policy outcomes, not an untouched domain: static results and aggregate prior predictability were already known\. Every method completes every task, and semantic GAVA’s four errors per cohort are recoverable\. The study measures checking cost and declared decision loss, not prevention of task failure or physical harm\.

## 8\.Conclusion

GAVA and the paired protocol make factual correction handling an executable choice among acceptance, rejection, inspection, and clarification\. A semantic prior reduces interaction and declared joint cost against uniform GAVA and always verify across exploratory and internally replicated cohorts; the full policy also lowers joint cost against identical\-prior fixed decisions\. Four errors per cohort expose the trade\-off\. Added value over matched calibrated clarification remains inconclusive\. The contribution is selective checking under explicit evidence and cost assumptions\. Broader correction distributions, calibrated uncertainty, and human or physical evaluation are required for claims about dialogue competence or safety\.

## References

- Beigiet al\.\(2025\)M\. Beigi, Y\. Shen, P\. Shojaee, Q\. Wang, Z\. Wang, C\. K\. Reddy, M\. Jin, and L\. HuangSycophancy mitigation through reinforcement learning with uncertainty\-aware adaptive reasoning trajectories\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 13079–13092\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.661)Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p2.1),[§2\.4](https://arxiv.org/html/2610.00282#S2.SS4.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.5.1.1.1)\.
- Bhatet al\.\(2024\)V\. Bhat, A\. U\. Kaypak, P\. Krishnamurthy, R\. Karri, and F\. KhorramiGrounding LLMs for robot task planning using closed\-loop state feedback\.External Links:2402\.08546Cited by:[§2\.2](https://arxiv.org/html/2610.00282#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.3.1.1.1)\.
- Cakmak and Thomaz \(2012\)M\. Cakmak and A\. L\. ThomazDesigning robot learners that ask good questions\.InProceedings of the Seventh Annual ACM/IEEE International Conference on Human\-Robot Interaction,pp\. 17–24\.External Links:[Document](https://dx.doi.org/10.1145/2157689.2157693)Cited by:[§2\.3](https://arxiv.org/html/2610.00282#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.4.1.1.1)\.
- Côtéet al\.\(2018\)M\. Côté, Á\. Kádár, X\. Yuan, B\. Kybartas, T\. Barnes, E\. Fine, J\. Moore, M\. Hausknecht, L\. El Asri, M\. Adada,et al\.TextWorld: a learning environment for text\-based games\.arXiv preprint arXiv:1806\.11532\.External Links:[Link](https://arxiv.org/abs/1806.11532)Cited by:[§4\.1](https://arxiv.org/html/2610.00282#S4.SS1.p1.1)\.
- Cuiet al\.\(2023\)Y\. Cui, S\. Karamcheti, R\. Palleti, N\. Shivakumar, P\. Liang, and D\. Sadigh“No, to the Right”—Online Language Corrections for Robotic Manipulation via Shared Autonomy\.External Links:2301\.02555Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p1.1),[§2\.1](https://arxiv.org/html/2610.00282#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.2.1.1.1)\.
- Deitset al\.\(2013\)R\. Deits, S\. Tellex, P\. Thaker, D\. Simeonov, T\. Kollar, and N\. RoyClarifying commands with information\-theoretic human\-robot dialog\.Journal of Human\-Robot Interaction2\(2\),pp\. 58–79\.External Links:[Document](https://dx.doi.org/10.5898/JHRI.2.2.Deits)Cited by:[§2\.3](https://arxiv.org/html/2610.00282#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.4.1.1.1)\.
- Dongreet al\.\(2025\)V\. Dongre, X\. Yang, E\. C\. Acikgoz, S\. Dey, G\. Tur, and D\. Hakkani\-TurReSpAct: harmonizing reasoning, speaking, and acting towards building large language model\-based conversational AI agents\.InProceedings of the 15th International Workshop on Spoken Dialogue Systems Technology,Bilbao, Spain,pp\. 72–102\.External Links:[Link](https://aclanthology.org/2025.iwsds-1.7/)Cited by:[§2\.3](https://arxiv.org/html/2610.00282#S2.SS3.p1.1)\.
- Huanget al\.\(2022\)W\. Huang, F\. Xia, T\. Xiao, H\. Chan, J\. Liang, P\. Florence, A\. Zeng, J\. Tompson, I\. Mordatch, Y\. Chebotar,et al\.Inner monologue: embodied reasoning through planning with language models\.External Links:2207\.05608Cited by:[§2\.2](https://arxiv.org/html/2610.00282#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.3.1.1.1)\.
- Liuet al\.\(2023\)H\. Liu, A\. Chen, Y\. Zhu, A\. Swaminathan, A\. Kolobov, and C\. ChengInteractive robot learning from verbal correction\.External Links:2310\.17555Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p1.1),[§2\.1](https://arxiv.org/html/2610.00282#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.2.1.1.1)\.
- Liuet al\.\(2022\)I\. Liu, X\. Yuan, M\. Côté, P\. Oudeyer, and A\. SchwingAsking for knowledge \(AFK\): training RL agents to query external knowledge using language\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 14073–14093\.External Links:[Link](https://proceedings.mlr.press/v162/liu22t.html)Cited by:[§2\.3](https://arxiv.org/html/2610.00282#S2.SS3.p1.1)\.
- Piet al\.\(2025\)R\. Pi, K\. Miao, L\. Peihang, R\. Liu, J\. Gao, J\. Zhang, and X\. ZhouPointing to a llama and call it a camel: on the sycophancy of multimodal large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 20166–20180\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1020)Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p2.1),[§2\.4](https://arxiv.org/html/2610.00282#S2.SS4.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.5.1.1.1)\.
- Rubaviciuset al\.\(2026\)R\. Rubavicius, P\. D\. Fagan, A\. Lascarides, and S\. RamamoorthySECURE: semantics\-aware embodied conversation under unawareness for lifelong robot learning\.InProceedings of the 4th Conference on Lifelong Learning Agents,Proceedings of Machine Learning Research, Vol\.330,pp\. 612–634\.External Links:[Link](https://proceedings.mlr.press/v330/rubavicius26a.html)Cited by:[§2\.3](https://arxiv.org/html/2610.00282#S2.SS3.p1.1)\.
- Seaborn and Yalçın \(2026\)K\. Seaborn and Ö\. N\. YalçınRobotic sycophancy: a scoping review\.InCompanion Proceedings of the 21st ACM/IEEE International Conference on Human\-Robot Interaction,pp\. 938–943\.External Links:[Document](https://dx.doi.org/10.1145/3776734.3794532)Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p2.1)\.
- Sharmaet al\.\(2022\)P\. Sharma, B\. Sundaralingam, V\. Blukis, C\. Paxton, T\. Hermans, A\. Torralba, J\. Andreas, and D\. FoxCorrecting robot plans with natural language feedback\.InRobotics: Science and Systems,External Links:[Link](https://arxiv.org/abs/2204.05186)Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p1.1),[§2\.1](https://arxiv.org/html/2610.00282#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.2.1.1.1)\.
- Shiet al\.\(2024\)L\. X\. Shi, Z\. Hu, T\. Z\. Zhao, A\. Sharma, K\. Pertsch, J\. Luo, S\. Levine, and C\. FinnYell at your robot: improving on\-the\-fly from language corrections\.External Links:2403\.12910Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p1.1),[§2\.1](https://arxiv.org/html/2610.00282#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.2.1.1.1)\.
- Shridharet al\.\(2021\)M\. Shridhar, X\. Yuan, M\. Côté, Y\. Bisk, A\. Trischler, and M\. HausknechtALFWorld: aligning text and embodied environments for interactive learning\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p3.1)\.
- Sinha \(2026\)D\. SinhaSycoBench\-600: measuring sycophancy and correction selectivity in LLM assistants\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 35278–35284\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1759),[Link](https://aclanthology.org/2026.findings-acl.1759/)Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p2.1),[§2\.4](https://arxiv.org/html/2610.00282#S2.SS4.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.5.1.1.1)\.
- Thieraufet al\.\(2024\)C\. Thierauf, R\. Thielstrom, B\. Oosterveld, W\. Becker, and M\. Scheutz“Do This Instead”—Robots That Adequately Respond to Corrected Instructions\.ACM Transactions on Human\-Robot Interaction13\(3\),pp\. 1–23\.External Links:[Document](https://dx.doi.org/10.1145/3623385)Cited by:[§1](https://arxiv.org/html/2610.00282#S1.p1.1),[§2\.1](https://arxiv.org/html/2610.00282#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2610.00282#S2.T1.2.2.1.1.1)\.

Similar Articles

Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents

arXiv cs.AI

This paper proposes ActionRating, a formulation that places clarification inside an agent's action space on a shared ordinal scale with navigation, enabling two information-seeking modes (mandatory and opportunistic). On hierarchical taxonomy classification benchmarks, experiments with 9 LLMs show that opportunistic clarification improves accuracy and information-seeking effectiveness.

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

Hugging Face Daily Papers

This paper proposes the Agent-Editing World Model (AEWM) to enhance LLM agents by modeling how reasoning and actions shape future task progress, achieving significant performance improvements on benchmarks across domains like search and software engineering.