Tag
The paper introduces GAVA, a framework that arbitrates whether an embodied agent should accept, reject, inspect the world, or ask for clarification when given a user correction, evaluated in text-only ALFWorld with training-derived object-location priors that reduce interaction and joint cost against uniform baselines.
The paper introduces MATE, a deterministic post-retrieval memory adaptation procedure that converts retrieved trajectories into execution-oriented memory for embodied agents, achieving 81.3% and 93.3% success on 134 ALFWorld tasks with Qwen2.5-14B and 72B while using about one-tenth the tokens of raw trajectories.
SciHorizon-eLab is an agentic protocol-to-task compiler that compiles scientific protocols into embodied tasks for scalable benchmarking, introducing a benchmark of 300 certified tasks.
RoboFoundry proposes a self-evolving system-as-policy framework for embodied agents, achieving state-of-the-art performance on benchmarks and enabling zero-shot transfer in real-world deployments.
RoboFollow introduces a diagnostic benchmark to expose the illusion of instruction-following in embodied agents by analyzing high scene entropy and perturbations, revealing gaps in current models despite strong initial performance.
OmniEcho introduces a spatially aware omni-modal model and a new benchmark for spatial audio-visual perception and navigation in embodied agents, achieving state-of-the-art performance.
This paper formalizes behavioural knowledge as a precondition Bayesian network and explores three placement strategies in neuro-symbolic RL for embodied agents, demonstrating improvements in solution quality and sample efficiency on benchmarks like MiniGrid and Fetch.
This paper proposes an audit protocol for update admission in continual embodied agents, emphasizing the need to balance error control with retained learning opportunities to prevent harmful updates and enable useful learning.
The project lets users prompt AI agents to play 1v1 football in a physics-based simulation, evaluating how well prompts translate to actions in embodied environments.
SafeBranch is a framework that aligns embodied agents to act safely using branch pairs from unsafe rollouts, significantly improving safety in interactive tasks without sacrificing task success.
This paper argues that representing agent skills as code, rather than natural language, yields the best cost reduction for LLM agents, and introduces SpeedRunner, a coding agent that learns programmatic skills from past trajectories to improve performance while cutting costs across embodied environments.
Presents Thea, a harness for embodied agents that orchestrates robot capabilities as callable tools, introducing Scene Graph as Context and Evaluation as Exit Codes to close the loop between agent and physical world.
This paper introduces Spatial Memory Agent (SMA), a runtime framework that improves frozen vision-language models' spatial reasoning through verifier-guided reflection and reusable memory without parameter updates or external tools, achieving strong results across five benchmarks and four base VLMs.
MetaSpace is a framework that applies metamorphic testing to evaluate spatial cognition in embodied agents, generating test cases from execution trajectories and encoding logical/physical constraints as Prolog rules. Benchmarking shows state-of-the-art MLLM-driven agents score far below human levels on spatial cognition, with directional tasks being especially weak.
Introduces a new paradigm called Combodied Agents that unify digital and embodied AI agents to model, predict, and support individual human-state trajectories over time, focusing on sustained human benefit rather than task completion.
This paper proposes CEAA, a modular cognitive architecture for embodied intelligent virtual agents that integrates high-level reasoning models with real-time embodied execution in interactive 3D environments.
360CityArena is a photorealistic urban navigation benchmark built from 360-degree videos of Tokyo's Akihabara district. It evaluates embodied agents' environment understanding, path reasoning, and spatial reasoning, showing that even strong LMM-based agents like Gemini 2.5 Flash perform far below human level.
A survey paper introducing a functional role taxonomy for language grounding in embodied agents, distinguishing five roles and auditing evidence to assess whether language's contribution is genuinely supported.
SpatialCLI presents a framework that trains vision-language models to use specialist spatial tools and then internalize those capabilities, boosting Qwen3-VL-8B-Instruct on the MindCube benchmark from 29.3% to 84.6% with tools.
This paper organizes embodied data sources into a five-level pyramid (real-robot, UMI, egocentric/exocentric, simulation, general vision-language), analyzing their trade-offs between scalability and robot alignment, and reviews recent embodied foundation models in terms of data recipes. It also discusses open challenges for building next-generation embodied systems.