Tag
Introduces SimLife, a scalable platform for simulating long-term household life, and SimLife-BP, a benchmark to evaluate AI agents' ability to infer behavioral patterns from extended observations, finding that current models often lack deep rule-based understanding.
DeliveryGym is a 3D reinforcement learning environment for long-horizon embodied agent planning with adaptive curriculum, demonstrating performance improvements through RL.
VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.
This paper introduces GPT-Policy, a framework for in-context robot learning using vision-language models, enabling robots to learn from demonstrations without gradient updates. It evaluates the framework in real-robot trials, showing improved task completion.
PhysBrain 1.5 is an open-source 8B parameter model for embodied AI that achieves state-of-the-art performance by understanding scenes, generating robot actions, and predicting future frames, scoring 72.5 on 28 benchmarks.
OM-1 is a fast embodied AI system designed for rapid performance in artificial intelligence applications related to physical environments.
PhysBrain 1.5 is a unified model that integrates physical environment understanding, action generation, and future state prediction via autoregressive training, achieving state-of-the-art open-source performance on 28 embodied benchmarks.
The article highlights the rapid progress in humanoid robotics from 2015 to 2026, showcasing advanced capabilities like running and industrial tasks, and speculates on future developments by 2036.
The article discusses the GPT-6 Astra model's capability to perform in-context learning in physical world environments, representing a significant advancement in embodied AI.
VSArena v0.6.0 is a major update to a browser-based studio for running and evaluating Vision-Language-Action policies in 3D physics environments, emphasizing reliable and reproducible assessment without local simulators or physical robots.
Xiaomi-Robotics-U0-4B is a compact 4B model for unified embodied world generation, ranked #1 on World Arena and significantly improving real-world success rates with open-source code.
This paper introduces Counterfactual Latent World Models (CLWM) to address counterfactual collapse in world models, improving planning success in embodied reasoning under partial observability through contrastive objectives.
Xiaomi Robotics has released a compact embodied world model, including a 4B version and sequence variants, on Hugging Face with Apache 2.0 license.
GPT-6 demonstrated the ability to autonomously spin up a VM, set up MuJoCo, and recreate a scene from video for embodied AI research, highlighting advancements in real-to-sim-to-real workflows.
ReactHuman is a physics-grounded benchmark that evaluates multimodal large language models for human-like reactive decision-making in simulated humanoid robots facing household hazards, revealing persistent safety failures and offering a scalable diagnostic tool.
Show-Harness is a method that enables vision-language models to control robots through discrete semantic actions, allowing zero-shot deployment and efficient fine-tuning across different robots and GUIs.
An AI model named Astra was given a robot, paintbrush, and camera to paint the Golden Gate Bridge in real life, demonstrating its ability to learn and improve over multiple attempts.
A video demonstrates an AI system called "GPT Astra" using chain-of-thought reasoning to direct a robot to clean a table.
The paper proposes a Risk-Informed World Model (RIWM) to enhance safety in critical embodied systems by prioritizing decision-relevant consequences and recovery over predictive likelihood.
ZDTaichu5.0-9B is a multimodal foundation model that combines a Qwen3.5-9B language backbone with a C-RADIOv4-H vision encoder, excelling in general visual understanding, spatial reasoning, and agent tasks among 10B-scale VLMs.