@rohanpaul_ai: A lot of embodied AI still feels like AI modules bolted onto a robot. TARS is taking a different architectural bet with…
Summary
TARS launches AWE 3.5, an embodied-native foundation model that integrates action, perception, geometry, and touch into one model for general-purpose physical AI, with claims of 2x task execution efficiency over PI0.5.
View Cached Full Text
Cached at: 08/22/26, 11:23 AM
A lot of embodied AI still feels like AI modules bolted onto a robot.
TARS is taking a different architectural bet with AI World Engine (AWE) 3.5, TARS’ embodied-native foundation model for physical AI.
Its “Born as One” approach puts action, perception, geometry, and touch into one model from the beginning rather than stitching those capabilities together later.
The same model-driven system is designed to generalize across different tasks, objects, environments and robot bodies.
The training recipe then implements and validates a full closed-loop methodology for embodied-native foundation models through pre-training and post-training.
During pre-training, 2 priors give the model a base understanding of action patterns, spatial structure and understanding of physical laws before it is adapted to a robot, while post-training uses the AI World Engine to roll possible future states forward inside the model, predict what different actions may lead to and use those predictions to choose better actions.
TARS describes the full loop as 5 connected parts: embodied-native architecture, dual-prior pre-training, World Engine-driven post-training, scaling validation and continuous data feedback.
TARS positions AWE 3.5 as one of the most powerful embodied-native foundation models for general-purpose physical AI, with several minutes of long-horizon closed-loop reasoning and roughly 2x task execution efficiency versus PI0.5.
@TARSRobotics #AWE35 #TARS #tarsrobotics
Similar Articles
bytedance/UI-TARS-desktop
ByteDance released TARS, a multimodal AI agent stack comprising Agent TARS (a CLI/Web UI-based general AI agent for GUI, browser, and terminal tasks) and UI-TARS Desktop (a native desktop application powered by the UI-TARS model for local and remote computer/browser automation). The stack integrates multimodal LLMs with MCP tools for human-like task completion.
@rohanpaul_ai: Robotics is slow because every change needs physical setup, people, space, and repeated field runs. Physical AI needs t…
Antioch introduces Antioch Agent, a browser-based robotics simulator that lets developers test robot software in a closed agentic loop without physical hardware, accelerating development cycles.
@rohanpaul_ai: Language had a strange advantage robotics does not: Text is already a compressed, shared interface for human thought, w…
Discusses the challenges facing embodied AI and robotics, including a 100,000-year data gap and lack of shared benchmarks, and highlights startup opportunities in data loops, eval systems, and deployment.
@rohanpaul_ai: Thinking Machines is replacing turn-taking AI with always-present AI. They just announced TML-Interaction-Small, a 276B…
Thinking Machines announced TML-Interaction-Small, a 276B MoE model designed for real-time, always-on interaction with sub-0.4s latency and integrated multimodal processing.
tencent/HY-Embodied-0.5
Tencent releases HY-Embodied-0.5, a suite of foundation models designed for embodied AI agents featuring a Mixture-of-Transformers (MoT) architecture with efficient 2B and powerful 32B variants for real-world robot control and spatial-temporal reasoning.