@rohanpaul_ai: A lot of embodied AI still feels like AI modules bolted onto a robot. TARS is taking a different architectural bet with…

X AI KOLs Timeline Models

Summary

TARS launches AWE 3.5, an embodied-native foundation model that integrates action, perception, geometry, and touch into one model for general-purpose physical AI, with claims of 2x task execution efficiency over PI0.5.

A lot of embodied AI still feels like AI modules bolted onto a robot. TARS is taking a different architectural bet with AI World Engine (AWE) 3.5, TARS’ embodied-native foundation model for physical AI. Its "Born as One" approach puts action, perception, geometry, and touch into one model from the beginning rather than stitching those capabilities together later. The same model-driven system is designed to generalize across different tasks, objects, environments and robot bodies. The training recipe then implements and validates a full closed-loop methodology for embodied-native foundation models through pre-training and post-training. During pre-training, 2 priors give the model a base understanding of action patterns, spatial structure and understanding of physical laws before it is adapted to a robot, while post-training uses the AI World Engine to roll possible future states forward inside the model, predict what different actions may lead to and use those predictions to choose better actions. TARS describes the full loop as 5 connected parts: embodied-native architecture, dual-prior pre-training, World Engine-driven post-training, scaling validation and continuous data feedback. TARS positions AWE 3.5 as one of the most powerful embodied-native foundation models for general-purpose physical AI, with several minutes of long-horizon closed-loop reasoning and roughly 2x task execution efficiency versus PI0.5. @TARSRobotics #AWE35 #TARS #tarsrobotics 1.
Original Article
View Cached Full Text

Cached at: 08/22/26, 11:23 AM

A lot of embodied AI still feels like AI modules bolted onto a robot.

TARS is taking a different architectural bet with AI World Engine (AWE) 3.5, TARS’ embodied-native foundation model for physical AI.

Its “Born as One” approach puts action, perception, geometry, and touch into one model from the beginning rather than stitching those capabilities together later.

The same model-driven system is designed to generalize across different tasks, objects, environments and robot bodies.

The training recipe then implements and validates a full closed-loop methodology for embodied-native foundation models through pre-training and post-training.

During pre-training, 2 priors give the model a base understanding of action patterns, spatial structure and understanding of physical laws before it is adapted to a robot, while post-training uses the AI World Engine to roll possible future states forward inside the model, predict what different actions may lead to and use those predictions to choose better actions.

TARS describes the full loop as 5 connected parts: embodied-native architecture, dual-prior pre-training, World Engine-driven post-training, scaling validation and continuous data feedback.

TARS positions AWE 3.5 as one of the most powerful embodied-native foundation models for general-purpose physical AI, with several minutes of long-horizon closed-loop reasoning and roughly 2x task execution efficiency versus PI0.5.

@TARSRobotics #AWE35 #TARS #tarsrobotics

Similar Articles

bytedance/UI-TARS-desktop

GitHub Trending (daily)

ByteDance released TARS, a multimodal AI agent stack comprising Agent TARS (a CLI/Web UI-based general AI agent for GUI, browser, and terminal tasks) and UI-TARS Desktop (a native desktop application powered by the UI-TARS model for local and remote computer/browser automation). The stack integrates multimodal LLMs with MCP tools for human-like task completion.

tencent/HY-Embodied-0.5

Hugging Face Models Trending

Tencent releases HY-Embodied-0.5, a suite of foundation models designed for embodied AI agents featuring a Mixture-of-Transformers (MoT) architecture with efficient 2B and powerful 32B variants for real-world robot control and spatial-temporal reasoning.