Tag
The article discusses how AI is used in bot farms to create fake social media accounts that generate hyper-realistic comments and viral posts, potentially influencing online debates.
The article lists multiple ways to earn income in robotics through selling systems, services, artifacts, and owning verticals, highlighting the industry's reliance on practical engineering skills over pure AI development.
TacForcing is a streaming action-generation framework that integrates real-time tactile feedback during execution for contact-rich manipulation tasks, achieving higher success rates than existing methods in simulated and real-world settings.
ForeTime-VLA introduces a causal future-token distillation method from a world action model to improve dynamic conveyor-belt manipulation, achieving higher grasp success rates compared to existing VLA policies.
A hierarchical vision-language-action model improves long-horizon robot manipulation by using world-model-guided test-time search to scale computation for high-level subtask decisions.
H2R-Bench is a new benchmark that evaluates video world models on transforming human manipulation videos into robot-centric demonstrations, testing embodiment constraints and interaction fidelity across six manipulation families and two robot embodiments.
Dyna Robotics introduces Dyna-2, a world-action model pre-trained on one million hours of human video, discovering new scaling laws for robot manipulation.
RynnValue introduces an open-source robotic value foundation model that uses temporal distance as a scalable supervision target for reward learning, surpassing preference-supervised state-of-the-art on RBM-EVAL-OOD and improving real-world policy success rates.
Argues that the most valuable robotics acquisition target is the company owning the best real-world manipulation dataset, as models converge but data remains scarce and concentrated among a few players.
Ego2Robot is a scalable pipeline that converts egocentric human manipulation videos into robot training data via action retargeting and visual synthesis, producing 18,561 hours of data across 15 robot morphologies. Experiments show that joint pretraining on this synthesized data improves out-of-distribution generalization for vision-language-action models, including on real-robot deployment.
Push-Wiper is a robotics framework that reformulates viscous stain cleaning as an aggregation problem, using segmented pushing trajectories and a Diffusion Policy to generalize across stains, surfaces, and geometries, achieving up to 130% higher cleaning scores and zero-shot transfer.
ST-WAM proposes a semantic-temporal world action model that uses DINOv3 features as a shared semantic representation to improve robot manipulation robustness under visual distribution shifts, achieving 98.7% on LIBERO and 92.8% on RoboTwin 2.0, with significant gains over Fast-WAM in zero-shot settings.
Introduces WCM, a World Critic Model that jointly predicts future latent states and estimates values to improve temporal modeling for Vision-Language-Action reinforcement learning, achieving state-of-the-art results across robotic manipulation benchmarks.
SCOUT introduces per-context reset curricula for sparse-reward reinforcement learning, adapting assistance removal to each context's learning progress, outperforming global pacing methods in navigation and manipulation tasks.
A tweet highlighting that benchmarking robot policies is broken and sharing results from thousands of evaluations over 12 manipulation tasks to determine which policy to use.
HiFi-UMI introduces a portable data-production system for robot-free UMI data that achieves high trajectory accuracy using stereo-inertial SLAM and wide-angle cameras. Training manipulation policies on this data alone enables zero-shot deployment on real robots, matching or exceeding teleoperation baselines across several model families, and the authors open-source a 2,000-hour high-fidelity dataset.
N₀-TWAM is a tactile-native world-action model for contact-rich manipulation, trained at scale on visuo-tactile data from 6 embodiments and 450 tasks. The authors release code and pretrained checkpoints, positioning it as the first tactile world-action model trained at scale.
Introduces N_0-VTLA, a vision-tactile-language-action foundation model for contact-rich manipulation, featuring large-scale tactile pretraining and advantage-conditioned offline policy improvement, with strong results on real-robot and simulation benchmarks.
MiniCPM launches a new series of robot models: MiniCPM-RobotManip and MiniCPM-RobotTrack, focusing on manipulation and tracking tasks respectively.
Xiaomi Robotics-1, a robot foundation model trained on 100,000 hours of real-world manipulation data, has been released on Hugging Face. The model can autonomously perform household tasks like folding laundry, loading a washer, and washing dishes.