Tag
A research collaboration introduces ARLI, a method to enable RL fine-tuning of Vision-Language-Action models with asynchronous inference using real-time chunking, improving robot behaviors by allowing the RL policy to observe more recent images.
This paper introduces a method to split action generation in vision-language-action models into slow and fast layers, enabling real-time reactive robot control and improving success rates from 42% to 97% with just 10 minutes of online data.
The paper introduces THAW-VLA, a method that distills world-model representations into Vision-Language-Action models for robotics, enhancing robustness and performance on simulation and real hardware without additional inference overhead.
CARE is a framework for Vision-Language-Action policies that enhances robotic manipulation by learning from execution failures to generate corrective actions, demonstrating improved success rates in simulations and real-world tasks.
GeminiSpace, an autonomous indoor navigation system using Google Gemini, won first place at the Gemini 3 Seoul Hackathon by converting panoramic photos into navigable maps and robotic trajectories.
This paper presents HuRo, a pipeline for robotizing human videos to create scalable VLA pretraining data, showing significant improvements in task completion and robustness on real-world manipulation tasks.
This paper introduces ActionPiece, a novel action tokenization method for autoregressive vision-language-action models that uses physical rank consistency to improve the fidelity of action relationships, evaluated on benchmarks like LIBERO and SimplerEnv.
This paper proposes Time–Frequency Geometric Cross-Attention (TFGCA), a drop-in module for chunked vision-language-action models that improves action trajectory prediction by decomposing chunks into time-frequency representations and capturing geometric relationships, resulting in significant performance gains on benchmarks and real-robot tasks.
Dynin-Robotics is an omnimodal unified diffusion model that integrates vision, language, and action for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.
LayerRoute introduces an action-conditioned routing interface for Vision-Language-Action policies that dynamically adapts access to VLM layer representations, improving robot manipulation performance with minimal additional parameters.
The article critiques the oversimplified view that robotics is merely an integration problem after AI breakthroughs, highlighting challenges in training data scarcity and generalization for vision-language-action models. It contrasts these limitations with the reliability of specialized industrial robots achieved through traditional automation.
This study explores adding Greek to a Cosmos3 vision-language-action robot policy, revealing challenges in measurement and showing that bilingual training improves performance but overfits to translator phrasing. The results emphasize the need for null baselines and seed replication in low-resource localization.
MobileVLA-R1 2.0 is an RL-enhanced vision-language-action framework that couples structured reasoning with mobile robot control, achieving improvements in long-horizon instruction following on real-world platforms.
ShieldVLA proposes a safety-aligned fine-tuning framework for Vision-Language-Action models using Hamilton-Jacobi reachability to estimate safe operating regions, reducing safety costs by 57% and improving task success in robotics benchmarks.
SimpleMemVLA introduces a simple memory mechanism for Vision-Language-Action models by feeding intact timestamped video history into a pretrained VLM backbone, achieving state-of-the-art results on long-horizon manipulation tasks without dedicated memory modules.
StreamPI introduces a streaming multimodal temporal modeling framework for vision-language-action models, improving robot manipulation through instruction-anchored attention and randomized interval training without additional parameters.
MA-VLA is a unified framework for multi-arm robot collaboration that decomposes cooperative behavior into atomic prompts and uses training-time permutations (Arm Shuffle) to enable generalization to unseen coordination patterns.
ForeTime-VLA introduces a causal future-token distillation method from a world action model to improve dynamic conveyor-belt manipulation, achieving higher grasp success rates compared to existing VLA policies.
This paper proposes Intention Distillation (INDI) to distill behavior intent into the action decoder of Vision-Language-Action models, improving performance on benchmarks like SimplerEnv-Bridge and real-world tasks.
This paper introduces Role-Conditioned Sub-Token Routing (RoleSub), a method to efficiently compress vision-language-action models by routing sub-token groups, reducing computational costs while maintaining strong performance on robotic tasks.