vision-language-action

Tag

Cards List
#vision-language-action

@svlevine: How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large r…

X AI KOLs Following · yesterday Cached

A research collaboration introduces ARLI, a method to enable RL fine-tuning of Vision-Language-Action models with asynchronous inference using real-time chunking, improving robot behaviors by allowing the RL policy to observe more recent images.

0 favorites 0 likes
#vision-language-action

@askalphaxiv: “Reinforcement Learning for Real-Time Vision-Language-Action Policies” VLA models are usually too slow for reactive rob…

X AI KOLs Following · 2d ago Cached

This paper introduces a method to split action generation in vision-language-action models into slow and fast layers, enabling real-time reactive robot control and improving success rates from 42% to 97% with just 10 minutes of online data.

0 favorites 0 likes
#vision-language-action

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

Hugging Face Daily Papers · 2d ago Cached

The paper introduces THAW-VLA, a method that distills world-model representations into Vision-Language-Action models for robotics, enhancing robustness and performance on simulation and real hardware without additional inference overhead.

0 favorites 0 likes
#vision-language-action

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Hugging Face Daily Papers · 2d ago Cached

CARE is a framework for Vision-Language-Action policies that enhances robotic manipulation by learning from execution failures to generate corrective actions, demonstrating improved success rates in simulations and real-world tasks.

0 favorites 0 likes
#vision-language-action

@Official_m9g: Won Gemini 3 Seoul Hackathon (Hard Tech) with this: Photos → graph → text layout → map → waypoints. Perspective images …

X AI KOLs Timeline · 3d ago Cached

GeminiSpace, an autonomous indoor navigation system using Google Gemini, won first place at the Gemini 3 Seoul Hackathon by converting panoramic photos into navigable maps and robotic trajectories.

0 favorites 0 likes
#vision-language-action

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Hugging Face Daily Papers · 5d ago Cached

This paper presents HuRo, a pipeline for robotizing human videos to create scalable VLA pretraining data, showing significant improvements in task completion and robustness on real-world manipulation tasks.

0 favorites 0 likes
#vision-language-action

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Hugging Face Daily Papers · 2026-09-16 Cached

This paper introduces ActionPiece, a novel action tokenization method for autoregressive vision-language-action models that uses physical rank consistency to improve the fidelity of action relationships, evaluated on benchmarks like LIBERO and SimplerEnv.

0 favorites 0 likes
#vision-language-action

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

arXiv cs.AI · 2026-09-11 Cached

This paper proposes Time–Frequency Geometric Cross-Attention (TFGCA), a drop-in module for chunked vision-language-action models that improves action trajectory prediction by decomposing chunks into time-frequency representations and capturing geometric relationships, resulting in significant performance gains on benchmarks and real-robot tasks.

0 favorites 0 likes
#vision-language-action

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Hugging Face Daily Papers · 2026-09-11 Cached

Dynin-Robotics is an omnimodal unified diffusion model that integrates vision, language, and action for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.

0 favorites 0 likes
#vision-language-action

LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies

arXiv cs.AI · 2026-09-10 Cached

LayerRoute introduces an action-conditioned routing interface for Vision-Language-Action policies that dynamically adapts access to VLM layer representations, improving robot manipulation performance with minimal additional parameters.

0 favorites 0 likes
#vision-language-action

Machines that think: embodied intelligence (10 minute read)

TLDR AI · 2026-09-08 Cached

The article critiques the oversimplified view that robotics is merely an integration problem after AI breakthroughs, highlighting challenges in training data scarcity and generalization for vision-language-action models. It contrasts these limitations with the reliability of specialized industrial robots achieved through traditional automation.

0 favorites 0 likes
#vision-language-action

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

Hugging Face Daily Papers · 2026-09-07

This study explores adding Greek to a Cosmos3 vision-language-action robot policy, revealing challenges in measurement and showing that bilingual training improves performance but overfits to translator phrasing. The results emphasize the need for null baselines and seed replication in low-resource localization.

0 favorites 0 likes
#vision-language-action

MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

Hugging Face Daily Papers · 2026-09-05 Cached

MobileVLA-R1 2.0 is an RL-enhanced vision-language-action framework that couples structured reasoning with mobile robot control, achieving improvements in long-horizon instruction following on real-world platforms.

0 favorites 0 likes
#vision-language-action

ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models

Hugging Face Daily Papers · 2026-09-02 Cached

ShieldVLA proposes a safety-aligned fine-tuning framework for Vision-Language-Action models using Hamilton-Jacobi reachability to estimate safe operating regions, reducing safety costs by 57% and improving task success in robotics benchmarks.

0 favorites 0 likes
#vision-language-action

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Hugging Face Daily Papers · 2026-09-02 Cached

SimpleMemVLA introduces a simple memory mechanism for Vision-Language-Action models by feeding intact timestamped video history into a pretrained VLM backbone, achieving state-of-the-art results on long-horizon manipulation tasks without dedicated memory modules.

0 favorites 0 likes
#vision-language-action

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Hugging Face Daily Papers · 2026-08-26 Cached

StreamPI introduces a streaming multimodal temporal modeling framework for vision-language-action models, improving robot manipulation through instruction-anchored attention and randomized interval training without additional parameters.

0 favorites 0 likes
#vision-language-action

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Hugging Face Daily Papers · 2026-08-26 Cached

MA-VLA is a unified framework for multi-arm robot collaboration that decomposes cooperative behavior into atomic prompts and uses training-time permutations (Arm Shuffle) to enable generalization to unseen coordination patterns.

0 favorites 0 likes
#vision-language-action

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

arXiv cs.AI · 2026-08-24 Cached

ForeTime-VLA introduces a causal future-token distillation method from a world action model to improve dynamic conveyor-belt manipulation, achieving higher grasp success rates compared to existing VLA policies.

0 favorites 0 likes
#vision-language-action

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Hugging Face Daily Papers · 2026-08-24 Cached

This paper proposes Intention Distillation (INDI) to distill behavior intent into the action decoder of Vision-Language-Action models, improving performance on benchmarks like SimplerEnv-Bridge and real-world tasks.

0 favorites 0 likes
#vision-language-action

Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies

arXiv cs.LG · 2026-08-20 Cached

This paper introduces Role-Conditioned Sub-Token Routing (RoleSub), a method to efficiently compress vision-language-action models by routing sub-token groups, reducing computational costs while maintaining strong performance on robotic tasks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback