action-prediction

Tag

Cards List
#action-prediction

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

arXiv cs.AI · 2026-08-21 Cached

The paper introduces DECOWAM, a decoupled whole-body world-action model for legged mobile manipulation that improves video and action prediction performance over existing models like FastWAM through dedicated conditional interfaces and a new dataset.

0 favorites 0 likes
#action-prediction

Black Forest Lab's Flux 3: Omni-modality for image, video, audio & action prediction

Reddit r/singularity · 2026-07-23

Black Forest Lab's Flux 3 is a new omni-modal AI model capable of generating and predicting images, video, audio, and actions.

0 favorites 0 likes
#action-prediction

Speculate with Memory: Lossless Acceleration for LLM Agents

arXiv cs.LG · 2026-07-15 Cached

This paper introduces memory-augmented speculative execution for LLM agents, using three online memory systems to improve prediction accuracy by 19-39% on action prediction and up to 2.5x on observation prediction, all while being lossless with zero added wall-clock cost.

0 favorites 0 likes
#action-prediction

CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding

arXiv cs.AI · 2026-07-13 Cached

CLAP proposes a method to convert pretrained vision-language models (VLMs) into vision-language-action models (VLAs) by prepending natural-language action descriptions to action token sequences, preserving semantic capabilities without architectural changes. It achieves 90.8% on LIBERO and improves robustness.

0 favorites 0 likes
#action-prediction

An open model predicting a robot's actions from a control signal. The corner panels are the action and hand pose it was given, everything else is imagined. Is this a world model, or just a video generator?

Reddit r/singularity · 2026-07-12

An open model that predicts a robot's actions from a control signal, raising questions about whether it constitutes a world model or just a video generator.

0 favorites 0 likes
#action-prediction

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

Hugging Face Daily Papers · 2026-06-06 Cached

Light-WAM is a lightweight world action model for efficient robot manipulation that uses a compact video backbone and downsampled latent space for future-video supervision, achieving high performance with low inference latency.

0 favorites 0 likes
#action-prediction

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models

Hugging Face Daily Papers · 2026-06-01 Cached

RoboSemanticBench is a benchmark that diagnoses semantic grounding in action prediction for vision-language-action models, revealing that while robots can grasp objects, they fail to select semantically correct targets based on instruction semantics.

0 favorites 0 likes
#action-prediction

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark

arXiv cs.AI · 2026-05-29 Cached

This paper introduces the PiSAR benchmark for screen-conditioned action prediction and compares supervised fine-tuned models against frontier zero-shot baselines. Key findings show a fine-tuned Qwen3-VL-8B achieves 0.783 semantic similarity, significantly outperforming Claude Opus 4.7 and GPT-5.5 (0.459 and 0.482), but the same fine-tuning recipe on a larger reasoning-tuned Gemma model yields only 0.441, indicating a model-recipe mismatch.

0 favorites 0 likes
#action-prediction

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents

Hugging Face Daily Papers · 2026-05-18 Cached

MementoGUI introduces a plug-in agentic memory framework for GUI agents that uses learned controllers for selective memory management and retrieval, improving performance on long-horizon tasks with compressed visual and textual representations.

0 favorites 0 likes
← Back to home

Submit Feedback