Tag
SCOUT introduces a recovery-aware agentic framework with an adaptive exploration-exploitation policy for ultra-long egocentric video reasoning, trained via UPS-GRPO, a uncertainty-prioritized RL method with turn-level advantage decomposition. It achieves state-of-the-art results on ultra-long egocentric benchmarks and remains competitive on shorter long-video settings.
Presents SERUM, a multi-pass framework that extracts structured behavioral models of user actions and intents from raw egocentric video using hierarchical VLM annotation, reducing hallucinations and producing interpretable process models without manual annotation.
ReViV is a unified framework for holistic egocentric 4D reconstruction that simultaneously reconstructs viewer (body, hand, gaze) and view (depth, camera trajectory) dynamics from a single monocular RGB video using a Masked Generative Egocentric Transformer, achieving state-of-the-art accuracy and efficiency.
Vinci2 presents a proactive assistance system for continuous egocentric videos, introducing a new benchmark EgoServe and a memory-augmented agent EgoMemo that determines when to intervene based on temporal context.
Reka AI Labs releases CS2-10k, a large dataset of over 600,000 egocentric gameplay videos with 10,000+ hours of frame-level keyboard, mouse, and 3D position data, available on Hugging Face for world models, action-conditioned video generation, and egocentric navigation research.
This paper finds that egocentric human video, when processed with a filtering and labeling pipeline, can outperform teleoperated real-robot data for pretraining embodied foundation models, achieving lower validation loss and higher success rates on real-robot tasks.
ACE-EGO-0 is a unified Vision-Language-Action pretraining framework that leverages egocentric human videos and robot trajectories via a reliability-aware training objective, achieving state-of-the-art on embodied AI benchmarks.
EgoPhys introduces a framework to construct deformable physical digital twins from egocentric RGB video using generalizable priors and a compact codebook, enabling zero-shot generalization to unseen objects without per-spring optimization. The system is demonstrated on a real robot, showing that egocentric human play video can serve as internal world representation for deformable-object planning.
This paper introduces V-RAGBench, a benchmark for evaluating retrieval-augmented generation over long egocentric videos, and CARVE, a method that adaptively selects retrieval configurations per chunk to improve VideoRAG performance.
A training-free framework for spatial reasoning from egocentric videos that enables revisiting conclusions through synthesized novel-view videos generated from predicted 3D geometry.
ActiveMimic is a pretraining framework that recovers camera and wrist trajectories from egocentric human video to model active perception as a viewpoint action, enabling robot pretraining that matches the performance of models trained directly on robot data.
SuperMemory-VQA is a new egocentric VQA benchmark featuring 52.9 hours of AI-glasses footage and 4,853 QA pairs designed to evaluate AI assistants on long-horizon memory tasks spanning object recall, intent, timelines, and conversations. Benchmarking reveals existing agentic frameworks and LLMs remain far from reliable on these real-world memory challenges.
Human Archive, a Silicon Valley startup, has raised $8.2 million to collect first-person video data from Indian gig workers to train robots for physical tasks. Despite rejections from major home services platforms, the company is partnering with other firms in the sector.
This paper investigates using vision-language models to assess nursing competency from egocentric video during simulation, finding that recognition accuracy inversely relates to competency level, suggesting a pedagogically informative signal.
Ego2World converts egocentric cooking videos (HD-EPIC) into executable symbolic worlds with graph-transition rules, enabling evaluation of belief-state planning under partial observation. Experiments show that belief memory improves task completion, suggesting it should be a first-class target in embodied agent evaluation.
PhysBrain 1.0 is a technical report presenting a method that uses human egocentric video to generate physical commonsense supervision for vision-language-action models, achieving state-of-the-art results on embodied control benchmarks including ERQA, PhysBench, SimplerEnv-WidowX, LIBERO, and RoboCasa.