Tag
This paper introduces GEAR, a Geometry-Enabled Attention Routing framework that uses geometry as explicit token-level addresses for visual memory to enhance long-horizon camera-controlled video generation, achieving state-of-the-art results.
This paper proposes VisKG-LM, a method that compiles knowledge graphs into visual memory for efficient multiple-choice question answering, achieving performance gains over baselines by decoupling graph encoding from language reasoning.
A tweet praising CollovLabs' advancement from single-shot recognition to persistent visual memory in AI.
Introduces DMV-Bench, an interactive benchmark for evaluating visual memory in multimodal agents using incidental visual cues from product images, and proposes DualMem, a dual-coding memory architecture that outperforms text-only and other multimodal baselines across various chain lengths.