Tag
A 2B vision-language model distilled from a tool-using agent achieves first-place performance on long egocentric video question answering by pruning its multilingual embedding table to meet parameter limits, reaching 89% accuracy of the larger pipeline with only 1.1% of parameters.
ReMem introduces a dual-level memory-augmented keyframe selection framework for training-free long video understanding, achieving state-of-the-art zero-shot performance on multiple benchmarks.
This technical report presents Kwai Keye-VL-2.0, an open-source Mixture-of-Experts multimodal foundation model designed for long-video understanding and agentic intelligence, leveraging DeepSeek Sparse Attention and cross-modal distillation to achieve state-of-the-art performance among similar-scale models.
MemDreamer decouples perception and reasoning for long video understanding using hierarchical graph memory and agentic retrieval, achieving state-of-the-art performance with reduced computational overhead.