long-video-language-models

Tag

Cards List
#long-video-language-models

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

Hugging Face Daily Papers · 3d ago Cached

This paper presents a controlled study on visual-token allocation for long-video multimodal language models, finding that frame selection significantly drives accuracy, while spatial compression is nearly free when savings are reinvested into more frames, highlighting the need for a unified comparison harness.

0 favorites 0 likes
← Back to home

Submit Feedback