video-mllm

Tag

Cards List
#video-mllm

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Papers with Code Trending · 2026-07-16 Cached

VideoChat3 is a fully open, efficient, and generalist video-centric multimodal large language model that introduces Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for streaming video perception, along with scalable video data synthesis pipelines, achieving superior performance with only 4B parameters.

0 favorites 0 likes
#video-mllm

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

arXiv cs.CL · 2026-04-20 Cached

RefereeBench introduces the first large-scale benchmark with 925 curated sports videos and 6,475 QA pairs to evaluate whether video MLLMs can reliably act as multi-sport referees. Evaluation of state-of-the-art models shows current MLLMs fall short (≤60% accuracy), struggling with rule application and temporal grounding despite their generic video understanding capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback