video-language-models

Tag

Cards List
#video-language-models

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

Hugging Face Daily Papers · 4d ago Cached

MoTE introduces task-specific expert routing to replace dense decoder feed-forward networks in multi-task video understanding, improving accuracy and efficiency with interpretable, sparse computation.

0 favorites 0 likes
#video-language-models

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

Hugging Face Daily Papers · 2026-08-05 Cached

Introduces CLIP-CC-Bench, a benchmark for evaluating video-language models on paragraph-level descriptions of movie clips, with 200 clips and human reference paragraphs. Tests 17 models and finds that long-form video description remains unsolved.

0 favorites 0 likes
← Back to home

Submit Feedback