Tag
MoTE introduces task-specific expert routing to replace dense decoder feed-forward networks in multi-task video understanding, improving accuracy and efficiency with interpretable, sparse computation.
Introduces CLIP-CC-Bench, a benchmark for evaluating video-language models on paragraph-level descriptions of movie clips, with 200 clips and human reference paragraphs. Tests 17 models and finds that long-form video description remains unsolved.