Tag
TempCloze is a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs, requiring models to identify missing video segments from distractors to reduce linguistic shortcuts in semantics, alignment, and progression.
LiteFrame proposes a lightweight video encoder with Compressed Token Distillation training that reduces latency and enables processing 8x more frames for long-form video understanding in Video LLMs, improving accuracy while reducing compute.