video-captioning

Tag

Cards List
#video-captioning

Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders

arXiv cs.CL ↗ · 6d ago Cached

The paper introduces ZENDAYA, a closed-form bandwidth dial for streaming multimodal decoders that dynamically adjusts input reading to improve real-time text generation performance. It demonstrates that reducing input consumption can enhance quality and efficiency in streaming settings across video and audio benchmarks.

0 favorites 0 likes
#video-captioning

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

Introduces multi-reference image-grounded video captioning and proposes RefCaptioner, a two-stage post-training framework with mixed-data SFT and hierarchical coverage-discounted GRPO. The paper also presents MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding.

0 favorites 0 likes
#video-captioning

PEEK: Picking Essential frames via Efficient Knowledge distillation

Hugging Face Daily Papers ↗ · 2026-05-29 Cached

Introduces PEEK, an efficient dynamic frame sampling method that distills caption-conditioned frame relevance rankings from a teacher model into a lightweight temporal model, outperforming state-of-the-art methods in video captioning while maintaining computational efficiency.

0 favorites 0 likes
#video-captioning

NemoStation/Marlin-2B

Hugging Face Models Trending ↗ · 2026-05-13

NemoStation/Marlin-2B is a fine-tuned model based on Qwen3.5-2B for video-text-to-text tasks, supporting video captioning and temporal grounding.

0 favorites 0 likes
← Back to home

Submit Feedback