omnimodal-llm

Tag

Cards List
#omnimodal-llm

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

Hugging Face Daily Papers · 2026-07-28 Cached

OmniScope is a training-free token compression framework for omnimodal LLMs that estimates audio and video relevance separately using the query as a shared anchor, achieving up to 3.53x prefill speedup and over 15% GPU memory reduction with minimal accuracy loss.

0 favorites 0 likes
← Back to home

Submit Feedback