visual-tokens

Tag

Cards List
#visual-tokens

RenderRank: Learning to Rerank Text with Compressed Visual Tokens

Hugging Face Daily Papers ↗ · 3d ago Cached

RenderRank introduces a reranking method that uses compressed visual document representations to reduce token usage and improve efficiency while maintaining accuracy across multiple datasets.

0 favorites 0 likes
#visual-tokens

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

Hugging Face Daily Papers ↗ · 3d ago Cached

The paper proposes δ-Vision, a method to reduce computational overhead in multimodal language models by using lightweight low-rank adapters to reconstruct visual states efficiently without discarding visual tokens.

0 favorites 0 likes
#visual-tokens

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

arXiv cs.AI ↗ · 2026-09-16 Cached

VideoMM introduces an adaptive macro-micro inference framework that reduces visual token overhead in video MLLMs, achieving a 6.13× speedup and 7.4% accuracy gain for efficient long-form video understanding.

0 favorites 0 likes
#visual-tokens

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

CoVeR is a training-free spatial token selector that improves 3D reasoning in Vision-Language Models by enforcing exact token budgets and full scene coverage, outperforming prior state-of-the-art methods.

0 favorites 0 likes
#visual-tokens

A 460M VLM gets first-token latency down to 0.3s on an iPhone by using only 64 visual tokens

Reddit r/LocalLLaMA ↗ · 2026-08-05

VisionPsy-Nano-460M-Flash is a new 460M vision-language model that uses only 64 visual tokens per image, cutting first-token latency to 0.3s on iPhone while retaining roughly 99% of the full model's benchmark score. The optimization trades off some OCR and fine-detail quality.

0 favorites 0 likes
#visual-tokens

Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM

arXiv cs.CL ↗ · 2026-06-04 Cached

This paper proposes a query-based cross-modal projector that compresses visual tokens via cross-attention to improve Mamba-based multimodal LLMs, boosting both performance and throughput on vision-language benchmarks while eliminating the need for manual 2D scan order design.

0 favorites 0 likes
#visual-tokens

AdaCodec: A Predictive Visual Code for Video MLLMs

Hugging Face Daily Papers ↗ · 2026-06-01 Cached

AdaCodec reduces video encoding redundancy in multimodal LLMs by transmitting full visual tokens only when scene prediction fails, otherwise using compact inter-frame change descriptions. It outperforms per-frame RGB baselines at matched token budgets and achieves better or comparable results with significantly fewer tokens, reducing time-to-first-token from 9.26s to 1.62s.

0 favorites 0 likes
← Back to home

Submit Feedback