Tag
VisCo is a training-efficient self-compression framework that reuses a pretrained vision-language model as an intrinsic encoder for visual token compression, achieving superior performance across all compression ratios without external modules.
PARCEL introduces a novel vision-language model architecture that uses pool-anchored resampling and conditioned elastic queries to improve efficiency and performance across different visual-token budgets, outperforming existing matryoshka baselines.
AQuaUI is a training-free inference-time token reduction method for GUI agent models that uses adaptive quadtrees to reduce spatial redundancy in screenshots, achieving up to 13.22% speedup and 29.52% fewer visual tokens while retaining 99.06% of performance.