Tag
LatentPress introduces a method to compress conversational and document context into continuous memory tokens, enabling frozen decoders to read directly without text reconstruction, achieving higher compression ratios and improved performance on long-context tasks.
Proposes SOLAR, an auxiliary fine-tuning objective that aligns soft-token representations across languages to improve multilingual reasoning consistency, achieving up to +17.7 points accuracy gain.
SKIM is an adaptive multi-resolution soft token compression framework that compresses procedural skills for LLMs, maintaining task performance while reducing prefill cost and latency.