Tag
A new paper proposes sequential KV cache compression using probabilistic language tries and predictive delta coding, achieving theoretical compression ratios of ~914,000× beyond TurboQuant by exploiting the sequential structure of language model tokens rather than treating vectors independently.
A reflective essay on how lossy compression creates visible artifacts that influence digital aesthetics, forensics, and art.
This paper presents a systematic benchmark of token pruning—a compression technique that removes tokens and embeddings for irrelevant languages—applied to Korean-centric LLM tasks. The study evaluates popular multilingual models (Qwen3, Gemma-3, Llama-3, Aya) across different vocabulary configurations and finds that token pruning significantly improves generation stability and reduces memory footprint for domain-specific deployments.
Anthropic, Alibaba, Google and others unleash a wave of major model drops—Claude Opus 4.7, Qwen 3.6, emotion-rich Google TTS, plus tiny 1.58-bit phone LLMs and real-time 3-D world generators—alongside open tools for video, VR and character creation.