compression

Tag

Cards List
#compression

Latent Space as a New Medium

Hacker News Top · 2026-07-13 Cached

Kevin Kelly argues that the latent space within AI models represents a compressed form of all human knowledge, and that this latent space will become a new medium for creativity, enabling novel artistic and scientific exploration.

0 favorites 0 likes
#compression

@brianbellx: I removed 423 GB from GLM‑5.2 without changing the model. 1,403 GB → 980 GB. 753B weights. Bit for bit exact. No quanti…

X AI KOLs Timeline · 2026-07-12 Cached

A technique to remove 423 GB from GLM-5.2 (753B weights) without quantization or retraining, achieving bit-exact compression by keeping weights compressed in VRAM.

0 favorites 0 likes
#compression

Handsum: An LQIP Image File Format

Lobsters Hottest · 2026-07-11 Cached

Nigel Tao introduces Handsum, a new fixed-size LQIP image file format based on DCT, offering predictable file sizes and simplicity compared to JPEG.

0 favorites 0 likes
#compression

@VukRosic99: Most KV-cache compression applies SVD to the keys alone, or embeds queries and keys jointly. Both miss the obvious targ…

X AI KOLs Timeline · 2026-07-10 Cached

KQ-SVD is a new method for KV-cache compression that directly approximates the attention matrix via optimal low-rank decomposition, achieving 5-10x lower error than key-only SVD on LLaMA and Mistral models.

0 favorites 0 likes
#compression

It Takes a MAESTRO To Prune Bad Experts

arXiv cs.CL · 2026-07-10 Cached

This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.

0 favorites 0 likes
#compression

@omarsar0: A Visual Introduction to Information Theory (bookmark it) Information Theory is such an beautiful and powerful subject.…

X AI KOLs Following · 2026-07-09 Cached

An intuitive, visual introduction to information theory covering entropy, mutual information, and channel capacity, assuming only basic probability. The paper explains fundamental limits of compression and transmission.

0 favorites 0 likes
#compression

Empirical Minimal-Realisation Compression of Deep Neural Networks via Controllability-Observability Tests

arXiv cs.LG · 2026-07-08 Cached

This paper proposes a controllability–observability framework for compressing deep neural networks by reducing hidden-state redundancy, demonstrating significant compression with minimal accuracy loss on MNIST and CIFAR-10.

0 favorites 0 likes
#compression

Text Distance from Nested and Hierarchical Repetitions: A Compression-Based Perspective

arXiv cs.CL · 2026-07-08 Cached

This paper presents a new method for structural sequence analysis using the Ladderpath approach to extract nested and hierarchical repetitions, defining three distance measures that outperform gzip-based NCD and BERT in out-of-distribution and few-shot text classification tasks, offering a lightweight and interpretable alternative.

0 favorites 0 likes
#compression

Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving

arXiv cs.CL · 2026-07-08 Cached

This paper presents a workload-aware benchmark comparing KV-cache compression techniques (quantization, pruning, merging) on long-context LLM serving tasks, finding that compression ratio alone is a poor predictor of performance and advocating for workload-aware selection.

0 favorites 0 likes
#compression

Fable created novel 4D splat format

Hacker News Top · 2026-07-04 Cached

A novel 4D Gaussian splat format called .splat4d with tunable error bounds, offering 16-58x compression over raw splats and native HTTP Range streaming for dynamic scenes, with code and demo available.

0 favorites 0 likes
#compression

@JunchenJiang: Will this put an end to the debate that compressing KV caches might screw up LLM inference?

X AI KOLs Timeline · 2026-07-04 Cached

This tweet questions whether a new finding will resolve the debate on whether compressing KV caches harms LLM inference performance.

0 favorites 0 likes
#compression

The risk of KV cache compression

arXiv cs.LG · 2026-07-03 Cached

This paper theoretically characterizes the minimax risk of KV cache compression in transformers, providing design principles for accurate compression under causal masking, and instantiates them in a practical algorithm with promising results on LongBench.

0 favorites 0 likes
#compression

Can we use SLMs to compress data?

Reddit r/LocalLLaMA · 2026-07-03

Explores whether small language models (SLMs) can achieve lossless data compression by deliberately overfitting on training data, rethinking the typical aversion to overfitting in machine learning.

0 favorites 0 likes
#compression

MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression

arXiv cs.LG · 2026-07-02 Cached

MosaicKV introduces dynamic two-dimensional KV cache compression for long-context LLM serving, achieving up to 16x attention speedup and 3x memory reduction with minimal accuracy loss.

0 favorites 0 likes
#compression

@wikeygo: The Genius Design Hidden in WinRAR That Made a Russian Programmer a Legend for Thirty Years! The Russian programmer behind WinRAR, Eugene Roshal, came from the Soviet 'tank city' Chelyabinsk. In his early years, he programmed on an Apple clone with only 96KB of memory, developing a habit of extreme byte saving and efficient code since childhood. 19…

X AI KOLs Timeline · 2026-07-02 Cached

This article tells the story of Eugene Roshal, the Russian programmer behind WinRAR, including his programming background in the Soviet era, his self-developed compression algorithm in 1993, the launch of WinRAR in 1995, and the brilliant commercial strategy behind the 40-day trial design: no feature limitation, only pop-up reminders, allowing free personal use to build an ecosystem, forcing companies to purchase bulk licenses, while opening the decompression code but keeping the compression algorithm closed to maintain market barriers.

0 favorites 0 likes
#compression

Can you build a recognizable World Map in under 500 bytes?

Hacker News Top · 2026-07-01 Cached

A developer explores creating a recognizable world map in under 500 bytes by using ASCII characters and deflate-raw compression, achieving a 445-byte compressed map despite the uncompressed text being 8,523 bytes.

0 favorites 0 likes
#compression

I built a way to stop massive tool outputs from blowing up AI agent context windows

Reddit r/AI_Agents · 2026-06-30

LeanCTX is updated to act as a gateway managing tool outputs for AI agents, compressing large data, scrubbing secrets, and integrating external tools via MCP to prevent token blowup.

0 favorites 0 likes
#compression

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

arXiv cs.LG · 2026-06-30 Cached

Hard-KV introduces a Cascade Cache hierarchy and Logits Calibration mechanism to resolve the static-dynamic mismatch in head-adaptive KV cache compression, achieving up to 2x throughput improvement in long-context LLM inference.

0 favorites 0 likes
#compression

Compression-Driven Anomaly Detection in Brain MRI Using an Interpretable Quantum Autoencoder

arXiv cs.AI · 2026-06-29 Cached

This paper presents a quantum autoencoder for compression-driven anomaly detection in brain MRI, achieving high ROC-AUC scores and outperforming classical baselines while providing interpretable anomaly heatmaps.

0 favorites 0 likes
#compression

Compressed Whisper large-v3-turbo to 368 MB with Q3_K-matched QAT — multilingual WER results

Reddit r/openclaw · 2026-06-28

Whisper large-v3-turbo has been compressed to 368 MB using Q3_K-matched quantization-aware training, with multilingual word error rate results reported.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback