data-compression

Tag

Cards List
#data-compression

Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

arXiv cs.CL ↗ · 5d ago Cached

Introduces Uncheatable Eval, a dynamic benchmark using compression rates to evaluate language models and mitigate data contamination.

0 favorites 0 likes
#data-compression

What are you doing this week?

Lobsters Hottest ↗ · 2026-09-21 Cached

A developer shares progress on value block compression for an index using zigzag delta encoding and LZ4, aiming to save about 4 TB of disk space by enhancing compression efficiency.

0 favorites 0 likes
#data-compression

Compressing a Flag to 11 Bits

Hacker News Top ↗ · 2026-09-12

The article discusses a method for compressing a boolean flag into 11 bits, focusing on efficient data encoding techniques.

0 favorites 0 likes
#data-compression

Granular-Ball Quantum Clustering for Resource-Efficient and Robust Learning

arXiv cs.LG ↗ · 2026-09-10 Cached

Proposes Granular-Ball Quantum Clustering (GBQC), a framework that combines granular-ball structural abstraction with quantum feature learning to reduce computational overhead and improve clustering accuracy and robustness on various datasets.

0 favorites 0 likes
#data-compression

bzip3

Hacker News Top ↗ · 2026-09-07 Cached

BZip3 is a high-performance compression tool that offers superior compression ratios and speed compared to its predecessor BZip2, utilizing advanced algorithms like context mixing entropy coding and Burrows-Wheeler transform.

0 favorites 0 likes
#data-compression

InferenceFS: Never worry about data again! (Again!)

Lobsters Hottest ↗ · 2026-08-29 Cached

InferenceFS is a filesystem tool that uses the latent space of large language models to infer file contents from filenames, offering theoretical infinite compression by relying on external AI backends like Claude and Gemini.

0 favorites 0 likes
#data-compression

@lateinteraction: this is apparently a 64x compression ratio at effectively no loss in quality - really nice work @yjoonjang !

X AI KOLs Following ↗ · 2026-08-27

A tweet highlights a 64x compression ratio achieved with effectively no loss in quality, praising the work of @yjoonjang.

0 favorites 0 likes
#data-compression

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

KATok is an adaptive video tokenizer that selectively drops uninformative tokens for data-dependent compression, improving spatial consistency in diffusion-based video generation.

0 favorites 0 likes
#data-compression

Green BOA: Determining the environmental break-even point for ML-based data compression

arXiv cs.LG ↗ · 2026-08-21 Cached

The paper presents an analysis of the environmental break-even point for ML-based data compression, estimating carbon-equivalent costs for training and inference against savings from reduced disk storage.

0 favorites 0 likes
#data-compression

Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions

arXiv cs.LG ↗ · 2026-08-19 Cached

Introduces iterative tensor network transformations (ITNTs) for efficient element-wise evaluation of nonlinear functions on tensor trains, enabling data processing and optimization on large-scale datasets with applications in reactive flow fields and Max-SAT problems.

0 favorites 0 likes
#data-compression

ALP: Adaptive Lossless Floating-Point Compression

Lobsters Hottest ↗ · 2026-07-24 Cached

This paper introduces ALP, a state-of-the-art lossless compression algorithm for IEEE 754 floating-point data, exploiting decimal and high-precision patterns. It achieved superior performance in decoding speed, compression ratio, and compression speed, earning the SIGMOD Best Artifact Award.

0 favorites 0 likes
#data-compression

Building a World Map with only 500 bytes

Simon Willison's Blog ↗ · 2026-07-04 Cached

Iwo Kadziela created a method to generate a credible ASCII world map using only 445 bytes of data by leveraging deflate compression and JavaScript's DecompressionStream API.

0 favorites 0 likes
#data-compression

Suffix BWT vs cyclic shift BWT, and fast computation

Lobsters Hottest ↗ · 2026-07-04 Cached

Explains two variants of the Burrows-Wheeler transform (cyclic shift and suffix) and fast computation methods, aimed at data compression enthusiasts.

0 favorites 0 likes
#data-compression

Can we use SLMs to compress data?

Reddit r/LocalLLaMA ↗ · 2026-07-03

Explores whether small language models (SLMs) can achieve lossless data compression by deliberately overfitting on training data, rethinking the typical aversion to overfitting in machine learning.

0 favorites 0 likes
#data-compression

Overfitted a 900KB Transformer to Compress a 100MB CSV into 7MB

Hacker News Top ↗ · 2026-06-23 Cached

A developer overfits a small 900KB transformer model to compress a 100MB CSV file down to 7MB, demonstrating a novel approach to data compression using overfitted neural networks.

0 favorites 0 likes
#data-compression

PivCo-Huffman “merge” operations

Lobsters Hottest ↗ · 2026-06-22 Cached

This blog post analyzes the PivCo-Huffman paper, which introduces 'merge' operations for parallel Huffman decoding, enabling efficient vectorized and GPU-friendly decoding without interleaving overhead.

0 favorites 0 likes
#data-compression

@HowToPrompt__: Vector databases are officially cooked This repo shrinks 60 million text chunks from 201 GB to just 6 GB without any lo…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

A new open-source repo compresses 60 million text chunks from 201 GB to 6 GB with zero loss in accuracy, making vector databases potentially obsolete for many use cases.

0 favorites 0 likes
#data-compression

Data Compression Explained (2012)

Hacker News Top ↗ · 2026-06-16 Cached

A comprehensive book explaining data compression techniques including information theory, coding methods, modeling, and transforms, targeting programmers with math skills.

0 favorites 0 likes
#data-compression

Unlocking Latent Dimensions: Exploring Representations of Large-Scale X-ray Scattering Data using Variational Autoencoders

arXiv cs.LG ↗ · 2026-06-16 Cached

This paper explores the use of variational autoencoders to learn latent representations of large-scale X-ray scattering data, enabling efficient data compression and analysis.

0 favorites 0 likes
#data-compression

@HowToAI_: This repo shrinks 201GB of text down to 6GB without losing any accuracy. → 97% smaller than vector DBs → Runs locally →…

X AI KOLs Timeline ↗ · 2026-05-15 Cached

This repository compresses 201GB of text down to 6GB with no accuracy loss, making it 97% smaller than vector databases. It runs locally and offers a drop-in MCP for Claude, fully open source and private.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback