Tag
Introduces Uncheatable Eval, a dynamic benchmark using compression rates to evaluate language models and mitigate data contamination.
A developer shares progress on value block compression for an index using zigzag delta encoding and LZ4, aiming to save about 4 TB of disk space by enhancing compression efficiency.
The article discusses a method for compressing a boolean flag into 11 bits, focusing on efficient data encoding techniques.
Proposes Granular-Ball Quantum Clustering (GBQC), a framework that combines granular-ball structural abstraction with quantum feature learning to reduce computational overhead and improve clustering accuracy and robustness on various datasets.
BZip3 is a high-performance compression tool that offers superior compression ratios and speed compared to its predecessor BZip2, utilizing advanced algorithms like context mixing entropy coding and Burrows-Wheeler transform.
InferenceFS is a filesystem tool that uses the latent space of large language models to infer file contents from filenames, offering theoretical infinite compression by relying on external AI backends like Claude and Gemini.
A tweet highlights a 64x compression ratio achieved with effectively no loss in quality, praising the work of @yjoonjang.
KATok is an adaptive video tokenizer that selectively drops uninformative tokens for data-dependent compression, improving spatial consistency in diffusion-based video generation.
The paper presents an analysis of the environmental break-even point for ML-based data compression, estimating carbon-equivalent costs for training and inference against savings from reduced disk storage.
Introduces iterative tensor network transformations (ITNTs) for efficient element-wise evaluation of nonlinear functions on tensor trains, enabling data processing and optimization on large-scale datasets with applications in reactive flow fields and Max-SAT problems.
This paper introduces ALP, a state-of-the-art lossless compression algorithm for IEEE 754 floating-point data, exploiting decimal and high-precision patterns. It achieved superior performance in decoding speed, compression ratio, and compression speed, earning the SIGMOD Best Artifact Award.
Iwo Kadziela created a method to generate a credible ASCII world map using only 445 bytes of data by leveraging deflate compression and JavaScript's DecompressionStream API.
Explains two variants of the Burrows-Wheeler transform (cyclic shift and suffix) and fast computation methods, aimed at data compression enthusiasts.
Explores whether small language models (SLMs) can achieve lossless data compression by deliberately overfitting on training data, rethinking the typical aversion to overfitting in machine learning.
A developer overfits a small 900KB transformer model to compress a 100MB CSV file down to 7MB, demonstrating a novel approach to data compression using overfitted neural networks.
This blog post analyzes the PivCo-Huffman paper, which introduces 'merge' operations for parallel Huffman decoding, enabling efficient vectorized and GPU-friendly decoding without interleaving overhead.
A new open-source repo compresses 60 million text chunks from 201 GB to 6 GB with zero loss in accuracy, making vector databases potentially obsolete for many use cases.
A comprehensive book explaining data compression techniques including information theory, coding methods, modeling, and transforms, targeting programmers with math skills.
This paper explores the use of variational autoencoders to learn latent representations of large-scale X-ray scattering data, enabling efficient data compression and analysis.
This repository compresses 201GB of text down to 6GB with no accuracy loss, making it 97% smaller than vector databases. It runs locally and offers a drop-in MCP for Claude, fully open source and private.