fp4

Tag

Cards List
#fp4

New set of FP4 attention kernels for B300, achieving up to 1.69x speedup over FA4

Reddit r/LocalLLaMA · 2026-07-14 Cached

The FastVideo team releases new FP4 attention kernels for B300, achieving up to 1.69x speedup over FlashAttention 4.

0 favorites 0 likes
#fp4

@HuggingPapers: NVIDIA just released the NVFP4 quantized Kimi-K2.7-Code on Hugging Face A 1T-parameter Moonshot AI model quantized to F…

X AI KOLs Following · 2026-07-10 Cached

NVIDIA released the NVFP4 quantized Kimi-K2.7-Code, a 1 trillion-parameter Moonshot AI model quantized to FP4 for Blackwell GPUs, preserving accuracy with reduced memory usage.

0 favorites 0 likes
#fp4

@coffeecup2020: If your card support Blackwell, read this! https://github.com/turbo-tan/llama.cpp-tq3… updated with turbo4/turbo3 TQ3_4…

X AI KOLs Timeline · 2026-07-01 Cached

A llama.cpp fork introduces TurboQuant TQ3_4S quantization that maps to Blackwell FP4 tensor cores, achieving up to 221% faster prompt processing on GB10 while maintaining near Q4 quality at Q3 size.

0 favorites 0 likes
#fp4

@AaronWeiHuang: Our new blog looks at how FP4 is moving beyond compression into a practical primitive for training and inference across…

X AI KOLs Following · 2026-06-30 Cached

NVIDIA's blog details how FP4, with the NVFP4 format and Blackwell hardware, has evolved from a compression trick to a practical primitive for training and inference across LLMs and diffusion models, achieving near 16-bit accuracy.

0 favorites 0 likes
#fp4

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference

arXiv cs.LG · 2026-06-26 Cached

SharQ introduces a training-free method combining activation sparsity and FP4 quantization for LLM inference, using sparse-dense decomposition and a unified FP4 weight payload. It achieves significant latency reduction and accuracy recovery over FP4-only baselines.

0 favorites 0 likes
#fp4

@RayFernando1337: “The selected runtime uses NVFP4 weights for maximum performance. From the original FP8 weights, we performed an in-hou…

X AI KOLs Following · 2026-06-23

Discusses using NVFP4 4-bit floating point weights for maximum performance, achieved via in-house quantization from FP8 using NVIDIA ModelOpt, highlighting the data format's dual scale factors for high dynamic range.

0 favorites 0 likes
#fp4

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

Hugging Face Daily Papers · 2026-06-18 Cached

This paper identifies a fundamental limitation (shrinkage bias) in non-uniform FP4 quantization formats for LLM pretraining and proposes UFP4, a uniform 4-bit training recipe that outperforms existing E2M1-based methods.

0 favorites 0 likes
#fp4

@Italianclownz: Converted Qwen 3.6 35b a3b to ROCmfp4 and this is flying. Used the mtp version bc this ROCmfp4 can also incorporate the…

X AI KOLs Timeline · 2026-05-24 Cached

Converted the Qwen 3.6 35b a3b model to ROCmfp4 format, leveraging MTP benefits for improved performance on AMD hardware.

0 favorites 0 likes
#fp4

@charles_irl: another page for the @modal LLMEng Almanac: an explorer for low-precision floats, from bf16 to fp4 https://modal.com/ll…

X AI KOLs Following · 2026-05-18 Cached

A page from Modal's LLM Engineer's Almanac that provides an interactive explorer for understanding low-precision floating-point formats like bf16 and fp4.

0 favorites 0 likes
← Back to home

Submit Feedback