mxfp4

Tag

Cards List
#mxfp4

I tried running a 1.56TB MoE model on a 6GB RTX 4050 Laptop, Here’s the result

Reddit r/LocalLLaMA · 2026-07-29

Testing a 1.56TB Mixture-of-Experts model on a 6GB RTX 4050 laptop, requiring patched memory streaming with NVMe to achieve 0.106 tokens/s decode speed.

0 favorites 0 likes
#mxfp4

A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.

Reddit r/LocalLLaMA · 2026-07-27 Cached

A user runs Kimi K3, a 2.8T-parameter open-weight MoE model, on 80 RTX 5090 GPUs using only GDDR7 and Ethernet—no HBM—achieving 20 tok/s single stream, the first frontier model to run on consumer hardware.

0 favorites 0 likes
#mxfp4

Any idea why bartowski claims DeepSeek-V4-Flash is MXFP4?

Reddit r/LocalLLaMA · 2026-07-03

A user questions why bartowski claims DeepSeek-V4-Flash is in MXFP4 format when the original model page lists BF16, F32, and other tensor types, not MXFP4.

0 favorites 0 likes
#mxfp4

What is the biggest dense model that would fit into 128 GB RAM (at MXFP4)?

Reddit r/LocalLLaMA · 2026-07-01

Discusses the largest dense model that can be loaded in 128 GB RAM using MXFP4 quantization.

0 favorites 0 likes
#mxfp4

@charles_irl: Low-precision floats are weird. I have been building up my intuition by playing with them outside of inference/training…

X AI KOLs Following · 2026-06-22 Cached

A tweet thread introduces a visualizer for micro-scaling/block quant formats like NVFP4 and MXFP4, explaining how these low-precision floats work and their use in LLM inference to reduce memory bandwidth demands.

0 favorites 0 likes
#mxfp4

@zcbenz: nvfp4 vs mxfp4 is not just different choices of block size and scale format, nvfp4 uses an additional tensor-wise scale…

X AI KOLs Timeline · 2026-06-17 Cached

A technical comparison between nvfp4 and mxfp4 formats, highlighting that nvfp4 uses an additional tensor-wise scale factor to overcome fp4's range limit, allowing more precision in block-wise scale factors.

0 favorites 0 likes
#mxfp4

@dealignai: Qwen3.6-27b and 35b MXFP4 MXFP8 CRACK is out now with MTP. Enjoy uncensored speediness! 35b mxfp4: https://huggingface.…

X AI KOLs Timeline · 2026-05-24 Cached

DealignAI releases CRACK-abliterated and MXFP4/MXFP8 quantized versions of Qwen3.6-27B and 35B models, preserving MTP for faster speculative decoding on Apple Silicon.

0 favorites 0 likes
#mxfp4

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor

arXiv cs.LG · 2026-05-21 Cached

This paper decomposes MXFP4 quantization error into three additive components—scale bias, deadzone truncation, and grid noise—and proposes targeted corrections that recover BF16 accuracy to within 0.7 pp on Qwen2.5-3B and 3.0 pp on Qwen3-30B-A3B-Base for LLM reinforcement learning post-training.

0 favorites 0 likes
← Back to home

Submit Feedback