nvfp4

Tag

Cards List
#nvfp4

@nrehiew_: For the visual learners

X AI KOLs Timeline · 2026-06-05 Cached

A thread reviewing the paper 'Pretraining Large Language Models with NVFP4' and discussing NVFP4 pre-training, especially for NVIDIA Blackwell.

0 favorites 0 likes
#nvfp4

@TheAhmadOsman: My pal Jensen is delivering Frontier Opensource Intelligence (that is extremely cost efficient) just like he said he wo…

X AI KOLs Following · 2026-06-01 Cached

Jensen Huang hints at more Nemotron model releases, highlighting open-source frontier intelligence and cost efficiency enabled by NVFP4 training.

0 favorites 0 likes
#nvfp4

@vllm_project: vLLM v0.22.0 is out! 459 commits from 230 contributors (63 new). Highlights: DeepSeek V4 hardening (NVFP4 fused MoE, fu…

X AI KOLs Timeline · 2026-05-30 Cached

vLLM v0.22.0 released with 459 commits, featuring DeepSeek V4 hardening, experimental Rust frontend, and batch-invariant Cutlass FP8, reducing end-to-end latency by 28.9%.

0 favorites 0 likes
#nvfp4

@mr_r0b0t: Official @NVIDIAAI GLM5.1-NVFP4 spotted on @huggingface

X AI KOLs Timeline · 2026-05-28 Cached

NVIDIA releases GLM-5.1-NVFP4, a quantized version of ZAI's GLM-5.1 model with 754B total parameters (40B activated), available on Hugging Face under MIT license.

0 favorites 0 likes
#nvfp4

@mr_r0b0t: 16 local AI agents streaming at once! MiniMax M2.7 NVFP4 — 2x GB10, no cloud APIs.

X AI KOLs Timeline · 2026-05-25 Cached

A demonstration shows 16 local AI agents streaming simultaneously using MiniMax M2.7 NVFP4 on two Nvidia GB10 chips, with no cloud APIs required.

0 favorites 0 likes
#nvfp4

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

arXiv cs.CL · 2026-05-21 Cached

Mix-Quant proposes a phase-aware quantization framework for agentic LLMs, using NVFP4 quantization for the prefilling stage to accelerate computation while preserving BF16 precision for decoding to maintain accuracy. The method achieves up to 3x speedup in prefilling with minimal performance degradation on agentic benchmarks.

0 favorites 0 likes
#nvfp4

Real-Time Long Video Generation (GitHub Repo)

TLDR AI · 2026-05-20 Cached

NVlabs releases LongLive 2.0, a parallel infrastructure for real-time long video generation using NVFP4 quantization, supporting both training and inference. It achieves 45.7 FPS and is accepted at ICLR 2026.

0 favorites 0 likes
#nvfp4

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

Hugging Face Daily Papers · 2026-05-18 Cached

LongLive-2.0 introduces an NVFP4-based parallel infrastructure for long video generation, achieving up to 2.15x training speedup and 1.84x inference speedup with a 5B model reaching 45.7 FPS.

0 favorites 0 likes
#nvfp4

@ctnzr: We've gone even farther: Nemotron 3 Super is 120B and pretrained on 25T tokens in NVFP4. Nemotron 3 Ultra is ~500B and …

X AI KOLs Following · 2026-05-15 Cached

NVIDIA announces Nemotron 3 Super (120B) and Nemotron 3 Ultra (~500B) models, pretrained on 25T tokens using NVFP4 precision, emphasizing accelerated computing and efficiency improvements.

0 favorites 1 likes
#nvfp4

@HowToAI_: NVIDIA has done the impossible and nobody's talking about it. They trained a 12 BILLION parameter LLM in 4-bit precisio…

X AI KOLs Timeline · 2026-05-15

NVIDIA trained a 12-billion parameter LLM in 4-bit precision using the new NVFP4 format with micro-scaling, achieving near-zero intelligence loss while halving memory usage and tripling arithmetic speed, marking a major breakthrough in efficient AI training.

0 favorites 0 likes
#nvfp4

NVFP4 Kimi2.6 and Kimi 2.5 released by Nvidia

Reddit r/LocalLLaMA · 2026-05-14

Nvidia released NVFP4 quantized versions of Moonshot AI's Kimi-K2.6 and Kimi-K2.5 language models, maintaining high accuracy and available for commercial and non-commercial use.

0 favorites 0 likes
#nvfp4

Blackwell LLM Toolkit - NVFP4 Config +Wheels + Benchmarks for Blackwell GPUs via TensorRT-LLM - 270 tk/s Nemotron 3 Omni

Reddit r/LocalLLaMA · 2026-05-12

A developer toolkit providing configurations, wheels, and benchmarks for running large language models with NVFP4 precision on Nvidia Blackwell GPUs using TensorRT-LLM.

0 favorites 0 likes
#nvfp4

@0xSero: Just added 2 new model compressions: Hy3-FP8 & NVFP4 I recommend trying this model it's very strong and fits on 256gb o…

X AI KOLs Following · 2026-05-10 Cached

0xSero has released new FP8 and NVFP4 quantized versions of the Tencent Hy3-preview model, enabling it to run on 256GB VRAM with full context.

0 favorites 0 likes
#nvfp4

unsloth/Qwen3.6-27B-NVFP4

Hugging Face Models Trending · 2026-04-23 Cached

Unsloth releases an NVFP4 quantized checkpoint of Qwen3.6-27B, claiming 2.5x faster throughput and accuracy comparable to FP8 and BF16, with instructions for running on a 24GB GPU via vLLM.

0 favorites 0 likes
#nvfp4

Qwen3.6-27B KLDs - INTs and NVFPs

Reddit r/LocalLLaMA · 2026-04-22

Reddit post compares quantized Qwen3.6-27B variants (INT4, NVFP4, BF16-INT4) showing trade-offs between memory size and accuracy for different use-cases.

0 favorites 0 likes
#nvfp4

@0xSero: GLM-5.1-478B-NVFP4 Running on: - 4x RTX Pro 6000 - Sglang - 370,000 max tokens (1.75x full context) - p10 27.7 | p90 45…

X AI KOLs Timeline · 2026-04-21 Cached

A quantized 478B-parameter GLM-5.1 model runs on 4×RTX Pro 6000 GPUs via SGLang, delivering 370k-token context at up to 45 tok/s decode and 1340 tok/s prefill, and is demoed driving Figma.

0 favorites 0 likes
#nvfp4

@0xSero: Finally GLM-5.1-505B-REAP-NVFP4 45 tokens/s decode 1350 tokens/s prefill 32% prune This was the hardest I ever worked t…

X AI KOLs Timeline · 2026-04-20 Cached

Developer @0xSero achieved high-performance inference on an optimized GLM-5.1-505B variant using NVFP4 quantization and 32% pruning, reaching 45 tokens/s decode and 1350 tokens/s prefill speeds.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback