Tag
NVIDIA announces DFlash, an open source block diffusion model for speculative decoding that achieves up to 15x higher inference throughput on Blackwell GPUs while maintaining interactivity.
Ollama doubled GPU capacity for GLM 5.2 on its US cloud, using NVIDIA B300 Blackwell GPUs, emphasizing privacy and open models.
NVIDIA's Blackwell platform achieved fastest training times across all MLPerf Training 6.0 benchmarks, scaling to 8,192 GPUs and showcasing up to 1.6x performance gains with the GB300 NVL72 over the GB200 NVL72.
NVIDIA's Blackwell GB300 NVL72 platform leads the first agentic AI infrastructure benchmark, AgentPerf from Artificial Analysis, delivering up to 20x more agents per megawatt than the previous Hopper generation.
A benchmark of NVFP4 on an RTX 5090 with Qwen3.6-27B shows prefill speed gains of 32-42% over equal-bit Q4_K_M and 52-68% over Q6_K, but decode gains are modest (+9% vs Q4) as decode is memory-bandwidth bound. The quality loss compared to Q6 is minimal (-0.8 average), making NVFP4 a good choice for local inference.
A user shares performance benchmarks comparing the Nvidia RTX Pro 4500 Blackwell 32GB GPU against the RTX 5060 Ti 16GB for AI inference, showing 1.6-6x speed improvements depending on model size and quantization.
A thread reviewing the paper 'Pretraining Large Language Models with NVFP4' and discussing NVFP4 pre-training, especially for NVIDIA Blackwell.
BIS issued guidance requiring licenses for advanced AI chip exports to Chinese-headquartered firms located outside China, highlighting previous enforcement gaps for overseas subsidiaries like Tencent Malaysia buying Nvidia Blackwell chips.
A developer documents the extensive hardware and firmware hacking required to run an NVIDIA RTX Pro 6000 Blackwell GPU in a legacy Dell PowerEdge R730 server, achieving 650K context length for local AI inference.
Nvidia teases a new ARM-based PC laptop chip (likely N1X) to be announced at Computex on June 2. The chip is a lower-power version of the GB10 Superchip used in DGX Spark, featuring 20 ARM cores and 6144 CUDA cores, targeting Windows laptops.
The author presents SM1, a variant of Mamba1 with d_state=1, using two native PyTorch ops to replace the selective scan, reducing memory by 16x compared to d_state=16. The closed-form solution eliminates the state dimension, enabling efficient inference with constant memory per token.
Llama.cpp now supports Nvidia's Programmatic Dependent Launch (PDL) for Blackwell GPUs, offering a 5-10% performance boost on token generation. The feature is not enabled by default and requires a build flag.
A developer toolkit providing configurations, wheels, and benchmarks for running large language models with NVFP4 precision on Nvidia Blackwell GPUs using TensorRT-LLM.
A user demonstrates successfully running the DeepSeek V4 Pro model on a local workstation using a modified llama.cpp CUDA repository, highlighting performance metrics and hardware requirements.
Michael Goin reviews the vLLM v0.20.0 release, highlighting 752 commits and new features like DeepSeek V4 support, TurboQuant, and PyTorch 2.11 integration.
NVIDIA announces 16 games joining GeForce NOW cloud streaming in May, including new AAA titles like Forza Horizon 6 and 007 First Light, and expands RTX 5080-class performance across the library for Ultimate members.
Supermicro and NVIDIA unveil turnkey “AI Factory” reference architectures combining Blackwell GPUs, certified servers, networking, storage and deployment services to let enterprises spin up cluster-scale AI infrastructure faster.
Yahoo Finance reports on Nvidia CEO Jensen Huang's long-term vision for the company over the next decade, as Nvidia and Supermicro advance turnkey AI Factory infrastructure built on Blackwell systems.
NVIDIA CEO Jensen Huang highlighted an inflection point in AI inference during the GTC keynote, while Supermicro is partnering with NVIDIA to deliver turnkey 'AI Factory' infrastructure solutions built around the Blackwell platform.