Tag
Firebird launched the CIS region's largest AI factory in Armenia, powered by NVIDIA accelerated computing and Dell infrastructure, with plans to deploy over 70,000 NVIDIA GPUs and 300 MW of capacity by 2027. NVIDIA also intends to invest in Firebird as part of a broader 2-gigawatt roadmap across Armenia, Kazakhstan, and other markets.
Cursor open-sources Mixture-of-Kittens (MoK), a deterministic MoE training megakernel for NVIDIA Blackwell GPUs that fuses computation and communication, delivering up to 2.37x speedup over baseline implementations.
The tweet shares evidence that NVIDIA intentionally shapes its software ecosystem to make newer workflows appear Blackwell-exclusive while leaving Ampere paths unsupported, advising users to wait before upgrading and suggesting NVIDIA's moat is weakening in favor of Intel and AMD.
Krasis, a MoE-focused runtime, enables running the 397B-parameter Ornith model on a single RTX PRO 6000 Blackwell 96GB GPU with ~20-24 tok/s decode by dynamically managing expert residency in VRAM.
Technical post detailing how to run DeepSeek V4 Flash on two Nvidia 4090d GPUs using custom Triton kernels and vLLM, achieving ~105 tokens/second with 262k context.
NVIDIA details how its full-stack inference software, co-developed with open-source ecosystems like PyTorch, reduces token costs by up to 5x on Blackwell GPUs, with real-world deployments from Baseten, Cognition, Deep Infra, and others demonstrating performance gains.
Baseten releases GLM-5.2-Vision, a vision-language model that adds MoonViT vision encoder to GLM-5.2 via a trained PatchMerger projector, keeping the text backbone and vision tower frozen. The model is quantized to NVFP4 for efficient inference on Blackwell hardware.
User asks the community for recommendations on which large language model to deploy on a high-end local setup with 4 RTX PRO 6000 GPUs (384GB total), primarily for internal company policy management and thinking tasks.
NVIDIA argues that performance per watt is the key metric for AI infrastructure efficiency, highlighting how its Blackwell and Vera Rubin platforms achieve up to 25x improvement over Hopper for MoE models.
This article provides an in-depth interpretation of NVIDIA's newly released 'AI Model Co-Design' paper, pointing out that in AI inference scenarios, storage (memory bandwidth, weight reading) has replaced GPU compute as the primary bottleneck. It elaborates on the design strategies of TensorRT-LLM and Blackwell architecture around the Roofline model, emphasizing that reducing data movement is more critical than improving compute power.
FastAFD is an open-source serving system for Attention-FFN Disaggregation of MoE models on Blackwell NVL72, achieving 1.35-1.45× per-GPU decode throughput improvement over colocated MoE serving.
A detailed report on optimizing a production vLLM serving configuration on NVIDIA's DGX Spark, correcting flags that were costing 34% MTP acceptance after reviewing 90+ official NVIDIA documents and running a 69-scenario tool evaluation.
NVIDIA reported that its Blackwell inference stack reduced DeepSeek V4 token costs by up to 5x in one month.
NVIDIA's blog details how FP4, with the NVFP4 format and Blackwell hardware, has evolved from a compression trick to a practical primitive for training and inference across LLMs and diffusion models, achieving near 16-bit accuracy.
NVIDIA's full-stack inference software, codesigned with hardware, has reduced token costs by up to 5x on the Blackwell platform in just one month, enabling lower cost per token for AI factories. Companies like Baseten, Cognition, Deep Infra, and Together AI are using the stack to optimize inference performance.
A detailed analysis of how NVIDIA GPU programming evolved from Volta to Blackwell, highlighting the shift from synchronous thread models to asynchronous dataflow and the challenges of feeding Tensor Cores. The article discusses new hardware features like TMA, TMEM, and tcgen05 MMA, and shows how modern kernels like FlashAttention-3 and FlashMLA exploit these changes for higher utilization.
Announcing Orinth 1.0 AEON ULTIMATE UNCENSORED, a model with BF16 and NVFP4 quantization for DGX Spark/Blackwell architecture, claiming 200-300% performance improvement with working DFlash.
A user asks about hardware requirements for serving GLM-5.2 in NVFP4 format, which vLLM now supports with reduced memory footprint and maintained accuracy.
Reports of a 96GB VRAM modded RTX 5090 (Blackwell RTX 6000) are confirmed from Shenzhen's Huaqiangbei market, priced around $8,200 total for the hacked card.
A user discusses a locked Dell quote for 6x RTX PRO 6000 Max-Q GPUs at a discounted price to build an inference cluster for GLM 5.2, asking the community for advice on purchasing strategy before the quote expires.