Tag
AntLing open-sources Ling-3.0-flash-dspark, a DSpark draft model for Ling-3.0-flash, which achieves 1,120 tokens per second with low latency on NVIDIA Blackwell GPUs.
Nota AI releases a 4-bit quantized version of Upstage's Solar Open2 250B MoE model, using proprietary NVFP4 quantization that requires NVIDIA Blackwell GPUs.
Apple now shows a popup notifying users that AI features in iWork and Freeform send data to Google Cloud. The company also revealed its top-tier AI model, AFM Cloud Pro, runs on Nvidia Blackwell GPUs hosted in Google Cloud, marking a departure from its previous privacy promises of keeping AI data on Apple silicon in Apple's own data centers.
Meta open-sources TLX Block Attention, a warp-specialized Triton kernel that achieves 2.3x speedup for block-diagonal self-attention on NVIDIA Blackwell GPUs, with up to 3.5x speedup when fused with rotary embeddings.
llama.cpp build b9095 introduces NCCL-free tensor parallelism for dual Blackwell PCIe GPUs, enabling efficient multi-GPU inference without relying on NCCL.
A user proposes building a heterogeneous AI cluster using Blackwell GPUs and high-memory servers connected via RDMA, seeking collaboration on Tinygrad driver development.
At Nvidia GTC 2026, CEO Jensen Huang introduced the next-gen Vera Rubin system while Supermicro unveiled a full-stack AI Factory portfolio built on Nvidia Blackwell GPUs for turnkey enterprise AI deployment.