Tag
The article details optimization techniques for FlashAttention-4, including an S/P ping-pong method for decode to overlap operations, achieving up to 16% performance gain on NVIDIA Blackwell GPUs.
A llama.cpp fork introduces TurboQuant TQ3_4S quantization that maps to Blackwell FP4 tensor cores, achieving up to 221% faster prompt processing on GB10 while maintaining near Q4 quality at Q3 size.
A tweet highlights that the abliterated, NVFP4 quantized Gemma-4-12B model (7.7 GB) can rival Qwen 3.6-35B in practical tasks while running fast on Blackwell GPUs, demonstrating significant efficiency gains.
Nvidia announced RTX Spark, an Arm-based chip for Windows PCs combining a 20-core Grace CPU, up to 6,144 Blackwell GPU cores, and up to 128GB unified memory, aiming to bring high performance and AI capabilities to slim laptops and compact desktops.
LongLive-2.0 introduces an NVFP4-based parallel infrastructure for long video generation, achieving up to 2.15x training speedup and 1.84x inference speedup with a 5B model reaching 45.7 FPS.