I'm eager for a 15x speedup on my strix halo

Reddit r/LocalLLaMA News

Summary

Nvidia claims a 15x speedup in text generation using a diffusion model, generating entire blocks at once.

Nvidia says 15x speed up possible with diffusion model. Entire block of text generated at once. https://x.com/NVIDIAAI/status/2069465510790545761
Original Article

Similar Articles

DiffusionGemma: 4x Faster Text Generation

Hacker News Top

Google introduces DiffusionGemma, an experimental 26B MoE open model that achieves up to 4x faster text generation on GPUs using text diffusion, targeting speed-critical interactive local workflows.

2x Strix Halo speed-up with an R9700

Reddit r/LocalLLaMA

A user shares how they achieved a 2x performance boost in AI inference by splitting a large MoE model between a Strix Halo APU and an R9700 GPU, detailing configurations and code modifications.

@charles_irl: dflash go brr

X AI KOLs Timeline

NVIDIA announces DFlash, an open source block diffusion model for speculative decoding that achieves up to 15x higher inference throughput on Blackwell GPUs while maintaining interactivity.