@RisingSayak: Published my first kernel to go the last mile to optimize LTX-2.3 from @Lightricks! torch.compile + cuDNN attn already …

X AI KOLs Following Tools

Summary

Published a custom kernel to further optimize LTX-2.3 from Lightricks, achieving 1.52x speedup on GB10, building upon previous torch.compile and cuDNN attention optimizations.

Published my first kernel to go the last mile to optimize LTX-2.3 from @Lightricks! torch.compile + cuDNN attn already gave a 1.42x boost. W/ the custom kernel added, I got 1.52x on a GB10 🔥 This was my systematic exploration of a simple agentic kernel dev workflow. More 👇 https://t.co/u4iDpzSir0
Original Article
View Cached Full Text

Cached at: 06/13/26, 02:27 PM

Published my first kernel to go the last mile to optimize LTX-2.3 from @Lightricks!

torch.compile + cuDNN attn already gave a 1.42x boost. W/ the custom kernel added, I got 1.52x on a GB10 🔥

This was my systematic exploration of a simple agentic kernel dev workflow.

More 👇 https://t.co/u4iDpzSir0

Similar Articles