Tag
The article discusses the integration of compute shaders into p5.js to simplify teaching GPU programming through scaffolded learning, addressing educational challenges in computer graphics.
NVIDIA introduces CUDA Rust with two tracks (SIMT and Tile) for writing GPU kernels natively in Rust, enabling performance and developer experience improvements in AI systems.
A social media post highlighting talks by Cerebras CTO on GPU programming and AI hardware architecture, noting their low view counts despite being informative.
This article details the process of GPU memory writes, tracing the STG.E instruction through various hardware components like the load/store unit, coalescer, and L1 cache on an RTX 4090.
Pngine is a declarative format and runtime for WebGPU that uses S-expressions to simplify declaring and validating WebGPU configurations, enabling cross-platform sharing and export to formats like HTML or PNG with embedded runtime.
Mojo programming language has been open-sourced under an Apache 2 license, fulfilling a long-standing promise with the release of its compiler and toolchain.
The article explains CUDA shared memory swizzling techniques to optimize GPU memory access patterns, with code examples demonstrating performance improvements.
A summary of William Brandon's (performance engineer at Anthropic) GPU programming fundamentals lecture, emphasizing that understanding the streaming multiprocessor (SM) structure of GPU hardware is key to predicting performance, rather than starting solely from the software abstraction of thread blocks/threads.
An updated beginner-friendly tutorial on CUDA programming, covering how to write a simple kernel to add arrays on the GPU.
Tweet by @TheAhmadOsman pointing to a resource for learning how AI kernels work.
Modular announces Week 2 of Mojo 101, a free four-week live course teaching Mojo programming from the engineers who built it, covering value ownership and metaprogramming in the upcoming session.
This article provides a detailed walkthrough of what happens when a CUDA kernel is compiled and executed on an NVIDIA GPU, covering compilation to PTX and SASS, and the underlying hardware interaction.
Modular is hosting a session called Mojo 101 covering syntax to GPU programming, aimed at teaching the Mojo language.
The CUDA Handbook, a comprehensive guide to GPU programming with CUDA, is now available online for free reading with ads or ad-free via membership. The book covers architecture, APIs, and algorithms, with open-source code.
Claude Fable used the pyptx DSL to write a FlashAttention forward kernel for NVIDIA B200 that achieves near-parity performance with the hand-tuned CUTLASS kernel, demonstrating the potential for AI agents in compiler and DSL design.
Manning Books promotes Elliot Arledge's 'CUDA for Deep Learning,' a book that teaches GPU-level programming with CUDA to optimize deep learning model performance, available at a discount until Sunday.
A detailed thread summarizing the book 'Programming Massively Parallel Processors', focusing on CUDA and GPU programming concepts, optimization techniques, and parallel patterns.
A highly recommended 4.5-hour GPU programming lesson on CUDA and ThunderKittens by Ben Spector, offering an in-depth, behind-the-scenes look at kernel optimization.
A CMU PhD who developed the kernels now used by NVIDIA in TensorRT-LLM explains fast attention, covering fused CUDA kernels, FlashInfer, Triton, and paged-KV attention, enabling more tokens per second on the same GPU.
A PyTorch core engineer at Meta demonstrated a fast CUDA kernel optimization loop that outperforms expensive bootcamps, with the winning code merged into PyTorch via the KernelBot competition.