gpu-programming

Tag

Cards List
#gpu-programming

@ZhihuFrontier: GPU programming changed because Tensor Cores became too fast to feed Zhihu contributor THU-PACMAN实验室 shared a sharp bre…

X AI KOLs Timeline ↗ · 2026-06-30 Cached

A detailed analysis of how NVIDIA GPU programming evolved from Volta to Blackwell, highlighting the shift from synchronous thread models to asynchronous dataflow and the challenges of feeding Tensor Cores. The article discusses new hardware features like TMA, TMEM, and tcgen05 MMA, and shows how modern kernels like FlashAttention-3 and FlashMLA exploit these changes for higher utilization.

0 favorites 0 likes
#gpu-programming

@charles_irl: https://x.com/charles_irl/status/2071606346844442871

X AI KOLs Timeline ↗ · 2026-06-29 Cached

This article explains the entire process of compiling and launching a CUDA kernel, from source code to hardware execution, using a simple vector addition example and detailing the role of nvcc, PTX, SASS, and ioctls.

0 favorites 0 likes
#gpu-programming

@YoussefHosni951: Most engineers don't fail at CUDA because it's hard. They fail because they read the right books in the wrong order. CU…

X AI KOLs Timeline ↗ · 2026-06-23 Cached

A thread recommending the optimal order to read CUDA books, starting with CUDA by Example to build intuition before diving into more advanced texts.

0 favorites 0 likes
#gpu-programming

@neural_avb: TIL about "GPU Mode" They got a youtube series to learn CUDA. Plus a github repo with slides/notebooks. Some lectures a…

X AI KOLs Timeline ↗ · 2026-06-23 Cached

GPU Mode is a learning resource featuring a YouTube series, GitHub repo with slides/notebooks, and a practice website for mastering CUDA programming.

0 favorites 0 likes
#gpu-programming

Modern GPU Programming for MLSys

Hacker News Top ↗ · 2026-06-23 Cached

A new book from CMU's Machine Learning Systems course teaches modern GPU programming for ML systems, covering Blackwell architecture, GEMM, and FlashAttention using the TIRx Python DSL.

0 favorites 0 likes
#gpu-programming

@Modular: Free 4-week Mojo course, taught by the team that built the language. See why Mojo is the ideal language for agentic dev…

X AI KOLs Timeline ↗ · 2026-06-22 Cached

Modular announces a free 4-week Mojo course taught by the Mojo team, covering language fundamentals to GPU programming, starting July 9th on YouTube.

0 favorites 0 likes
#gpu-programming

@reprompting: Naive CUDA softmax using shared memory reduction. Reduction seems to be a pretty straightforward concept.

X AI KOLs Timeline ↗ · 2026-06-17 Cached

A tweet sharing a naive CUDA softmax implementation using shared memory reduction, noting that reduction is straightforward.

0 favorites 0 likes
#gpu-programming

Show HN: cuTile Rust: Safe, data-race-free GPU kernels in Rust

Hacker News Top ↗ · 2026-06-16 Cached

NVIDIA Labs releases cuTile Rust, a tile-based system for writing memory-safe, data-race-free GPU kernels in idiomatic Rust. It extends Rust's ownership model to GPU kernels, JIT-compiles Rust AST to GPU code, and achieves performance close to native CUDA.

0 favorites 0 likes
#gpu-programming

@levidiamode: 163/365 of GPU Programming Looking at a few different agentic GPU kernel optimization systems today. The two I'm most i…

X AI KOLs Timeline ↗ · 2026-06-15 Cached

A tweet discussing two agentic GPU kernel optimization systems: Auto GPU Kernel by @dogacel0 and Kernel Design Agents from @songhan_mit's lab, both winners at the MLSys Sparse Attention FlashInfer competition. The thread highlights different approaches using subagents and Claude skills for GPU programming.

0 favorites 0 likes
#gpu-programming

@pradheepraop: implemented the top-k kernel from the kernel design section in the msa paper. https://github.com/Mantissagithub/learn_c…

X AI KOLs Timeline ↗ · 2026-06-15 Cached

Implemented a top-k kernel from the kernel design section of the MSA paper, using exp-free comparison and warp-level tree merging with CUDA shuffles. The code is available on GitHub.

0 favorites 0 likes
#gpu-programming

@levidiamode: 158/365 of GPU Programming I think I understand the high level differences between the FlashAttention 2, 3 and 4 forwar…

X AI KOLs Timeline ↗ · 2026-06-10 Cached

The author documents their progress in learning GPU programming, focusing on understanding the high-level differences between FlashAttention 2, 3, and 4 forward passes, and lists several low-level concepts they need to explore further.

0 favorites 0 likes
#gpu-programming

@levidiamode: 157/365 of GPU Programming Another FlashAttention4 resource that's been really helpful for me is the talk @charles_irl …

X AI KOLs Following ↗ · 2026-06-09 Cached

A daily GPU programming thread highlights a talk by Charles_irl that reverse-engineers FlashAttention4 code before the paper release, praising the Modal team's deep code dissection and inferences about the forward pass.

0 favorites 0 likes
#gpu-programming

What about OpenCL and CUDA C++ alternatives?

Hacker News Top ↗ · 2026-06-09 Cached

This article examines the history of CUDA alternatives like OpenCL and SYCL, explaining why they failed to become dominant in AI compute due to slow committee-driven development and the challenges of open coopetition.

0 favorites 0 likes
#gpu-programming

@kazukifujii: Tech Blog Release Day5 This is the first installment of a blog series that explains CUDA Programming from the basics, w…

X AI KOLs Timeline ↗ · 2026-06-04 Cached

Kazuki Fujii announces the first installment of a blog series on CUDA Programming basics, written in an accessible way, essential for understanding FlashAttention and hardware-aware acceleration techniques.

0 favorites 0 likes
#gpu-programming

@elliotarledge: https://x.com/elliotarledge/status/2059409567805816872

X AI KOLs Timeline ↗ · 2026-05-26 Cached

CUDA 13.3 introduces significant enhancements including Tile C++ support, C++23 standard, improved NVRTC, stable CUDA Python 1.0 APIs, and PTX 9.3 with new fabric instructions and async multimem operations, targeting kernel developers and runtime engineers.

0 favorites 0 likes
#gpu-programming

@charles_irl: New articles in the GPU Glossary for CuTe DSL, CUTLASS, and CuTe -- the tools used to write some of the highest-perform…

X AI KOLs Following ↗ · 2026-05-26 Cached

New articles in the GPU Glossary cover CuTe DSL, CUTLASS, and CuTe – tools for writing high-performance GPU kernels on data center GPUs, with examples in Python.

0 favorites 0 likes
#gpu-programming

@levidiamode: Day 138/365 of GPU Programming One of my favorite lectures I've watched this year is Stanford's CS336 lecture 7 on GPU …

X AI KOLs Timeline ↗ · 2026-05-21 Cached

A learner shares enthusiasm for Stanford CS336 lecture 7 on GPU parallelism, which covers fundamental operations and connects them to multi-GPU setups and parallelism techniques like tensor, data, and pipeline parallelism.

0 favorites 0 likes
#gpu-programming

@ManningBooks: PyTorch gets you pretty far, but when performance becomes the problem, understanding what's happening at the GPU level …

X AI KOLs Timeline ↗ · 2026-05-19 Cached

Promotional post for the book 'CUDA for Deep Learning' by Elliot Arledge, offering a first chapter summary video that explains GPU performance, the CUDA programming model, and when to write custom CUDA kernels.

0 favorites 0 likes
#gpu-programming

CUDA Books

Hacker News Top ↗ · 2026-05-17 Cached

A curated list of major books on CUDA programming covering beginner to advanced topics, including C++ and Python, with focus on practical resources for NVIDIA GPU parallel computing.

0 favorites 0 likes
#gpu-programming

CUDA-oxide: Nvidia's official Rust to CUDA compiler

Hacker News Top ↗ · 2026-05-11 Cached

CUDA-oxide is an experimental Rust-to-CUDA compiler developed by NVIDIA that enables writing safe GPU kernels in idiomatic Rust, compiling directly to PTX without requiring domain-specific languages or foreign bindings.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback