ptx

Tag

Cards List
#ptx

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

arXiv cs.CL · 2026-08-19 Cached

PTXBench is introduced as a benchmark to evaluate and adapt large language models for optimizing GPU kernels using architecture-specific PTX, showing uneven performance and fine-tuning insights.

0 favorites 0 likes
#ptx

@matthewjgunton: There are 3 basic levels in the NVIDIA software stack CUDA C++, PTX, and SASS Understanding all 3 helps you know why CU…

X AI KOLs Timeline · 2026-08-07 Cached

An educational tweet explaining the three levels of NVIDIA's software stack (CUDA C++, PTX, SASS) and how CUDA's abstraction creates a moat, while mentioning Luminal's automatic compiler search.

0 favorites 0 likes
#ptx

@elliotarledge: Claude Fable 5 [max] on KernelBench-Hard. The main kernel that impressed me was a B200 fp8 GEMM: it HAND WROTE raw SM10…

X AI KOLs Timeline · 2026-07-03 Cached

Claude Fable 5 achieves top results on KernelBench-Hard by hand-writing PTX code for B200 fp8 GEMM, outperforming other models and reaching 44-59% of peak performance on compute-bound shapes.

0 favorites 0 likes
#ptx

@charles_irl: https://x.com/charles_irl/status/2071606346844442871

X AI KOLs Timeline · 2026-06-29 Cached

This article explains the entire process of compiling and launching a CUDA kernel, from source code to hardware execution, using a simple vector addition example and detailing the role of nvcc, PTX, SASS, and ioctls.

0 favorites 0 likes
#ptx

@bingxu_: I started INT21 two months ago, and I’m proud to announce that we’re coming out of stealth today with our first product…

X AI KOLs Timeline · 2026-06-16 Cached

INT21 announced PTX Kernel Factory, a self-improving agent swarm that autonomously generates expert-level PTX GPU kernels, with open-source proof-of-concept implementations and beta access.

0 favorites 0 likes
#ptx

@elliotarledge: https://x.com/elliotarledge/status/2059409567805816872

X AI KOLs Timeline · 2026-05-26 Cached

CUDA 13.3 introduces significant enhancements including Tile C++ support, C++23 standard, improved NVRTC, stable CUDA Python 1.0 APIs, and PTX 9.3 with new fabric instructions and async multimem operations, targeting kernel developers and runtime engineers.

0 favorites 0 likes
← Back to home

Submit Feedback