cuda-graphs

Tag

Cards List
#cuda-graphs

@PyTorch: The posters are live for #PyTorchCon North America! From compiler optimization & CUDA Graphs to AI architectures, RL, &…

X AI KOLs Timeline · 6h ago Cached

The posters for PyTorch Con North America are now available, showcasing research in areas like compiler optimization, CUDA Graphs, AI architectures, and reinforcement learning. The conference will be held in San Jose, CA on October 20-21, with early registration discounts.

0 favorites 0 likes
#cuda-graphs

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

arXiv cs.AI · 2026-07-24 Cached

SonicSampler presents a unified suite of tile-aware Triton kernels that vertically fuse the entire LLM sampling pipeline, supporting dynamic per-request behaviors and speculative verification, achieving up to 16x speedup over state-of-the-art baselines.

0 favorites 0 likes
#cuda-graphs

@JaydevTonde: https://x.com/JaydevTonde/status/2068361821002846418

X AI KOLs Timeline · 2026-06-20 Cached

A detailed tutorial on implementing CUDA Graphs in an LLM inference server Tokn, covering FastAPI server setup, engine initialization, and CUDA Graph capture for optimized decode phases.

0 favorites 0 likes
#cuda-graphs

Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode

Hugging Face Daily Papers · 2026-05-28 Cached

This paper investigates the performance gap in batch-1 LLM decode for physical AI systems, finding that faster memory bandwidth does not proportionally reduce latency due to launch overheads, and that quantization efficiency varies significantly across hardware.

0 favorites 0 likes
← Back to home

Submit Feedback