cuda-graph

Tag

Cards List
#cuda-graph

Enable CUDA graph for MTP draft by gaugarg-nv · Pull Request #28549 · ggml-org/llama.cpp

Reddit r/LocalLLaMA ↗ · 2026-09-16 Cached

This pull request enables CUDA graph for MTP draft in llama.cpp to improve performance of LLM inference on GPUs.

0 favorites 0 likes
#cuda-graph

DominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decoding

arXiv cs.CL ↗ · 2026-07-10 Cached

DominoTree introduces a training-free best-first draft tree for speculative decoding that uses conditional (non-factorized) correction from Domino to achieve up to 6.6x speedup over autoregressive decoding and the highest mean accept length across evaluated methods on Qwen3 models.

0 favorites 0 likes
#cuda-graph

Someone just ran a 744B parameter model at 30 tok/s across 6 consumer GPUs in 6 different US states over the open internet

Reddit r/ArtificialInteligence ↗ · 2026-06-20

A researcher debuted Shard, achieving 30 tok/s inference on a 744B parameter model distributed across 6 consumer GPUs over the open internet, a 15-20x improvement over previous methods.

0 favorites 0 likes
#cuda-graph

Every AI researcher should grasp inference acceleration—CUDA Graph is the heart of vLLM's GPU efficiency

X AI KOLs Timeline ↗ · 2026-04-21 Cached

A tweet urging AI researchers to learn inference-acceleration basics and spotlighting CUDA Graph as the key to vLLM’s GPU utilization.

0 favorites 0 likes
← Back to home

Submit Feedback